Instruction scheduling method, device and system, product and medium
By obtaining the data dependency and execution delay of the instruction, combined with the delay cache predicting dynamic delay, the accurate scheduling of the instruction transmission time is achieved, the problem of low instruction scheduling efficiency in the prior art is solved, and the processor performance is improved.
Patent Information
- Application Number
- CN202510600780.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing out-of-order scheduling strategies cannot accurately schedule dynamic changes in command transmission time, resulting in waste of resources, pipeline pauses, and inability to fully utilize processor performance.
By obtaining the data dependencies and instruction execution delays of each instruction, the predicted transmission time is determined, and the execution order of the scheduling instructions is scheduled based on the predicted transmission time, data dependencies and hardware resource availability information. Dynamic delay is predicted by historical delay information in the delay cache.
Accurate prediction of instruction transmission time is achieved, the execution efficiency of processor instructions is improved, pipeline stagnation time is reduced, and processor performance is fully utilized.
Smart Images

Figure CN120104191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of processor design, and in particular to an instruction scheduling method, device, system, product and medium. Background Art
[0002] In modern processor design, instruction scheduling efficiency directly affects processor performance, especially in high-performance processor design, where the scheduling mechanism directly determines pipeline utilization and system throughput. Currently, the mainstream design scheme is to adopt an out-of-order execution mechanism, dynamically scheduling instructions through an out-of-order scheduling strategy to fill pipeline gaps, thereby significantly improving execution efficiency.
[0003] However, most common out-of-order scheduling strategies build scheduling orders based on static delays or dependencies, without considering dynamic changes in instruction issuance time. In particular, for scenarios with uncertain delays such as memory access delays and cache misses, current scheduling strategies are difficult to accurately predict the best issuance time, resulting in resource waste, pipeline stalls, and processor performance that cannot be fully utilized.
[0004] Therefore, technicians in this field are in urgent need of an instruction scheduling method to solve the problem that the current out-of-order scheduling strategy cannot accurately schedule the dynamic changes in instruction issuance time and the scheduling effect needs to be further improved. Summary of the invention
[0005] The purpose of the present invention is to provide an instruction scheduling method, device, system, product and medium to further improve the execution efficiency of processor instructions and reduce pipeline stalls.
[0006] In order to solve the above technical problems, the present invention provides an instruction scheduling method, comprising:
[0007] Get the data dependencies and instruction execution latency of each instruction.
[0008] A predicted issue time of each of the instructions is determined according to the data dependency and the instruction execution delay.
[0009] Hardware resource availability information is obtained, and the execution order of each of the instructions is scheduled based on the predicted emission time, the data dependency and the hardware resource availability information.
[0010] The types of the instruction execution delay include dynamic delay and static delay; when the instruction execution delay is dynamic delay, obtaining the instruction execution delay includes:
[0011] The historical delay information corresponding to the current instruction is obtained from the delay cache, and the dynamic delay is predicted according to the historical delay information.
[0012] The delay cache is used to record the actual delay time of each of the instructions in the past as the historical delay information.
[0013] In a possible embodiment, obtaining the instruction execution delay includes:
[0014] It is determined whether the instruction type of the instruction is a load instruction or a non-load instruction.
[0015] If it is a load instruction, the status of the data accessed by the instruction in the cache is queried.
[0016] If the state is a cache hit, a static delay corresponding to the instruction is obtained, and the static delay is used as the instruction execution delay.
[0017] If the state is a cache miss, the corresponding historical delay information in the delay cache is queried, a dynamic delay is predicted based on the historical delay information, and the dynamic delay is used as the instruction execution delay.
[0018] If it is a non-load instruction, a static delay corresponding to the instruction is obtained, and the static delay is used as the instruction execution delay.
[0019] In a possible embodiment, determining the predicted issuance time of each instruction according to the data dependency and the instruction execution delay includes:
[0020] A data dependency chain between the instructions is determined according to the data dependency relationship, and the first instruction in each of the data dependency chains is used as a producer instruction, and the remaining instructions are used as consumer instructions.
[0021] For the producer instruction, the predicted transmission time is the current time.
[0022] For the consumer instruction, the predicted transmission time is the maximum value of the instruction completion times corresponding to the predecessor instructions of the current consumer instruction.
[0023] The preceding level instructions are all the instructions in the data dependency chain that are located one level before the current consumer instruction; and the instruction completion time is determined by the predicted emission time and the instruction execution delay.
[0024] In a possible embodiment, for the producer instruction, after determining the predicted transmission time, the method further includes:
[0025] It is determined whether the instruction execution delay corresponding to the producer instruction is the static delay or the dynamic delay.
[0026] If it is the dynamic delay, the instruction completion time is determined according to the predicted issuance time, the dynamic delay and a stability weighting factor.
[0027] In a possible embodiment, determining the instruction completion time according to the predicted emission time, the dynamic delay and the stability weighting factor includes:
[0028] A margin stabilization delay is determined based on the product of the stability weighting factor and the dynamic delay.
[0029] The instruction completion time is determined according to the sum of the margin stabilization delay and the predicted issue time.
[0030] The stability weighting factor is greater than or equal to 1, and the smaller the delay fluctuation corresponding to the instruction is, the closer the value of the stability weighting factor is to 1.
[0031] In a possible embodiment, the historical delay information includes a plurality of the actual delay times;
[0032] Then, before determining the margin stabilization delay, the method further includes:
[0033] The average and range of the actual delay times are determined.
[0034] The stability weighting factor is determined according to the ratio of the range to the average value.
[0035] In a possible embodiment, the method further includes:
[0036] When the command is issued, the passing time counter starts counting.
[0037] When the instruction is completed, the counting of the time counter is stopped, and the current count value is obtained as the actual delay time of the instruction and written into the delay cache.
[0038] In a possible embodiment, the delay cache retains at most N actual delay times for the same instruction, where N is any integer greater than 1.
[0039] When a new actual delay time is written into the instruction for which N actual delay times are retained in the delay cache, the earliest actual delay time written into the delay cache corresponding to the instruction is replaced.
[0040] In a possible embodiment, when writing the actual delay time into the delay buffer, the method further includes:
[0041] Determine whether the delay buffer is full.
[0042] If so, based on a least recently used algorithm, the earliest written one of all the actual delay times corresponding to the least frequently used instruction is deleted.
[0043] The actual delay time is written into an idle position, and a mapping relationship between the actual delay time and the corresponding instruction is established.
[0044] In a possible embodiment, scheduling the execution order of each of the instructions based on the predicted transmission time, the data dependency, and the hardware resource availability information includes:
[0045] The order of each instruction in the priority queue is adjusted according to the order of the predicted emission time, wherein the instruction with an earlier predicted emission time has a higher order in the priority queue.
[0046] For the highest-ranked instruction in the priority queue:
[0047] It is determined whether the instruction has data dependency according to the data dependency relationship.
[0048] If so, the order of the instruction in the priority queue is lowered.
[0049] If not, it is determined whether the hardware resource corresponding to the instruction is available according to the hardware resource availability information.
[0050] If not available, the order of the instruction in the priority queue is lowered.
[0051] If available, the instruction is transmitted.
[0052] In a possible embodiment, before determining whether the instruction has data dependency according to the data dependency relationship, the method further includes:
[0053] It is determined whether the instruction type of the instruction is a load instruction.
[0054] If so, it is determined whether the actual execution time of the load instruction at the current moment reaches the corresponding instruction execution delay.
[0055] If not reached, wait until the corresponding instruction execution delay is reached.
[0056] If so, the step proceeds to the step of determining whether the instruction has data dependency according to the data dependency relationship.
[0057] In order to solve the above technical problems, the present invention further provides an instruction scheduling device, comprising:
[0058] The acquisition module is used to obtain the data dependency and instruction execution delay of each instruction.
[0059] A prediction module is used to determine a predicted issuance time of each of the instructions according to the data dependency and the instruction execution delay.
[0060] The scheduling module is used to obtain hardware resource availability information and schedule the execution order of each of the instructions based on the predicted transmission time, the data dependency and the hardware resource availability information.
[0061] The types of the instruction execution delay include dynamic delay and static delay, and the acquisition module includes a dynamic delay acquisition unit, and the dynamic delay acquisition unit is used to acquire the instruction execution delay of the dynamic delay type.
[0062] The dynamic delay acquisition unit is further used to: acquire historical delay information corresponding to the current instruction from the delay cache, and predict the dynamic delay according to the historical delay information.
[0063] The delay cache is used to record the actual delay time of each instruction in the past as the historical delay information.
[0064] In order to solve the above technical problems, the present invention further provides an instruction scheduling system, comprising:
[0065] Memory, used to store computer programs.
[0066] A processor is used to implement the steps of the instruction scheduling method described above when executing the computer program.
[0067] In order to solve the above technical problem, the present invention also provides a computer program product, including a computer program / instruction, which implements the steps of the instruction scheduling method described above when executed by a processor.
[0068] In order to solve the above technical problem, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the instruction scheduling method described above are implemented.
[0069] An instruction scheduling method provided by the present invention provides a dynamic delay prediction mechanism for dynamic delays in instruction execution delays, where the specific delay time cannot be determined in advance. Specifically, the historical delay information corresponding to each instruction is stored in a delay cache. When the instruction execution delay corresponding to the instruction is a dynamic delay, the dynamic delay is predicted through the historical delay information stored in the delay cache, thereby obtaining an accurate dynamic delay for subsequent instruction scheduling. Based on the above-mentioned dynamic delay prediction mechanism, this method incorporates dynamic delays whose delay time cannot be determined in advance into the consideration scope of instruction scheduling. When instruction scheduling is performed based on this method, for dynamically changing scenarios in instruction emission time such as memory access delays and cache misses, the optimal emission timing of instructions can also be accurately predicted, thereby further improving the effect of instruction scheduling on processor performance improvement.
[0070] The instruction scheduling device, system, computer program product and computer-readable storage medium provided by the present invention correspond to the above method and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0072] Figure 1 A flowchart of an instruction scheduling method provided by the present invention.
[0073] Figure 2 A scheduling flow chart of a load instruction provided by the present invention.
[0074] Figure 3 A scheduling flow chart of a non-load instruction provided by the present invention.
[0075] Figure 4 A control flow chart of a delay cache provided by the present invention.
[0076] Figure 5 A structural diagram of an instruction scheduling architecture provided by the present invention.
[0077] Figure 6 A structural diagram of an instruction scheduling system provided by the present invention. DETAILED DESCRIPTION
[0078] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0079] The core of the present invention is to provide an instruction scheduling method, device, system, product and medium.
[0080] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0081] In modern processor design, out-of-order execution mechanism is a mainstream design solution. The out-of-order execution mechanism can disrupt the execution order of instructions in the pipeline and dynamically schedule instructions to fill pipeline gaps, thereby significantly improving execution efficiency. Specifically, in the processor's instruction scheduling, the instruction issuance time depends on three major factors: data dependency, hardware resource availability, and instruction execution latency.
[0082] The data dependency refers to the order dependency between instructions due to the read-write relationship between operands. Only when the dependent instruction is executed and the corresponding operand is valid, the instruction that depends on the operand can be issued for execution.
[0083] Hardware resource availability refers to whether the hardware units (such as execution units) required for instruction execution are idle in the current cycle. If the resources are not available, the instruction cannot be issued for execution even if the data dependencies are met.
[0084] Instruction execution delay refers to the length of time required for an instruction to be issued and completed (generally expressed in terms of the number of cycles), which is affected by factors such as the type of instruction and whether the cache hits. Among them, instruction execution delay can be further divided into static delay and dynamic delay. Static delay refers to a fixed delay that can be determined during the processor design phase, such as the delay in loading instructions when the cache hits. Dynamic delay refers to a delay in which the execution time of an instruction is affected by changes in the cache state or memory hierarchy at runtime, resulting in an execution delay that cannot be determined at the design or compilation phase, and generally needs to be determined at runtime based on actual conditions.
[0085] Based on the above, it can be seen that since it is impossible to obtain accurate dynamic delays during the processor design phase, and in actual applications, it is also impossible to obtain accurate dynamic delays before the instruction is executed. Therefore, the current out-of-order execution mechanisms mostly build scheduling sequences based on static delays or dependencies, without considering the dynamic changes in instruction emission time. Especially in scenarios such as memory access delays and cache misses, the current scheduling strategy is difficult to accurately predict the best emission time and perform corresponding instruction scheduling, resulting in resource waste, pipeline pauses, and the inability to fully utilize processor performance.
[0086] To solve the above problems, the present invention provides an instruction scheduling method, such as Figure 1 As shown, including:
[0087] S1: Get the data dependency and instruction execution delay of each instruction.
[0088] S2: Determine the predicted issuance time of each instruction based on data dependency and instruction execution latency.
[0089] S3: Obtain hardware resource availability information, and schedule the execution order of each instruction based on the predicted launch time, data dependency and hardware resource availability information.
[0090] The types of instruction execution delay include dynamic delay and static delay. When the instruction execution delay is dynamic delay, obtaining the instruction execution delay includes:
[0091] The historical delay information corresponding to the current instruction is obtained from the delay cache, and the dynamic delay is predicted based on the historical delay information; the delay cache is used to record the actual delay time of each instruction in the past as the historical delay information.
[0092] Step S1 is the key step of this method, which accurately predicts the dynamic delay that cannot be accurately determined in the design or compilation stage, so as to include these instructions in the instruction scheduling range, thereby achieving a wider range, more comprehensive and more accurate instruction scheduling, thereby better improving the execution efficiency of the processor pipeline and reducing the pipeline stagnation time.
[0093] For step S1, the focus of this method is how to obtain dynamic delays that cannot be accurately determined in the design or compilation stage. In this method, a delay cache mechanism is introduced. The delay cache is a part of the area in the cache that is specially divided for storing information related to delays. Specifically, the delay cache in this method is used to store historical delay information of instructions. The historical delay information corresponding to a certain instruction may include one or more actual delay times when the instruction was executed in the past, which is used to characterize the delay of the instruction in history. It should be noted that although this method does not limit which instructions are specifically recorded in the delay cache for historical delay information. However, from the above description of instruction execution delay, it can be seen that static delay is a fixed delay that can be determined in the design stage, and there is no need to record and predict historical delay information through a delay cache. Therefore, the instruction object for recording historical delay information in the delay cache in this method refers to some instructions whose instruction execution delay is a dynamic delay, so as to reduce the requirements for delay cache capacity and improve scheduling efficiency.
[0094] Furthermore, this embodiment does not limit how to predict dynamic delay based on the historical delay information recorded in the delay cache. In a simpler implementation scheme, the historical delay information of each instruction only contains one actual delay time in the past. In this instruction scheduling, the actual delay time can be directly used as the dynamic delay predicted for this time to participate in the instruction scheduling. It is not difficult to understand that other processing can also be performed on this actual delay time to achieve more accurate dynamic delay prediction. Moreover, when the historical delay information contains multiple actual delay times, other schemes for predicting dynamic delays can also be used, such as averaging multiple actual delay times, weighted summation, etc., and this embodiment does not limit this.
[0095] As for steps S2 and S3, it can be seen from the description of the above embodiment that the key to achieving accurate and efficient instruction scheduling is to accurately predict the optimal launch time of each instruction. The optimal launch time of the instruction is related to three major factors: data dependency, hardware resource availability information and instruction execution delay. Among them, the main factors that determine whether the instruction can be launched are whether the operands on which the instruction depends are satisfied and whether the hardware units required for the instruction execution are idle. Both of the above conditions can be known through real-time query. However, whether the data dependency is satisfied can be predicted through data dependency and instruction execution delay. According to the data dependency, it can be determined which instructions correspond to the operands on which each instruction depends. This part of the instructions is the instruction on which the current instruction depends, and this part of the instructions must be located at the previous level of the current instruction on the data dependency chain (a relationship chain determined based on the data dependency, and the execution of any level of instructions in the relationship chain directly depends on the operands of the previous level of instructions), so it is called the previous level instruction.
[0096] Based on data dependency and instruction execution delay, the expected completion time of all previous instructions of any instruction can be predicted, that is, the expected time for the instruction to meet data dependency. Therefore, step S2 in this method first determines the expected time for each instruction to meet data dependency through data dependency and instruction execution delay, that is, predicts the emission time, and provides a preliminary parameter basis for the scheduling of each instruction. Based on the predicted emission time, each instruction can be preliminarily scheduled before actually meeting data dependency and hardware resource availability, which is conducive to improving the accuracy and efficiency of subsequent actual instruction scheduling. In addition, it is too rough to schedule instructions based only on whether data dependency is met and whether hardware resources are available. Both of them only have two situations of satisfaction and non-satisfaction, and the reference features brought are too few for realizing the scheduling of massive instructions in the processor. The expected completion time determined in step S2 of this method is a specific time value, which can contain more data features, which is conducive to more accurate scheduling of massive instructions in the processor.
[0097] However, it should be noted that the data dependency satisfaction time represented by the predicted launch time is only an expected time after all. In actual applications, it cannot be guaranteed that the instruction will satisfy the data dependency when the predicted launch time is reached. Moreover, the predicted launch time cannot predict when the instruction will meet the availability of hardware resources. Therefore, when actually performing instruction scheduling in step S3, the predicted launch time is only one of the reference factors. It is also necessary to judge whether the instruction truly satisfies the data dependency based on the data dependency relationship, and to judge whether the instruction truly satisfies the hardware resource availability based on the hardware resource availability information. Because the availability of the hardware resources required for instruction execution changes in real time, the hardware resource availability information obtained in advance is useless. This is why this method does not obtain the hardware resource availability information until step S3.
[0098] However, it should be noted that, as can be seen from the above description of step S1, the focus of this method is to accurately predict the dynamic delay that was originally not considered in the instruction scheduling, so as to accurately predict the best time to issue these instructions and achieve more comprehensive instruction scheduling. That is, this effect can be achieved based on step S1 provided above in this method.
[0099] That is, the specific implementation of steps S2 and S3 in this embodiment is only a possible implementation scheme, which provides more comprehensive and rich data support for instruction scheduling by predicting the emission time, which is conducive to improving the efficiency and accuracy of instruction scheduling. In practical applications, other currently mature instruction scheduling schemes can also be adopted. As long as the dynamic delay prediction method provided in step S1 of this method is adopted, better instruction scheduling effect can be achieved.
[0100] As can be seen from the above, an instruction scheduling method provided by the present application introduces a delay cache mechanism to record the actual delay time in history for some instructions whose instruction execution delay is a dynamic delay in the delay cache. Furthermore, when these instructions are subsequently involved in instruction scheduling, their dynamic delays can be predicted based on the historical delay information recorded in the delay cache, and the optimal emission timing of these instructions can be accurately predicted, thereby incorporating these instructions into the instruction scheduling range. The instruction scheduling implemented based on this method can cover a wider range of instructions in the pipeline processing process of the processor, especially for scenarios where the instruction emission time changes dynamically, such as memory access delays and cache misses, and can also accurately predict the optimal emission timing, which can better utilize processor resources and reduce pipeline pauses, so that the performance of the processor can be more fully utilized.
[0101] On the other hand, in the above embodiment, it is mainly explained how to obtain the dynamic delay of the two types of instruction execution delay. Then, for how to determine whether the instruction execution delay corresponding to the instruction is a dynamic delay or a static delay, and how to obtain the static delay, this embodiment provides a possible implementation scheme. The above step S1 specifically includes:
[0102] S11: Determine whether the instruction type of the instruction is a load instruction or a non-load instruction.
[0103] S12: If it is a load instruction, query the status of the data accessed by the instruction in the cache.
[0104] S13: If the status is a cache hit, a static delay corresponding to the instruction is obtained, and the static delay is used as the instruction execution delay.
[0105] S14: If the status is a cache miss, query the corresponding historical delay information in the delay cache, predict the dynamic delay based on the historical delay information, and use the dynamic delay as the instruction execution delay.
[0106] S15: If it is a non-load instruction, obtain the static delay corresponding to the instruction, and use the static delay as the instruction execution delay.
[0107] First of all, it should be noted that the above-mentioned load instruction refers to an instruction that needs to read data from the cache or main memory, and the delay uncertainty is high, and there is a possibility that the corresponding instruction execution delay is a dynamic delay. Non-load instructions include arithmetic, logic and control operation instructions, and the instruction delay is relatively stable, and the corresponding instruction execution delay generally belongs to static delay. Therefore, in this embodiment, it can be determined whether the type of instruction belongs to a load instruction in the decoding stage; if it is a non-load instruction, the instruction execution delay of the instruction can be directly considered to be a static delay, and the corresponding fixed delay time can be obtained in advance; and if it is a load instruction, there is a possibility that the instruction execution delay is a dynamic delay. At this time, this embodiment further determines the delay type by querying the state of the data accessed by the instruction in the cache; if the cache hits, it means that the delay type of the instruction is a static delay, and the predetermined fixed delay time can be directly obtained; if the cache misses, it means that the delay type of the instruction is a dynamic delay, and the corresponding historical delay information needs to be queried from the delay cache, and a specific dynamic delay is predicted based on the historical delay information as the instruction execution delay of the instruction.
[0108] It should also be noted that this embodiment mainly provides further possible implementation plans for the part related to obtaining instruction execution delay in step S1, and does not impose any restrictions on the part related to obtaining data dependency and hardware resource availability information in step S1. The above steps S11~S15 do not represent all steps S1.
[0109] It can be seen from the above method flow that in this embodiment, by determining the instruction type, all instructions are divided into two categories: loading instructions and non-loading instructions, and the instruction execution delay corresponding to the instruction is preliminarily distinguished to see whether there is a possibility of dynamic delay. For loading instructions, there is a possibility of dynamic delay, so it is further determined whether it is a dynamic delay by judging the cache status of the data accessed by the instruction; if the cache hits, it is a static delay; if the cache misses, it is a dynamic delay, and it is necessary to predict the specific delay time based on the delay cache mechanism. This embodiment provides a two-layer query scheme based on instruction type and cache status to determine whether the delay type corresponding to the instruction is a dynamic delay or a static delay, and obtain the specific delay time. In addition, the two-layer query scheme provided in this embodiment uses the instruction type that can be determined in the decoding stage to reduce the number of subsequent steps of querying the cache status. For non-loading instructions with relatively stable and fixed delays, it can directly determine that its instruction execution delay is a static delay without querying the cache status, reducing additional query time to improve overall efficiency.
[0110] On the other hand, it can be seen from the above description that step S2 of the method predicts the time when each instruction is expected to meet the data dependency based on the data dependency relationship and the instruction execution delay, that is, obtains the predicted emission time. Based on the predicted emission time, preliminary scheduling can be performed before the instruction actually meets the data dependency and the hardware resource availability, and data support can be provided for instruction scheduling from other aspects, which is conducive to improving the accuracy and efficiency of instruction scheduling.
[0111] Therefore, this embodiment also provides a possible implementation scheme for determining the predicted transmission time in the above step S2. The above step S2 specifically includes:
[0112] S21: Determine the data dependency chain between the instructions according to the data dependency relationship, and use the first instruction in each data dependency chain as the producer instruction and the remaining instructions as the consumer instructions.
[0113] S22: For producer instructions, the predicted transmission time is the current time.
[0114] S23: For a consumer instruction, the predicted transmission time is the maximum value of the instruction completion times corresponding to the predecessor instructions of the current consumer instruction.
[0115] The previous level instructions are all instructions in the data dependency chain that are located one level before the current consumer instruction, and the instruction completion time is determined by the predicted emission time and the instruction execution delay.
[0116] In this embodiment, based on the position of the instruction in the data dependency chain, the instructions are divided into two different types: producer instructions and consumer instructions. The above-mentioned producer instructions refer to instructions that generate operands for use by other instructions, usually load write instructions or write back instructions. Consumer instructions refer to instructions that rely on the results of producer instructions as input, and their issuance time is constrained by the completion time of the producer instructions.
[0117] It should be noted that the above consumer instructions' dependence on producer instructions is not necessarily a direct dependence, but may be an indirect dependence with other consumer instructions in between. For example, assume that there is a simple data dependency chain of "instruction 1→instruction 2→instruction 3". Among them, the execution of instruction 3 depends on the operand of instruction 2, and the execution of instruction 2 depends on the operand of instruction 1, and the execution of instruction 1 does not depend on the operand of any other instruction. Then instruction 1 is the producer instruction, instructions 2 and 3 are both consumer instructions, and instruction 1 is the predecessor instruction of instruction 2, and instruction 2 is the predecessor instruction of instruction 3. It should also be noted that the above example is only one possible situation. In fact, the predecessor instruction of an instruction is not necessarily unique, and multiple predecessor instructions may exist at the same time.
[0118] Therefore, as in the above example, step S21 can establish a data dependency chain based on the data dependency relationship. It is not difficult to understand that the data dependency chain is established for a single instruction or a small number of instructions. When the data dependency chains corresponding to all instructions are determined, a data dependency network can be further formed. However, considering that the predicted emission time is determined in units of a single instruction, it is only necessary to determine the data dependency chain corresponding to the instruction. And the data dependency chain that determines the predicted emission time of the instruction only needs to be determined to the end of the instruction.
[0119] Afterwards, in step S22, the predicted emission time is determined for the producer instruction defined above. It is not difficult to understand that since the producer instruction does not depend on the operands of any other instructions, it satisfies the data dependency at any time. Therefore, step S22 takes the earliest time that can be obtained at present, that is, the current time, as the predicted emission time of the producer instruction.
[0120] For step S23, the predicted emission time is determined for the consumer instruction defined above. As can be seen from the above, when the consumer instruction satisfies the data dependency depends on whether all its previous instructions have been executed. Since the instruction completion time (moment) can be simply considered as the instruction emission time (moment) + instruction execution delay. Therefore, the predicted emission time of each instruction can be determined by the following formulas:
[0121] (1);
[0122] Where TComplete(p) represents the instruction completion time of producer instruction p; TIssue(p) represents the instruction emission time of producer instruction p (not equal to the predicted emission time); TDelay(p) represents the instruction execution delay of producer instruction p.
[0123] It should be noted that the producer instructions and consumer instructions defined in this embodiment do not correspond to the load instructions and non-load instructions in the above embodiments. The producer instructions can be load instructions or non-load instructions, and the corresponding instruction execution delay can be static delay or dynamic delay, which can be obtained through the instruction execution delay acquisition scheme provided in the above embodiments, and this embodiment does not limit this.
[0124] (2);
[0125] In the formula, TPredicted(c) represents the predicted transmission time of consumer instruction c; TComplete(i) represents the instruction completion time of the predecessor instruction i of consumer instruction c; N represents the total number of all predecessor instructions i corresponding to consumer instruction c. It is not difficult to understand that the predecessor instruction of a consumer instruction may be a producer instruction or another consumer instruction.
[0126] Based on this embodiment, the predicted issuance time of each instruction can be determined based on the data dependency relationship, thereby providing data support for subsequent instruction scheduling from another perspective, which is beneficial to improving the efficiency and accuracy of instruction scheduling.
[0127] Further, it can be seen from the above embodiment that when the instruction execution delay corresponding to the producer instruction is a dynamic delay, it is necessary to obtain the corresponding historical delay information from the delay cache to predict the dynamic delay. Considering the large volatility of the dynamic delay, in order to more accurately predict the predicted emission time of the instruction, this embodiment also provides a further implementation scheme. After step S22, the above method also includes:
[0128] S24: determining whether the instruction execution delay corresponding to the producer instruction is a static delay or a dynamic delay;
[0129] S25: If it is a dynamic delay, determine the instruction completion time according to the predicted transmission time, the dynamic delay and the stability weighting factor.
[0130] That is, in this embodiment, the predicted dynamic delay is further corrected by the stability weighting factor. The stability weighting factor is used to characterize the degree of delay fluctuation of the instruction during the historical execution process. Based on the stability weighting factor, if the delay fluctuation of the instruction is large, the dynamic delay can be appropriately adjusted upward to increase the transmission waiting time of the instruction, thereby reducing the risk of prediction failure.
[0131] Furthermore, this embodiment provides a possible implementation scheme for the value of the stability weighting factor and the specific method of adjusting the dynamic delay. The above step S25 specifically includes:
[0132] S251: Determine a margin stability delay according to the product of the stability weighting factor and the dynamic delay.
[0133] S252: Determine the instruction completion time according to the sum of the margin stabilization delay and the predicted issuance time.
[0134] The stability weighting factor is greater than or equal to 1, and the smaller the delay fluctuation of the corresponding instruction is, the closer the value of the stability weighting factor is to 1.
[0135] That is, the above formula 1 becomes the following formula (3):
[0136] (3);
[0137] Where β(p)·TDelay(p) represents the margin stabilization delay; β(p) represents the stability weighting factor of the producer instruction p.
[0138] In this embodiment, the stability weighting factor is used to reflect the degree of fluctuation of instruction delay in historical execution. When the delay fluctuation is relatively stable, the dynamic delay predicted based on the historical delay information recorded in the delay cache has a higher credibility, so the closer the value of the stability weighting factor is to 1, the less it is necessary to increase the dynamic delay. However, when the delay fluctuation is large, it means that the credibility of the predicted dynamic delay is low. At this time, it is necessary to appropriately increase the predicted dynamic delay as the instruction execution delay during instruction scheduling, that is, the value of the stability weighting factor is appropriately adjusted to reduce the risk of dynamic delay prediction failure by increasing the transmission waiting time of the instruction.
[0139] Furthermore, regarding how to determine the value of the stability weighting factor in the above embodiment, the above embodiment has already explained that the value of the stability weighting factor is related to the magnitude of the historical delay fluctuation of the instruction. The historical delay fluctuation of the instruction can be reflected by the historical delay information recorded in the delay cache, so the stability weighting factor can be determined by the historical delay information.
[0140] However, this embodiment also provides a possible stability weighting factor value solution:
[0141] The historical delay information includes multiple actual delay times.
[0142] Before step S251, the method further includes:
[0143] S250: Determine the average value and range of each actual delay time, and determine the stability weighting factor according to the ratio of the range to the average value.
[0144] Specifically, the above step S250 can be expressed by the following formula:
[0145] (4);
[0146] In the formula, γ represents a preset sensitivity coefficient, which characterizes the sensitivity of the stability weighting factor to delay fluctuations and is used to control the influence of delay fluctuations on the stability weighting factor. Indicates the range of actual delay time. Indicates the average value of the actual delay time. It reflects the relative fluctuation of historical delay. If the fluctuation is larger, it means that cache misses are more frequent, and the stability weighting factor increases.
[0147] In the above formula 4, the sensitivity coefficient is used to establish β(p) and In another possible implementation, the stable weighting factor can be determined directly by using interval mapping, as shown in Table 1 below.
[0148] Table 1 Stability weighted factor mapping table
[0149]
[0150] It is not difficult to see from Table 1 above that this embodiment directly presets several relatively reasonable and gradient-changing β(p) setting values based on experience or testing. The mapping relationship between the ranges. Once it falls into a certain range, the corresponding unique β(p) setting value can be directly determined through the mapping relationship.
[0151] It is not difficult to see that the above two solutions can both achieve the purpose of determining the stability weighting factor according to the ratio of the range to the average value as described in step S250. In practical applications, the appropriate implementation scheme can be freely selected as needed.
[0152] On the other hand, the above embodiment provides several possible implementation plans for how to implement step S2, and this embodiment also provides a more possible implementation plan for the implementation of step S3. In the above embodiment, it is explained that step S3 implements the scheduling of instructions based on three major factors: predicted emission time, data dependency, and hardware resource availability information. Among them, whether the instruction meets the data dependency and whether the hardware resource availability is met directly determines whether the instruction can be emitted. The predicted emission time represents the time when the instruction is expected to meet the data dependency, and more of a scheduling reference in the time dimension. Based on this, this embodiment provides a possible implementation plan, and the instruction scheduling in the above step S3 specifically includes:
[0153] S31: Adjust the order of each instruction in the priority queue according to the predicted issuance time.
[0154] Among them, the earlier the predicted issuance time of the instruction is, the higher the order in the priority queue is.
[0155] For the highest priority instruction in the priority queue:
[0156] S32: Determine whether the instruction has data dependency based on the data dependency relationship; if so, go to step S33; if not, go to step S34.
[0157] S33: Lower the order of the instruction in the priority queue.
[0158] S34: Determine whether the hardware resource corresponding to the instruction is available according to the hardware resource availability information; if not, go to step S35; if available, go to step S36.
[0159] S35: Lower the order of the instruction in the priority queue.
[0160] S36: Send command.
[0161] It should be noted that although both step S33 and step S35 are to lower the order of the instruction in the priority queue (i.e., the priority of emission), the specific adjustment schemes are different. In step S33, the instruction needs to be lowered because it does not meet the data dependency, so its order should be adjusted to after other instructions that meet the data dependency. Similarly, in step S35, the priority needs to be lowered because it does not meet the hardware resource availability, so its order should be adjusted to after other instructions that meet the hardware resource availability. When both the data dependency and the hardware resource availability are met, the instruction can be allowed to be emitted, and the instructions are emitted in sequence according to the order in the priority queue for execution.
[0162] This embodiment provides a specific instruction scheduling strategy, which performs preliminary scheduling of the priority of instructions based on the predicted emission time. Then, the order of each instruction in the priority queue is used to determine whether it meets the emission conditions (data dependency and hardware resource availability); if any condition is not met, its order is adjusted to be after other instructions that meet the condition; if all conditions are met, the instructions can be issued in sequence according to the order in the priority queue, thereby achieving effective scheduling of instructions.
[0163] Furthermore, from the above embodiment for step S1, it can be known that instructions can be divided into two types: load instructions and non-load instructions. Among them, the instruction execution delay is a dynamic delay, which belongs to the load instruction (but the load instruction is not necessarily a dynamic delay). And the load instruction refers to the instruction that needs to read data from the cache or main memory, so the transmission of the load instruction needs to consider whether it reads the data. If the data is not read, the instruction cannot be transmitted, that is, the scheduling and transmission judgment of the above steps S31~S36 are not required.
[0164] Based on this, this embodiment further provides an adaptive accurate inspection solution for the load instruction. Before the above step S32, the method further includes:
[0165] S37: Determine whether the instruction type of the instruction is a load instruction.
[0166] S38: If yes, determine whether the actual execution time of the load instruction at the current moment reaches the corresponding instruction execution delay; if not, wait until the corresponding instruction execution delay is reached; if reached, go to step S32.
[0167] However, it can be seen from the above embodiment of step S1 that the judgment of whether the instruction is a load instruction has been performed as early as the stage of delayed acquisition of instruction execution. Figure 2 The instruction scheduling process shown in FIG. 1 includes the entire process of obtaining instruction execution delay, instruction scheduling, and instruction issuance. Figure 2It is not difficult to see that the load instruction can still be preliminarily scheduled based on step S31. However, before making a subsequent judgment on whether to allow the instruction to be issued and updating the priority, it is necessary to ensure that the instruction has waited until the corresponding predicted issuance time is reached, and then perform an accurate check as performed in steps S32 to S36.
[0168] Similarly, for non-load instructions that do not need to read data from the cache or main memory, accurate checking can be started without waiting until the predicted issuance time. The entire instruction scheduling process for non-load instructions is as follows: Figure 3 shown.
[0169] In summary, the whole process of implementing instruction scheduling by the method is obtained. In this embodiment, instructions are divided into load instructions and non-load instructions and different instruction scheduling processes are provided. For load instructions that need to read data from the cache or main memory, before performing a precise check, it is first determined whether the load instruction reaches the predicted emission time; if it reaches, resources are consumed for a precise check; if it does not reach, wait until it reaches, and no resources are wasted for a precise check during this period, so as to further improve the efficiency of instruction scheduling.
[0170] On the other hand, it can be seen from the above that one of the keys to improve the efficiency and accuracy of instruction scheduling is the accurate prediction of dynamic delay. And the prediction of dynamic delay is realized by delay cache. Delay cache records the actual execution time of instructions in the historical execution process as historical delay information, which is used for dynamic delay prediction in the subsequent instruction scheduling process.
[0171] Therefore, this embodiment provides a possible implementation scheme for how the delay cache implements the recording of historical delay information. The above method also includes:
[0172] S41: When the instruction is sent, the time counter starts counting.
[0173] S42: When the instruction is completed, stop counting the time counter, obtain the current count value as the actual delay time of the instruction and write it into the delay cache.
[0174] That is, in this embodiment, the time counter starts counting after the instruction is issued and stops counting after the instruction is executed, so as to obtain the actual delay time of the instruction and record it in the delay cache. Then, when the dynamic delay of the instruction needs to be predicted in the subsequent instruction scheduling process, the dynamic delay can be accurately predicted based on the actual delay time recorded in the delay cache.
[0175] In addition, this embodiment also provides a possible implementation scheme for the delay cache structure:
[0176] The delay cache uses a direct-mapped memory structure. The data structure is shown in Table 2, which records the instruction number (ID), delay cycle (that is, the actual delay time is expressed in cycles), and valid bits.
[0177] Table 2 Delay cache data structure
[0178]
[0179] The instruction ID is used to uniquely identify an instruction (load instruction), and the program counter (PC) value is usually used. The execution time counter (i.e., the time counter used to count the instruction execution time in steps S41 and S42) records the number of execution delay cycles of the instruction in the last N times (N=5 is used as an example in Table 2), that is, the actual delay time. The valid bit is used to indicate the validity of the record, and is set to 1 only when the instruction cache misses, and the delay value is not recorded when it hits.
[0180] Furthermore, it can be seen from the above embodiments that when instructions are divided into two categories, load instructions and non-load instructions, generally only load instructions may have dynamic delays. It is not easy to determine whether the instruction execution delay of an instruction is a static delay or a dynamic delay in the decoding stage, but it can be determined whether the instruction type is a load instruction or a non-load instruction. And for a load instruction, it can be determined whether it corresponds to a dynamic delay by querying whether its access to the data cache hits.
[0181] So Figure 4 As shown, the instruction objects targeted by the delay cache control process provided by this embodiment only include load instructions. For load instructions, first query whether its cache hits. If it hits, it means that the load instruction corresponds to a static delay, and there is no need to provide historical delay information or record the historical delay information. If the cache misses, the delay cache is required to provide historical delay information. There are further two situations: first, the delay cache hits, which means that the historical delay time of the instruction exists in the delay cache and can be read out; second, the delay cache misses, which means that there is no historical delay information of the instruction in the delay cache, so it is necessary to record the historical delay information in the delay cache.
[0182] On the other hand, the above embodiment also illustrates that for each instruction in the delay cache, one or more actual delay times can be recorded as historical delay information. In practical applications, in order to improve the accuracy of dynamic delay prediction, and in some of the above embodiments, a solution for determining a stability weighting factor is provided, which also requires multiple actual delay times, so this embodiment provides a better implementation scheme:
[0183] The delay cache retains multiple actual delay times for the same instruction.
[0184] Furthermore, since the delay cache is a part of the area that is separately divided from the cache, if its capacity is too large, it will affect the normal function of the processor. When the delay cache capacity is full, it is impossible to continue to store new historical delay times for subsequent instruction scheduling. To solve this problem, this embodiment also provides a possible implementation scheme:
[0185] The delay cache retains at most N actual delay times for the same instruction, where N is any integer greater than 1. When a new actual delay time is written to an instruction for which N actual delay times are retained in the delay cache, the earliest actual delay time written in the delay cache corresponding to the instruction is replaced.
[0186] That is, in this embodiment, for the same instruction, at most N (such as 5) can be retained, and when N is reached, the earliest written actual delay time will be automatically overwritten when a new actual delay time is written. On the one hand, the data volume of the historical delay information of each instruction is controlled, and on the other hand, the automatic update of the historical delay information is realized. The historical delay information recorded in the delay cache can always reflect the actual delay situation of the delay fluctuation of the instruction in the recent period of time, which is conducive to improving the accuracy of dynamic delay prediction.
[0187] Furthermore, in addition to the automatic update solution when the historical delay information of a single instruction is written enough, this embodiment also provides a delay cache writing solution. When the actual delay time is written into the delay cache, the above method also includes:
[0188] S51: Determine whether the delay buffer is full; if so, go to step S52; if not, go to step S53.
[0189] S52: Based on the least recently used algorithm, delete the earliest written one among all actual delay times corresponding to the least frequently used instruction.
[0190] S53: Write the actual delay time into the idle position and establish a mapping relationship with the corresponding instruction.
[0191] The specific process is as follows Figure 4As shown, it can be known from the above embodiment that a delay cache miss means that there is no historical delay information of the current instruction in the delay cache, and it is necessary to time and record the actual delay time. At this time, based on step S51, it is necessary to first determine whether the delay cache is full; if it is full, corresponding to step S52, the least frequently used instruction is determined as the replacement target based on the least recently used (LRU) algorithm, and because the actual delay information is written, the earliest written one of all the actual delay times corresponding to the least frequently used instruction is used as the replacement object, and the replacement is written to ensure that the capacity of the delay cache is not exceeded to achieve data update. Further, in combination with the delay cache mapping structure provided in the corresponding embodiment of Table 2 above, when it is necessary to record the actual delay time, the ID of the instruction is recorded, and the valid bit is pulled high, and the technology is started by executing the time counter. Until the instruction is submitted, the execution time counter stops counting and records the current count value.
[0192] As can be seen from the above, this embodiment provides a replacement and update method for a delay cache when the data is full, ensuring that the dynamic delay can still be accurately predicted without requiring too much capacity for the delay cache.
[0193] On the other hand, with respect to the hardware architecture of an instruction scheduling method provided in the above embodiment, this embodiment also provides a possible implementation scheme. Figure 5 As shown, an instruction scheduling architecture 10 includes: a decoding module 11, a dependency record table 12, a delay cache 13, an emission time prediction module 14, a scheduling module 15 and a priority queue 16.
[0194] The decoding module is an existing module in the processor, which is used to decode instructions and obtain information such as instruction type and data dependency. The data dependency is recorded in the dependency record table, and when the instruction corresponds to dynamic delay and there is no historical delay information record in the delay cache, the delay cache records the historical delay information of the instruction.
[0195] The dependency record table is used to manage the dependency of physical registers in an out-of-order execution environment and maintain the ready state of registers, thereby managing the data dependency between instructions to facilitate the subsequent emission time prediction module to predict the emission time. A possible data structure of the dependency record table is shown in Table 3 below, which records the source register, the mapped physical register, and the ready flag.
[0196] Table 3 Data structure of dependency record table
[0197]
[0198] Specifically, the ready flag is used to indicate whether the data in the physical register is valid; if the ready flag is 1, it means that the value of the source register has been calculated and the instruction can be used directly; if the ready flag is 0, it means that the value of the physical register depends on the calculation result of the previous instruction, forming a read-after-write (RAW) dependency, and the subsequent instructions must wait until the previous instruction is executed before executing, that is, a real dependency. In actual execution, if a load instruction waits for the main memory to return data due to a cache miss, its corresponding target register will be in the "not ready" state for a long time, which will affect the scheduling of consumer instructions.
[0199] The emission time prediction module is a functional module for implementing the above steps S1 and S2, which obtains data dependencies from the dependency record table and predicts dynamic delays through the delay cache to determine the predicted emission time of each instruction.
[0200] The scheduling module corresponds to the above step S3. After determining the predicted emission time, it preliminarily schedules each instruction based on the predicted emission time. The priority queue is a queue that manages the execution order of each instruction. Based on the order of the instruction in the priority queue, it manages and schedules the instruction emission time and execution order.
[0201] Furthermore, in order to better illustrate an instruction scheduling method provided by the present method, this embodiment also provides a specific example in combination with a specific implementation scenario:
[0202] Assume that the processor is a single-issue processor with 13 instructions, some of which have data dependencies and include multiple load instructions. Assume that L1 cache hits delay 1 cycle, L2 cache hits delay 2 cycles, and cache misses delay 8 cycles. The issue time prediction uses equations 1 to 4. If the load instruction misses, use the correction value in the delay cache To perform delay adjustment, the instructions and instruction descriptions are shown in Table 4. Assume that instruction 1 hits the L1 cache, instruction 3 hits the L2 cache, and instruction 5 misses the cache.
[0203] Table 4 Priority queue and instruction description before adjustment
[0204]
[0205] From the above, we can see that instruction 1 hits the L1 cache and has a static delay of 1 cycle; instruction 3 hits the L2 cache and has a static delay of 2 cycles; instruction 5 does not hit the cache. Assuming that the historical delays in the delay cache are 8, 9, 7, and 8, the actual delay times (expressed in cycles) are calculated as the average value. , fluctuation range According to formula 4, take ,get , which is used to correct the execution delay to 8*1.25=10 cycles. That is, the dynamic delay of the corrected predicted instruction 5 is 10 cycles.
[0206] Furthermore, based on the above method, the instructions shown in Table 4 are scheduled to obtain the optimized Table 5.
[0207] Table 5 Adjusted priority queue and instruction description
[0208]
[0209] In the above-mentioned emission timing, it can be seen that this method optimizes the efficiency of instruction scheduling. First, the load instruction is predicted and corrected through the delay cache in the scheduling stage. Taking instruction 5 (LOAD R6) as an example, due to cache miss, its delay is no longer directly set to a constant, but combined with historical execution records, the actual delay is corrected by a weighted factor to obtain 10 cycles. The processor predicts its completion time based on this, avoiding resource waiting and execution failure caused by the early occurrence of consumer instruction R7.
[0210] Secondly, in scenarios where there is data dependency, such as instruction 2 (ADD R2), instruction 4 (MUL R5), and instruction 7 (XOR R8), the timing at which they can be issued is dynamically pushed through equations 1 and 2 to ensure that consumer instructions do not enter the execution stage too early when the operands are not ready.
[0211] Again, for non-load instructions without direct dependencies, such as instruction 11 (AND R11) and instruction 12 (NORR12), the scheduling module automatically identifies whether they can be issued in the current cycle, and flexibly inserts them during waiting for dependent data resources or resource conflicts, effectively filling pipeline gaps and improving single-cycle issuance efficiency.
[0212] In addition, in the second half of instruction scheduling, instruction R13 (NAND R13) and instruction R10 (STORE R10) depend on multiple upstream operands respectively. When calculating their predicted emission times, the system integrates the completion times of multiple producers to ensure that the final emission time meets the latest dependency readiness requirement, thereby achieving refined scheduling control and avoiding resource preemption.
[0213] At the same time, taking instruction 6 (SUB R7) as an example, its predicted launch time is the 12th cycle. This time is derived by the system based on the completion time of producer instructions R6 and R5, where R6 cache misses, and the delay is 10 cycles after weighted correction. It should be completed in the 12th cycle after being launched in the 2nd cycle; R5 is completed in the 6th cycle and is ready earlier. When the system runs to the 12th cycle, the scheduling module determines that P6 can be launched based on the prediction results, but in order to avoid scheduling failures caused by delay fluctuations or resource occupation, the scheduling module performs precise checks. Only when R6 and R5 are ready and the execution unit is idle and available, the instruction is launched. If any condition is not met, the instruction remains in the scheduling queue, and the priority is updated and waits for re-checking in the next cycle.
[0214] It can be seen that this method optimizes and improves instruction scheduling in dynamic delay change scenarios by accurately predicting instruction issuance time. While improving instruction scheduling efficiency and reducing idle cycles, it further improves the accuracy of instruction issuance time, and has certain energy efficiency optimization and system expansion functions, providing an effective optimization method for modern high-performance processor design.
[0215] In the above embodiment, an instruction scheduling method is described in detail. The present invention also provides an embodiment corresponding to an instruction scheduling device. This embodiment provides an instruction scheduling device, including:
[0216] An acquisition module, used to acquire data dependencies and instruction execution delays of each instruction;
[0217] A prediction module, for determining a predicted issuance time of each instruction based on data dependencies and instruction execution latency;
[0218] A scheduling module, used to obtain hardware resource availability information and schedule the execution order of each instruction based on the predicted launch time, data dependency and hardware resource availability information;
[0219] Wherein, the types of instruction execution delay include dynamic delay and static delay, and the acquisition module includes a dynamic delay acquisition unit, and the dynamic delay acquisition unit is used to acquire the instruction execution delay of the dynamic delay type;
[0220] The dynamic delay acquisition unit is further used to: acquire historical delay information corresponding to the current instruction from the delay cache, and predict the dynamic delay through the historical delay information;
[0221] The delay cache is used to record the actual delay time of each instruction in the past as historical delay information.
[0222] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.
[0223] In addition to the embodiment of an instruction scheduling method provided in the above embodiment, the present invention also provides an embodiment corresponding to a computer program product. A computer program product includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the instruction scheduling method described in any of the above embodiments can be implemented.
[0224] Since the embodiments of the computer program product part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the computer program product part, which will not be repeated here.
[0225] Figure 6 A structural diagram of an instruction scheduling system provided by another embodiment of the present invention is shown as follows: Figure 6 As shown, an instruction scheduling system includes: a memory 20 for storing a computer program;
[0226] The processor 21 is used to implement the steps of an instruction scheduling method in the above embodiment when executing a computer program.
[0227] The instruction scheduling system provided in this embodiment may include but is not limited to a mobile terminal, a personal computer, a workstation, etc.
[0228] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0229] The memory 20 may include one or more computer-readable storage media, which may be non-transitory. The memory 20 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201, wherein, after the computer program is loaded and executed by the processor 21, it can implement the relevant steps of an instruction scheduling method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 20 may also include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. Data 203 may include, but is not limited to, an instruction scheduling method, etc.
[0230] In some embodiments, a command scheduling system may further include a display screen 22 , an input / output interface 23 , a communication interface 24 , a power supply 25 , and a communication bus 26 .
[0231] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation of an instruction scheduling system, and may include more or less components than those shown in the figure.
[0232] An instruction scheduling system provided by an embodiment of the present invention includes a memory and a processor. When the processor executes a program stored in the memory, it can implement the following method: an instruction scheduling method.
[0233] Finally, the present invention also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.
[0234] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes.
[0235] The above is a detailed introduction to an instruction scheduling method, device, system, product and medium provided by the present invention. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referenced to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
[0236] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
Claims
1. An instruction scheduling method, characterized in that: include: Obtain the data dependencies and instruction execution delays of each instruction; Determining a predicted issue time of each of the instructions according to the data dependency and the instruction execution delay; Acquire hardware resource availability information, and schedule the execution order of each of the instructions based on the predicted transmission time, the data dependency and the hardware resource availability information; The types of the instruction execution delay include dynamic delay and static delay; when the instruction execution delay is dynamic delay, obtaining the instruction execution delay includes: Acquire historical delay information corresponding to the current instruction from a delay cache, and predict the dynamic delay based on the historical delay information; The delay cache is used to record the actual delay time of each of the instructions in the past as the historical delay information.
2. The instruction scheduling method according to claim 1, characterized in that: Obtaining the instruction execution delay includes: Determining whether the instruction type of the instruction is a load instruction or a non-load instruction; If it is a load instruction, query the status of the data accessed by the instruction in the cache; If the state is a cache hit, a static delay corresponding to the instruction is obtained, and the static delay is used as the instruction execution delay; If the state is a cache miss, querying the corresponding historical delay information in the delay cache, predicting a dynamic delay based on the historical delay information, and using the dynamic delay as the instruction execution delay; If it is a non-load instruction, a static delay corresponding to the instruction is obtained, and the static delay is used as the instruction execution delay.
3. The instruction scheduling method according to claim 1, characterized in that: Determining the predicted issuance time of each of the instructions according to the data dependency and the instruction execution delay comprises: Determine a data dependency chain between the instructions according to the data dependency relationship, and use the first instruction in each data dependency chain as a producer instruction and the remaining instructions as consumer instructions; For the producer instruction, the predicted transmission time is the current time; For the consumer instruction, the predicted transmission time is the maximum value of the instruction completion times corresponding to the predecessor instructions of the current consumer instruction; The preceding level instructions are all the instructions in the data dependency chain that are located one level before the current consumer instruction; and the instruction completion time is determined by the predicted emission time and the instruction execution delay.
4. The instruction scheduling method according to claim 3, characterized in that: For the producer instruction, after determining the predicted transmission time, the method further includes: Determining whether the instruction execution delay corresponding to the producer instruction is a static delay or a dynamic delay; If it is the dynamic delay, the instruction completion time is determined according to the predicted issuance time, the dynamic delay and a stability weighting factor.
5. The instruction scheduling method according to claim 4, characterized in that: Determining the instruction completion time according to the predicted launch time, the dynamic delay and the stability weighting factor comprises: determining a margin stabilization delay according to a product of the stability weighting factor and the dynamic delay; Determining the instruction completion time according to the sum of the margin stabilization delay and the predicted launch time; The stability weighting factor is greater than or equal to 1, and the smaller the delay fluctuation corresponding to the instruction is, the closer the value of the stability weighting factor is to 1.
6. The instruction scheduling method according to claim 5, characterized in that: The historical delay information includes a plurality of actual delay times; Then, before determining the margin stabilization delay, the method further includes: Determining the average and range of each of the actual delay times; The stability weighting factor is determined according to the ratio of the range to the average value.
7. The instruction scheduling method according to claim 1, characterized in that: The method also includes: When the instruction is transmitted, the time counter starts counting; When the instruction is completed, the counting of the time counter is stopped, and the current count value is obtained as the actual delay time of the instruction and written into the delay cache.
8. The instruction scheduling method according to claim 7, characterized in that: The delay cache retains at most N actual delay times for the same instruction, where N is any integer greater than 1; When a new actual delay time is written into the instruction for which N actual delay times are retained in the delay cache, the earliest actual delay time written into the delay cache corresponding to the instruction is replaced.
9. The instruction scheduling method according to claim 7, characterized in that: When writing the actual delay time into the delay buffer, the method further includes: Determining whether the delay buffer is full; If so, based on a least recently used algorithm, deleting the earliest written one of all the actual delay times corresponding to the least frequently used instruction; The actual delay time is written into an idle position, and a mapping relationship between the actual delay time and the corresponding instruction is established.
10. The instruction scheduling method according to claim 1, characterized in that: Scheduling the execution order of each of the instructions based on the predicted emission time, the data dependency, and the hardware resource availability information includes: Adjusting the order of each instruction in the priority queue according to the order of the predicted emission time; wherein the instruction with an earlier predicted emission time has a higher order in the priority queue; For the highest-ranked instruction in the priority queue: Determining whether the instruction has data dependency according to the data dependency relationship; If so, lower the order of the instruction in the priority queue; If not, determining whether the hardware resource corresponding to the instruction is available according to the hardware resource availability information; If not available, lowering the order of the instruction in the priority queue; If applicable, the instruction is transmitted.
11. The instruction scheduling method according to claim 10, characterized in that: Before determining whether the instruction has data dependency according to the data dependency relationship, the method further includes: Determining whether the instruction type of the instruction is a load instruction; If so, determining whether the actual execution time of the load instruction at the current moment reaches the corresponding instruction execution delay; If not reached, wait until the corresponding instruction execution delay is reached; If so, the step proceeds to the step of determining whether the instruction has data dependency according to the data dependency relationship.
12. An instruction scheduling device, characterized in that: include: An acquisition module, used to acquire data dependencies and instruction execution delays of each instruction; A prediction module, configured to determine a predicted issuance time of each of the instructions according to the data dependency and the instruction execution delay; A scheduling module, configured to obtain hardware resource availability information, and schedule the execution order of each of the instructions based on the predicted transmission time, the data dependency and the hardware resource availability information; Wherein, the types of the instruction execution delay include dynamic delay and static delay, and the acquisition module includes a dynamic delay acquisition unit, and the dynamic delay acquisition unit is used to acquire the instruction execution delay of the dynamic delay type; The dynamic delay acquisition unit is further used to: acquire historical delay information corresponding to the current instruction from the delay cache, and predict the dynamic delay through the historical delay information; The delay cache is used to record the actual delay time of each instruction in the past as the historical delay information.
13. An instruction scheduling system, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the instruction scheduling method as claimed in any one of claims 1 to 11 when executing the computer program.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the instruction scheduling method according to any one of claims 1 to 11 are implemented.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the instruction scheduling method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Method and apparatus for executing processor instructions based on a dynamically alterable delay
CN101501636A
Method and system for predicting Load instruction execution delay
CN111221579A
State reading instruction sending method and device, storage equipment and readable storage medium
CN115269468A
Instruction scheduling method and device based on timeliness priority, computer equipment and medium
CN117742795A
Instruction awakening method, device and equipment
CN117742796A
Cited By
Instruction transmitting method and device, electronic equipment and readable storage medium
CN120723312A
Instruction transmission method and device, electronic equipment and readable storage medium
CN120723312B
Tensor instruction optimization method and device and storage medium
CN121072629A
Instruction scheduling method and device, electronic equipment, storage medium and computer program product
CN121187653A