An instruction scheduling method, a computer program product, an electronic device, and a medium
By collecting instruction information and weights to generate priority information, dynamically adjusting instruction weights, and using reinforcement learning algorithms to optimize instruction scheduling, it solves the problems of high hardware costs and poor adaptability in multi-core and high-load scenarios, and improves equipment performance and throughput.
Patent Information
- Application Number
- CN202510502889.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, instruction scheduling methods are difficult to flexibly adapt to scheduling in multi-core and high-load scenarios, resulting in high hardware costs, large power consumption, and difficult to scale with complex hardware structures that are dependent on out-of-order execution.
By collecting instruction information, determining instruction weights, generating priority information, selecting target instructions for processing, and generating target reward values based on the processing information, dynamically adjusting instruction weights to adapt to different scenarios, and using reinforcement learning algorithms to optimize scheduling strategies.
It realizes flexible adaptive scheduling in multi-core and high-load scenarios, reduces pipeline pauses, optimizes key path instruction execution, improves overall equipment performance and throughput, and avoids dependence on complex hardware structures.
Smart Images

Figure CN120010928B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more particularly, to an instruction scheduling method, a computer program product, an electronic device, and a medium. Background Art
[0002] In a computer, instruction scheduling is one of the factors affecting performance. Therefore, it is necessary to optimize the instruction issue order, make full use of pipeline and execution unit resources, reduce pipeline stall time, and improve instruction throughput. For example, instructions are scheduled through out-of-order execution, that is, by dynamically tracking instruction dependencies and preferentially scheduling instructions whose operands are ready.
[0003] However, out-of-order execution relies on complex hardware structures such as reservation stations and reorder buffers, resulting in high instruction scheduling power consumption and high hardware costs, making it difficult to further expand in multi-core and high-load scenarios. In addition, with the popularization of multi-tasking and multi-threading, instruction scheduling needs to face more complex dynamic loads, but out-of-order scheduling has difficulty maintaining good adaptability.
[0004] In summary, how to flexibly perform adaptive scheduling of instructions is an urgent problem to be solved by those skilled in the art currently. Summary of the Invention
[0005] The object of the present invention is to provide an instruction scheduling method, which can, to a certain extent, solve the technical problem of how to flexibly perform adaptive scheduling of instructions. The present invention also provides a computer program product, an electronic device, and a computer-readable storage medium.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] In a first aspect, an instruction scheduling method is provided, including:
[0008] Collecting instruction information of instructions to be scheduled;
[0009] Determining the instruction weight corresponding to the instruction information currently;
[0010] Generating priority information of the instructions to be scheduled according to the instruction information and the instruction weight;
[0011] Selecting a target instruction from the instructions to be scheduled for processing according to the priority information;
[0012] Collecting processing information of processing the target instruction;
[0013] Generating a target reward value for evaluating instruction processing performance according to the processing information;
[0014] Adjust the instruction weight according to the target reward value, and process the to-be-scheduled instruction based on the adjusted instruction weight.
[0015] On the other hand, collect the instruction information of the to-be-scheduled instruction, including:
[0016] Predict the instruction delay of the to-be-scheduled instruction to obtain the estimated delay;
[0017] Collect the type information and the target path depth of the to-be-scheduled instruction;
[0018] Collect the availability information of the resources corresponding to the to-be-scheduled instruction;
[0019] Use the type information, the estimated delay, the target path depth, and the availability information as the instruction information of the to-be-scheduled instruction;
[0020] Determine the current instruction weight corresponding to the instruction information, including;
[0021] Determine the current instruction weight corresponding to the instruction information, where the instruction weight includes an instruction type weight, a smoothing factor, a path depth weight, and a resource availability weight.
[0022] On the other hand, generate the priority information of the to-be-scheduled instruction according to the instruction information and the instruction weight, including:
[0023] According to the relationship that the priority is positively correlated with the target path depth, generate path depth information based on the target path depth and the path depth weight;
[0024] Smooth the path depth information to obtain the path priority;
[0025] According to the relationship that the priority is negatively correlated with the estimated delay, generate a delay priority based on the estimated delay and the smoothing factor;
[0026] According to the relationship that the priority is positively correlated with the availability information, generate an availability priority based on the availability information and the resource availability weight;
[0027] According to the relationship that the priority is positively correlated with the instruction type weight, generate the priority information of the to-be-scheduled instruction based on the instruction type weight, the path priority, the delay priority, and the availability priority.
[0028] On the other hand, according to the relationship that the priority is positively correlated with the instruction type weight, generate the priority information of the to-be-scheduled instruction based on the instruction type weight, the path priority, the delay priority, and the availability priority, including:
[0029] According to the instruction type weight, perform a positively correlated operation on the path priority and the latency priority to generate an initial priority;
[0030] Perform a positively correlated operation on the initial priority and the availability priority to generate the priority information of the instruction to be scheduled.
[0031] On the other hand, collect and process the processing information of the target instruction, including:
[0032] Collect and process the number of successful transmissions of the target path of the target instruction;
[0033] Collect and process the number of stall cycles of the target instruction;
[0034] Collect and process the number of resource conflicts of the target instruction;
[0035] Use the number of successful transmissions of the target path, the number of stall cycles, and the number of resource conflicts as the processing information for processing the target instruction.
[0036] On the other hand, according to the processing information, generate a target reward value for evaluating the instruction processing performance, including:
[0037] Obtain the set reward weight corresponding to the processing information;
[0038] Based on the reward weight, perform a weighted operation on the processing information to obtain a target reward value for evaluating the instruction processing performance.
[0039] On the other hand, obtain the set reward weight corresponding to the processing information, including:
[0040] Obtain the set first adjustment weight for adjusting the influence of path transmission on the reward;
[0041] Obtain the set second adjustment weight for adjusting the influence of stalls on the reward;
[0042] Obtain the set third adjustment weight for adjusting the influence of resource conflicts on the reward;
[0043] Use the first adjustment weight, the second adjustment weight, and the third adjustment weight as the reward weight corresponding to the processing information.
[0044] On the other hand, based on the reward weight, perform a weighted operation on the processing information to obtain a target reward value for evaluating the instruction processing performance, including:
[0045] Determine the first product value of the first adjustment weight and the number of successful transmissions of the target path;
[0046] Determine the second product value of the second adjustment weight and the number of stall cycles;
[0047] Generate a first difference value between the first product value and the second product value;
[0048] Determine a third product value of the third adjustment weight and the number of resource conflicts;
[0049] Generate a second difference value between the first difference value and the third product value;
[0050] Use the second difference value as the target reward value for evaluating the instruction processing performance.
[0051] On the other hand, according to the target reward value, adjust the instruction weight, including:
[0052] Obtain the actual latency of the target instruction;
[0053] Detect whether the difference between the actual latency and the estimated latency is greater than a preset threshold;
[0054] In response to the difference between the actual latency and the estimated latency being greater than the preset threshold, adjust the smoothing factor and, according to the target reward value, adjust the path depth weight or the resource availability weight;
[0055] In response to the difference between the actual latency and the estimated latency being less than or equal to the preset threshold, adjust the path depth weight or the resource availability weight according to the target reward value.
[0056] On the other hand, before adjusting the path depth weight or the resource availability weight according to the target reward value, further include:
[0057] Detect whether the target reward value is less than a first set value;
[0058] In response to the target reward value being less than the first set value, perform the step of adjusting the path depth weight or the resource availability weight according to the target reward value.
[0059] On the other hand, adjusting the path depth weight or the resource availability weight according to the target reward value includes:
[0060] In response to the target reward value being less than the first set value and greater than or equal to a second set value, adjust the path depth weight.
[0061] On the other hand, adjusting the path depth weight or the resource availability weight according to the target reward value includes:
[0062] In response to the target reward value being less than the second set value, the resource availability weight is adjusted.
[0063] In a second aspect, a computer program product is provided, including computer programs / instructions which, when executed by a processor, implement the steps of any of the above-mentioned instruction scheduling methods.
[0064] In a third aspect, an electronic device is provided, including:
[0065] A memory for storing computer programs;
[0066] A processor for implementing the steps of any of the above-mentioned instruction scheduling methods when executing the computer programs.
[0067] In a fourth aspect, a computer-readable storage medium is provided, in which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any of the above-mentioned instruction scheduling methods.
[0068] An instruction scheduling method provided by the present invention includes: collecting instruction information of instructions to be scheduled; determining the instruction weight corresponding to the instruction information currently; generating priority information of the instructions to be scheduled according to the instruction information and the instruction weight; selecting a target instruction from the instructions to be scheduled for processing according to the priority information; collecting processing information of processing the target instruction; generating a target reward value for evaluating the instruction processing performance according to the processing information; adjusting the instruction weight according to the target reward value, so as to process the instructions to be scheduled based on the adjusted instruction weight.
[0069] The beneficial effects of the present invention are as follows: Priority information of the instructions to be scheduled can be generated according to the instruction information and the current instruction weight, and a target instruction is selected for processing according to the priority information, so that the instruction processing process is controlled by the instruction weight; and then a target reward value is generated according to the processing information of the target instruction, which reflects the processing performance of the target instruction from the data level, so as to realize the quantification of the processing process of the target instruction; then the instruction weight is adjusted according to the target reward value, which is equivalent to adjusting the instruction weight from the data level according to the processing performance of the target instruction, so that the adjusted instruction weight is adapted to the processing process of the target instruction. When selecting a target instruction for processing later, it is equivalent to selecting an instruction by applying the adjusted instruction weight, so as to realize the dynamic adaptive scheduling of the instructions, without relying on a complex hardware structure and having good flexibility. A computer program product, an electronic device and a computer-readable storage medium provided by the present invention also solve the corresponding technical problems. Description of the Drawings
[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0071] Figure 1 It is a flowchart of an instruction scheduling method provided by an embodiment of the present invention;
[0072] Figure 2 It is a schematic diagram of the reinforcement learning model of the processor;
[0073] Figure 3 It is a schematic diagram of the input of processor parameters;
[0074] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0075] Figure 5 It is another schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0076] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0077] Please refer to Figure 1 , Figure 1 It is a flowchart of an instruction scheduling method provided by an embodiment of the present invention.
[0078] An instruction scheduling method provided by an embodiment of the present invention may include the following steps:
[0079] Step S101: Collect the instruction information of the instruction to be scheduled.
[0080] In practical applications, when scheduling and processing instructions, it is necessary to know the corresponding information of the instructions in order to determine the processing method of the instructions according to the corresponding information of the instructions later. Therefore, it is necessary to collect the instruction information of the instructions to be scheduled. The quantity and type of the instructions to be scheduled can be flexibly determined according to the application scenario. For example, the instructions to be scheduled can be the instructions to be processed in the server or the instructions to be processed in the processor, etc.
[0081] In a specific application scenario, considering that the instruction type, instruction latency, target path depth, resource status, etc. all affect the scheduling of all instructions, during the process of collecting the instruction information of the instructions to be scheduled, the instruction latency of the instructions to be scheduled can be predicted to obtain an estimated latency, which can be the latency time when predicting a miss in the L1 cache. Collect the type information and target path depth of the instructions to be scheduled. The type information can be obtained by analyzing the instructions to be scheduled through a decoding module. For example, the type information can be a load instruction, a store instruction, an ALU (Arithmetic And Logic Unit) instruction, etc. The target path depth can be the critical path depth, which refers to the number of instructions on the dependency chain between the starting instruction that does not depend on other instructions and a certain instruction on the critical path in the data dependency relationship of the program. As shown in Table 1, it records the dependency relationships of all instructions in the instruction queue, including the critical path depth, instruction address, producer instruction address, and consumer instruction address. Collect the availability information of the resources corresponding to the instructions to be scheduled. This availability information refers to the current available status of the resources required by the instructions, such as the resources required by execution units, etc. Priority scheduling can be used to schedule instructions that can utilize idle resources to avoid resource waste. The type information, estimated latency, target path depth, and availability information are used as the instruction information of the instructions to be scheduled.
[0082] Table 1 Critical Path Depth Data Structure
[0083]
[0084] It should be noted that the reason for taking the critical path depth as the target path depth is that the critical path is unique during the program execution stage, and the critical path depth reflects the importance of the instructions in the overall program. Instructions with a large critical path depth are at the end of the dependency chain and have a greater impact on the overall program execution. Therefore, ensuring the priority scheduling of critical path instructions can reduce the latency on the critical path and improve the processor performance. So, the critical path depth can be taken as the target path depth.
[0085] Step S102: Determine the instruction weight corresponding to the current instruction information.
[0086] Step S103: Generate the priority information of the instructions to be scheduled according to the instruction information and the instruction weight.
[0087] Step S104: Select the target instruction from the instructions to be scheduled for processing according to the priority information.
[0088] In practical applications, when there are multiple instructions to be scheduled, how to select appropriate instructions for scheduling becomes a problem to be solved. During the instruction scheduling process, considering that instruction information can affect the instruction processing process. For example, if there is a dependency relationship between two instructions, the latter instruction needs to wait until the former instruction is processed before it can be executed. In other words, if the latter instruction is scheduled first at this time, it will only block the instruction processing process. Therefore, the instruction scheduling process can be controlled according to the type of instruction information. Further, the instruction weight corresponding to the instruction information can be determined, that is, corresponding instruction weights can be set for various types of instruction information. This instruction weight can be used to represent the impact of instruction information on instruction scheduling. In this way, by simply performing operations on the instruction information and the instruction weight, the priority information of the instruction to be scheduled obtained by comprehensively considering all instruction information can be generated. Subsequently, when selecting a target instruction for processing from the instructions to be scheduled according to the priority information, it is equivalent to performing instruction scheduling by comprehensively considering the impact of all instruction information on instruction scheduling, ensuring the rationality of instruction scheduling.
[0089] In a specific application scenario, during the process of selecting a target instruction for processing from the instructions to be scheduled according to the priority information, the instruction to be scheduled with the highest priority and not dependent on other instructions can be selected first as the target instruction for processing. Of course, the priority can also be flexibly applied according to the actual scenario to select the target instruction for processing.
[0090] In a specific application scenario, during the process of determining the instruction weight corresponding to the instruction information, the instruction weight corresponding to the instruction information can be determined. The instruction weight includes instruction type weight, smoothing factor, path depth weight, and resource availability weight. Among them, the instruction type weight is used to determine the priority of different types of instructions. As shown in Table 2, load instructions, prefetch instructions, floating-point operation instructions, and atomic operation instructions have relatively high weights. The reason is that such instructions have strong data dependencies and have a greater impact on the pipeline performance. Store instructions, branch instructions, multimedia instructions, and memory management instructions have medium weights. The reason is that such instructions are closely related to the scenario. Arithmetic operation instructions, logical operation instructions, and shift instructions have relatively low weights. The reason is that such instructions have less dependence on other instructions in terms of latency. The smoothing factor is used to smooth the value of the reciprocal of the emission time to prevent the priority corresponding to the instruction delay from becoming infinitely large when the predicted delay time is very small, thus causing the priority to be extremely biased towards certain instructions unreasonably. The path depth weight is used to determine the priority of different target path depths. The resource availability weight is used to represent the importance of resource availability in the calculation of scheduling priority.
[0091] Table 2 Instruction Type Weight List
[0092]
[0093] In a specific application scenario, during the process of generating the priority information of the to-be-scheduled instruction according to the instruction information and the instruction weight, the path depth information can be generated according to the relationship that the priority is positively correlated with the target path depth, based on the target path depth and the path depth weight; smooth the path depth information, such as smoothing the path depth information through a logarithmic function, etc., to obtain the path priority; according to the relationship that the priority is negatively correlated with the estimated delay, generate the delay priority based on the estimated delay and the smoothing factor; according to the relationship that the priority is positively correlated with the availability information, generate the availability priority based on the availability information and the resource availability weight; according to the relationship that the priority is positively correlated with the instruction type weight, generate the priority information of the to-be-scheduled instruction based on the instruction type weight, the path priority, the delay priority, and the availability priority. And during the process of generating the priority information of the to-be-scheduled instruction according to the relationship that the priority is positively correlated with the instruction type weight, based on the instruction type weight, the path priority, the delay priority, and the availability priority, the path priority and the delay priority can be positively correlatedly calculated according to the instruction type weight to generate the initial priority; and the initial priority and the availability priority can be positively correlatedly calculated to generate the priority information of the to-be-scheduled instruction.
[0094] To facilitate the understanding of the generation process of the priority information, assume that the instruction type weight is represented by the smoothing factor is represented by the path depth weight is represented by the resource availability weight is represented by the estimated delay is represented by the target path depth is represented by the availability information is represented by the priority information is represented by Then the generation formula for the priority information of the to-be-scheduled instruction c can be:
[0095] ;
[0096] where indicates that the resource is available and the instruction c can be immediately launched, indicates that the resource is unavailable and the instruction c needs to wait for the resource to be released.
[0097] From the above implementation process, it can be seen that in the process of generating the priority information of the to-be-scheduled instruction, the target path depth, availability information, and instruction type weight are set to positively affect the priority, and the estimated delay is set to negatively affect the priority. This is adapted to the scheduling rule that the greater the value of the target path depth, the better the availability, and the smaller the estimated delay, the higher the instruction scheduling priority, which can ensure the accuracy of priority generation and facilitate subsequent accurate scheduling of instructions based on the priority. Moreover, in this process, the path depth information can be smoothed to avoid the extreme situation where the value of the path depth is too large, resulting in a sharp increase in the priority and overly concentrated scheduling of certain instructions, thus ensuring the stability of instruction scheduling.
[0098] Step S105: Collect and process the processing information of the target instruction.
[0099] Step S106: Generate a target reward value for evaluating the instruction processing performance according to the processing information.
[0100] Step S107: Adjust the instruction weight according to the target reward value, and process the to-be-scheduled instruction based on the adjusted instruction weight.
[0101] In practical applications, after the to-be-scheduled instruction is determined, the instruction information will no longer change. If the instruction weight is set inappropriately for the instruction processing process, it will lead to unreasonable instruction scheduling, and then to unreasonable instruction processing. This unreasonableness can be reflected in affecting the instruction processing progress or not meeting the user's instruction processing requirements, etc. To avoid this situation, it is necessary to collect and process the processing information of the target instruction to reflect the situation of processing the instruction through the processing information. Then, the reinforcement learning method is applied to generate a target reward value for evaluating the instruction processing performance according to the processing information. Finally, the instruction weight is adjusted according to the target reward value, so that the instruction weight can be adjusted according to the instruction processing situation. If the to-be-scheduled instruction is scheduled based on the adjusted instruction weight subsequently, it realizes scheduling the to-be-scheduled instruction with reference to the previous round of instruction processing process. Repeating this process can make the instruction scheduling process adapt to the processing process and ensure the rationality of instruction scheduling.
[0102] In specific application scenarios, factors such as whether an instruction can be successfully launched, pauses during the instruction processing, and resource conflicts with other instructions will all affect the progress of instruction processing. Therefore, during the process of collecting the processing information of the target instruction, the number of successful launches of the target path of the target instruction can be collected. For example, during the instruction launch phase, check whether the launched instruction belongs to the target path. If it belongs to the target path and is successfully launched, increment the counter by one; collect the number of pause cycles of the target instruction. For example, check the instruction queue in each cycle. If the instruction queue is not empty but no instruction is successfully launched, increment the counter by one each time it occurs; collect the number of resource conflicts of the target instruction. For example, check in each cycle the situation where the data in the instruction queue is ready but cannot be launched because the required resources are occupied. Increment the resource counter by one each time it occurs; use the number of successful launches of the target path, the number of pause cycles, and the number of resource conflicts as the processing information for evaluating the performance of the target instruction.
[0103] In specific application scenarios, during the process of generating a target reward value for evaluating the instruction processing performance based on the processing information, weighted operations can be used to comprehensively consider all the processing information to obtain the target reward value. For example, the set reward weights corresponding to the processing information can be obtained. Specifically, obtain the set first adjustment weight for adjusting the impact of path launch on the reward, obtain the set second adjustment weight for adjusting the impact of pauses on the reward, obtain the set third adjustment weight for adjusting the impact of resource conflicts on the reward, and use the first adjustment weight, the second adjustment weight, and the third adjustment weight as the reward weights corresponding to the processing information; perform weighted operations on the processing information based on the reward weights to obtain the target reward value for evaluating the instruction processing performance.
[0104] In specific application scenarios, during the process of performing weighted operations on the processing information based on the reward weights to obtain the target reward value for evaluating the instruction processing performance, the first product value of the first adjustment weight and the number of successful launches of the target path can be determined; determine the second product value of the second adjustment weight and the number of pause cycles; generate the first difference between the first product value and the second product value; determine the third product value of the third adjustment weight and the number of resource conflicts; generate the second difference between the first difference and the third product value; use the second difference as the target reward value for evaluating the instruction processing performance.
[0105] For ease of understanding, assume that the first adjustment weight is represented by the second adjustment weight is represented by the third adjustment weight is represented by the number of pause cycles is represented by the number of resource conflicts is represented by the number of successful launches of the target path is represented by the target reward value is represented by It can be expressed that the generation formula of the target reward value of instruction c can be:
[0106] .
[0107] In a specific application scenario, during the process of adjusting the instruction weight according to the target reward value, when the gap between the estimated delay and the actual delay is too large, it indicates that the impact of the estimated delay on the priority is inaccurate. Therefore, it is necessary to adjust the smoothing factor corresponding to the estimated delay, that is, it is necessary to obtain the actual delay of the target instruction; detect whether the difference between the actual delay and the estimated delay is greater than the preset threshold; in response to the difference between the actual delay and the estimated delay being greater than the preset threshold, adjust the smoothing factor. During this process, the actual delay may also cause processing pauses. Therefore, when the target reward value indicates that there are many pipeline stall cycles, the current reliance on the predicted delay is too strong, and the prediction deviation causes scheduling misjudgment. The smoothing factor should be appropriately increased to suppress the role of the predicted delay in the priority calculation, avoid over-reliance on inaccurate delay predictions, and thus reduce stalls and improve the pipeline execution efficiency; in response to the difference between the actual delay and the estimated delay being less than or equal to the preset threshold, the smoothing factor may not be adjusted. Similarly, when the type of the target instruction to be processed first does not meet the user's requirements, the instruction type weight can be adjusted according to the user's requirements. For example, when the processing priority of a certain type of instruction is low, the instruction type weight of this type of instruction can be increased to improve the priority of this instruction for priority scheduling.
[0108] In a specific application scenario, it is also necessary to adjust the path depth weight or the resource availability weight according to the target reward value. During this process, dynamically adjusting the path depth weight can optimize the behavior of the scheduler in different operating scenarios. For example, in a delay-sensitive scenario, increasing the target path depth weight can give priority to scheduling instructions on the critical path. In a throughput-priority scenario, reducing the target path depth weight can improve the overall scheduling flexibility. Therefore, when the target reward value indicates that the instructions on the critical path have a greater impact on the overall processing performance, increase the critical path weight to improve the scheduling priority of the instructions on the critical path and better reduce the delay of the critical path. When the target reward value indicates that there are many resource conflicts or the resource utilization rate is low, increase the resource weight to strengthen the impact of the resource state on the instruction scheduling decision, reduce resource conflicts, and improve the resource usage efficiency. By dynamically adjusting the above instruction weights, the reinforcement learning algorithm can adapt to different operating scenarios in real time, optimize the dynamic scheduling of instructions, and significantly improve the overall performance of the device.
[0109] In a specific application scenario, before adjusting the path depth weight or resource availability weight according to the target reward value, if the target reward value is greater than or equal to the first setting value, it indicates that there is no need to adjust the path depth weight or resource availability weight, so it can be detected whether the target reward value is less than the first setting value; in response to the target reward value being less than the first setting value, the step of adjusting the path depth weight or resource availability weight according to the target reward value is performed. Correspondingly, in response to the target reward value being less than the first setting value and greater than or equal to the second setting value, the path depth weight is adjusted; in response to the target reward value being less than the second setting value, the resource availability weight is adjusted.
[0110] It can be seen from the above implementation process that the instruction scheduling algorithm based on reinforcement learning of the present invention can optimize the scheduling strategy in real time according to the device operation status, such as target path depth, resource utilization, predicted delay, etc., by dynamically adjusting the instruction priority weight, thereby improving the flexibility and efficiency of instruction scheduling. Compared with traditional static rule scheduling, it has stronger adaptability in complex dynamic scenarios, such as multi-task loads, cache misses, resource competition, etc., and can effectively reduce pipeline pauses, optimize the execution order of critical path instructions, reduce resource conflicts, and ultimately improve the overall performance and throughput of the device.
[0111] In actual applications, in the process of adjusting the instruction weight, for any instruction weight, the historical value of the instruction weight can be collected in chronological order, and the required weight value can be selected from the historical value as the candidate weight value. For example, if the weight value needs to be increased, the historical value greater than the current weight value is used as the candidate weight value. Conversely, if the weight value needs to be reduced, the historical value less than the current weight value is used as the candidate weight value; the candidate weight values are sorted in chronological order to obtain the sorting result, and the sliding window size value is determined according to the weight adjustment strength. For example, if the weight adjustment strength is large, the sliding window value is determined to be large, and if the weight adjustment strength is small, the sliding window value is determined to be small; traverse forward from the tail of the sorting result according to the sliding window size value, if the traversed weight value meets the set adjustment margin value, the traversed weight value is used as the adjusted weight value, if the traversed value does not meet the set adjustment margin value, return to execute the step of traversing forward from the tail of the sorting result according to the sliding window size value; if there is no suitable value after the traversal is completed, the adjusted weight value is determined based on the current weight value and the adjustment margin value. In this way, the present invention can use historical weight values as reference weight values to adjust current weight values. The adjustment of weight values can be completed by simply moving the sliding window. The data is stable and there is little noise interference. The set adjustment margin can avoid drastic changes in the adjustment of weight values, prevent overfitting of weight value adjustment, ensure the stability of weight adjustment, and then ensure the stability of instruction scheduling.
[0112] An instruction scheduling method provided by the present invention collects instruction information of instructions to be scheduled, determines the instruction weight corresponding to the current instruction information, generates priority information of the instructions to be scheduled according to the instruction information and the instruction weight, selects a target instruction from the instructions to be scheduled for processing according to the priority information, collects processing information of the target instruction, generates a target reward value for evaluating the instruction processing performance according to the processing information, and adjusts the instruction weight according to the target reward value to process the instructions to be scheduled based on the adjusted instruction weight. The beneficial effects of the present invention are as follows: Priority information of the instructions to be scheduled can be generated according to the instruction information and the current instruction weight, and a target instruction is selected for processing according to the priority information, so that the instruction processing process is controlled by the instruction weight; and then a target reward value is generated according to the processing information of the target instruction, which reflects the processing performance of the target instruction from the data level, so as to realize the quantification of the target instruction processing process; then the instruction weight is adjusted according to the target reward value, which is equivalent to adjusting the instruction weight according to the processing performance of the target instruction from the data level, so that the adjusted instruction weight is adapted to the processing process of the target instruction. When selecting a target instruction for processing later, it is equivalent to selecting an instruction using the adjusted instruction weight, so as to realize the dynamic adaptive scheduling of the instruction, without relying on a complex hardware structure and having good flexibility.
[0113] To facilitate the understanding of the instruction scheduling scheme provided by the present invention, the instruction scheduling process in the processor is described below. Assume that the instructions and related parameters in the processor instruction queue are shown in Table 3. Among them, the reinforcement learning model applied by the processor can be as Figure 2 shown. The state extraction module collects the running state of the processor in real time, including instruction type, predicted instruction latency, critical path depth, and resource status. The predicted instruction latency refers to the latency time when predicting a miss in the L1 cache. The scheduling execution module selects the instruction with the highest priority for emission. The reward feedback module calculates the reward value according to the number of successfully emitted critical paths, the number of pipeline stall cycles, and the number of resource conflicts. The reinforcement learning algorithm module dynamically adjusts the weights in the priority formula using the reward value to optimize the future scheduling strategy. And a decoder, a latency cache, a critical path depth table, a performance monitoring unit, and a reinforcement learning unit can be added to the processor hardware to assist in instruction scheduling. The decoder module analyzes the instruction to obtain the instruction type. The latency cache is used to record the historical latency data when the cache is missed. The critical path depth table is used to record the critical path depth and dependency relationship of the instruction to obtain the critical path depth. The performance monitoring unit is used to count the usage of hardware resources to obtain the resource availability status. The reinforcement learning unit is used to dynamically adjust the weight parameters according to the reward, as Figure 3 shown.
[0114] Table 3 Example instructions and related parameters in the instruction queue
[0115]
[0116] Assume that the initial critical path depth weight is 2.0, the resource availability weight is 0.5, and the smoothing factor is 1.0. Then the scheduling process is as follows:
[0117] In the first round of scheduling, instructions I1, I4, and I6 can be scheduled immediately. The priorities of instructions I1, I4, and I6 are:
[0118] ;
[0119] ;
[0120] ;
[0121] Instruction I6 has the highest priority. Therefore, issue instruction I6. Since instruction I6 is a non-critical path and does not increase the number of successfully issued instructions on the critical path, the reward calculation for this round , and assume no conflicts , , the reward for this round . The reward is very low. Increase the critical path depth weight from 2.0 to 2.3 to improve the priority of critical path instructions and preferentially schedule critical path instructions in the next round;
[0122] In the second round of scheduling, I6 has been issued, I1 and I4 are waiting to be scheduled, and I2 and I3 cannot be scheduled because the critical path depth weight has been increased. The priorities of instructions I1 and I4 are:
[0123] ;
[0124] ;
[0125] The priority of instruction I4 remains unchanged. Therefore, issue instruction I3. The reward calculation for this round , , , the reward for this round . The performance is good and the weight remains unchanged. However, in actual operation, the actual latency of I3 is 30, which is greater than the predicted value of 15.
[0126] In the fourth round of scheduling, I3, I6, and I1 have been issued, and I2, I4, and I5 are schedulable. However, because the actual latency of I3 is large and not completed, I5 is not ready. The instruction queue for this round is not empty, but there are no instructions to issue, and the pipeline stalls. Therefore, the reward calculation for this round , , , the reward for this round 。The pause is caused by inaccurate delay prediction. Adjust the smoothing factor, increasing it from 1.0 to 1.3.
[0127] During the fifth round of scheduling, I3, I6, and I1 have been issued, and I2, I4, and I5 are schedulable. Instruction I5 calculates the priority:
[0128] ;
[0129] Calculation of the reward for this round , , , the reward for this round . The performance is good, and the weight remains unchanged. However, the critical path ends, and the counter is cleared. Therefore, , , .
[0130] During the sixth round of scheduling, I5, 3, I6, and I1 have been issued, and I2 and I4 are schedulable. Assume that instruction I2 is issued and there is a resource conflict. Calculation of the reward for this round , , , the reward for this round . The reward is negative, and the resource availability weight is increased from 0.5 to 0.7.
[0131] For the seventh round of scheduling, the remaining instruction is I4, and the priority is recalculated:
[0132] ;
[0133] Issue instruction I4. Calculate the reward for this round , , , the reward for this round , since there are no critical instructions in the follow-up and no negative significance, the weight can be not adjusted.
[0134] The present invention also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of an instruction scheduling method provided by an embodiment of the present invention. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0135] An electronic device provided by an embodiment of the present invention includes a memory 201 and a processor 202. A computer program is stored in the memory 201, and when the processor 202 executes the computer program, the steps of the instruction scheduling method described in any of the above embodiments are implemented.
[0136] Please refer to Figure 5, another electronic device provided by an embodiment of the present invention may further include: an input port 203 connected to the processor 202, configured to transmit externally input commands to the processor 202; a display unit 204 connected to the processor 202, configured to display the processing result of the processor 202 to the outside; a communication module 205 connected to the processor 202, configured to implement communication between the electronic device and the outside. The display unit 204 may be a display panel, a laser scanning display, etc.; the communication methods adopted by the communication module 205 include but are not limited to Mobile High-Definition Link (MHL), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connection: Wireless Fidelity (WiFi), Bluetooth communication technology, low-power Bluetooth communication technology, and communication technology based on IEEE802.11s.
[0137] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program, and when the computer program is executed by a processor, it implements the steps of the instruction scheduling method described in any of the above embodiments.
[0138] The computer-readable storage medium involved in the present invention includes Random Access Memory (RAM), memory, Read-Only Memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium well-known in the technical field.
[0139] A computer program product provided by an embodiment of the present invention includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the following steps are implemented:
[0140] Collect the instruction information of the instruction to be scheduled;
[0141] Determine the instruction weight corresponding to the current instruction information;
[0142] Generate the priority information of the instruction to be scheduled according to the instruction information and the instruction weight;
[0143] Select a target instruction from the instructions to be scheduled for processing according to the priority information;
[0144] Collect the processing information of the processed target instruction;
[0145] Generate a target reward value for evaluating the instruction processing performance according to the processing information;
[0146] Adjust the instruction weight according to the target reward value, and process the instruction to be scheduled based on the adjusted instruction weight.
[0147] A computer program product provided by an embodiment of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, the following steps are implemented: predict the instruction delay of the instruction to be scheduled to obtain an estimated delay; collect the type information and target path depth of the instruction to be scheduled; collect the availability information of the resources corresponding to the instruction to be scheduled; use the type information, estimated delay, target path depth, and availability information as the instruction information of the instruction to be scheduled; determine the instruction weight corresponding to the instruction information currently, and the instruction weight includes an instruction type weight, a smoothing factor, a path depth weight, and a resource availability weight.
[0148] A computer program product provided by an embodiment of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, the following steps are implemented: generate path depth information according to the target path depth and the path depth weight according to the relationship that the priority is positively correlated with the target path depth; smooth the path depth information to obtain a path priority; generate a delay priority according to the estimated delay and the smoothing factor according to the relationship that the priority is negatively correlated with the estimated delay; generate an availability priority according to the availability information and the resource availability weight according to the relationship that the priority is positively correlated with the availability information; generate the priority information of the instruction to be scheduled according to the instruction type weight, path priority, delay priority, and availability priority according to the relationship that the priority is positively correlated with the instruction type weight.
[0149] A computer program product provided by an embodiment of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, the following steps are implemented: perform a positive correlation operation on the path priority and the delay priority according to the instruction type weight to generate an initial priority; perform a positive correlation operation on the initial priority and the availability priority to generate the priority information of the instruction to be scheduled.
[0150] A computer program product provided by an embodiment of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, the following steps are implemented: collect the number of successful target path launches for processing the target instruction; collect the number of stall cycles for processing the target instruction; collect the number of resource conflicts for processing the target instruction; use the number of successful target path launches, the number of stall cycles, and the number of resource conflicts as the processing information for processing the target instruction.
[0151] A computer program product provided by an embodiment of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the following steps are implemented: obtaining a set reward weight corresponding to processing information; performing a weighted operation on the processing information based on the reward weight to obtain a target reward value for evaluating the instruction processing performance.
[0152] A computer program product provided by an embodiment of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the following steps are implemented: obtaining a set first adjustment weight for adjusting the influence of path emission on the reward; obtaining a set second adjustment weight for adjusting the influence of stalls on the reward; obtaining a set third adjustment weight for adjusting the influence of resource conflicts on the reward; using the first adjustment weight, the second adjustment weight, and the third adjustment weight as the reward weight corresponding to the processing information.
[0153] A computer program product provided by an embodiment of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the following steps are implemented: determining a first product value of the first adjustment weight and the number of successful target path emissions; determining a second product value of the second adjustment weight and the number of stall cycles; generating a first difference between the first product value and the second product value; determining a third product value of the third adjustment weight and the number of resource conflicts; generating a second difference between the first difference and the third product value; using the second difference as the target reward value for evaluating the instruction processing performance.
[0154] A computer program product provided by an embodiment of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the following steps are implemented: obtaining the actual latency of a target instruction; detecting whether the difference between the actual latency and the estimated latency is greater than a preset threshold; in response to the difference between the actual latency and the estimated latency being greater than the preset threshold, adjusting the smoothing factor and adjusting the path depth weight or the resource availability weight according to the target reward value; in response to the difference between the actual latency and the estimated latency being less than or equal to the preset threshold, adjusting the path depth weight or the resource availability weight according to the target reward value.
[0155] A computer program product provided by an embodiment of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the following steps are implemented: before adjusting the instruction type weight or the resource availability weight according to the target reward value, detecting whether the target reward value is less than a first set value; in response to the target reward value being less than the first set value, performing the step of adjusting the path depth weight or the resource availability weight according to the target reward value.
[0156] A computer program product provided by an embodiment of the present invention includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the following steps are implemented: in response to the target reward value being less than the first set value and greater than or equal to the second set value, adjust the path depth weight.
[0157] A computer program product provided by an embodiment of the present invention includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the following steps are implemented: in response to the target reward value being less than the second set value, adjust the resource availability weight.
[0158] For the description of the relevant parts in an electronic device, a computer program product, and a computer-readable storage medium provided by an embodiment of the present invention, please refer to the detailed description of the corresponding parts in an instruction scheduling method provided by an embodiment of the present invention, which will not be elaborated here. In addition, in the above technical solutions provided by the embodiments of the present invention, the parts that are the same as the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0159] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0160] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An instruction scheduling method, characterized in that, It includes: Collect the instruction information of the to-be-scheduled instruction; Determine the current instruction weight corresponding to the instruction information; Generate the priority information of the to-be-scheduled instruction according to the instruction information and the instruction weight; Select a target instruction from the to-be-scheduled instructions for processing according to the priority information; Collect the processing information of processing the target instruction; Generate a target reward value for evaluating the instruction processing performance according to the processing information; Adjust the instruction weight according to the target reward value, and process the to-be-scheduled instruction based on the adjusted instruction weight; Among them, collecting the instruction information of the to-be-scheduled instruction includes: Predict the instruction delay of the to-be-scheduled instruction to obtain the estimated delay; Collect the type information and the target path depth of the to-be-scheduled instruction; Collect the availability information of the resources corresponding to the to-be-scheduled instruction; Use the type information, the estimated delay, the target path depth, and the availability information as the instruction information of the to-be-scheduled instruction; Among them, collecting the processing information of processing the target instruction includes: Collect the number of successful target path transmissions for processing the target instruction; Collect the number of pause cycles for processing the target instruction; Collect the number of resource conflicts for processing the target instruction; Use the number of successful target path transmissions, the number of pause cycles, and the number of resource conflicts as the processing information of processing the target instruction.
2. The instruction scheduling method according to claim 1, wherein Determining the current instruction weight corresponding to the instruction information includes; Determine the current instruction weight corresponding to the instruction information, and the instruction weight includes an instruction type weight, a smoothing factor, a path depth weight, and a resource availability weight.
3. The instruction scheduling method according to claim 2, wherein Generating the priority information of the to-be-scheduled instruction according to the instruction information and the instruction weight includes: According to the relationship that the priority is positively correlated with the target path depth, generate path depth information according to the target path depth and the path depth weight; Smooth the path depth information to obtain the path priority; According to the relationship that the priority is negatively correlated with the estimated delay, generate a delay priority according to the estimated delay and the smoothing factor; According to the relationship that the priority is positively correlated with the availability information, generate an availability priority according to the availability information and the resource availability weight; According to the relationship that the priority is positively correlated with the instruction type weight, generate the priority information of the to-be-scheduled instruction according to the instruction type weight, the path priority, the delay priority, and the availability priority.
4. The instruction scheduling method according to claim 3, wherein According to the relationship that the priority is positively correlated with the instruction type weight, generate the priority information of the to-be-scheduled instruction according to the instruction type weight, the path priority, the delay priority, and the availability priority, including: Perform a positive correlation operation on the path priority and the delay priority according to the instruction type weight to generate an initial priority; Perform a positive correlation operation on the initial priority and the availability priority to generate the priority information of the to-be-scheduled instruction.
5. The instruction scheduling method according to any one of claims 2 to 4, characterized in that Generating a target reward value for evaluating the instruction processing performance according to the processing information includes: Obtain the set reward weight corresponding to the processing information; Perform a weighted operation on the processing information based on the reward weight to obtain a target reward value for evaluating the instruction processing performance.
6. The instruction scheduling method according to claim 5, wherein Obtain the set reward weight corresponding to the processing information, including: Obtain the set first adjustment weight for adjusting the influence of path emission on the reward; Obtain the set second adjustment weight for adjusting the influence of stalls on the reward; Obtain the set third adjustment weight for adjusting the influence of resource conflicts on the reward; Use the first adjustment weight, the second adjustment weight, and the third adjustment weight as the reward weight corresponding to the processing information.
7. The instruction scheduling method according to claim 6, wherein Perform a weighted operation on the processing information based on the reward weight to obtain a target reward value for evaluating the instruction processing performance, including: Determine a first product value of the first adjustment weight and the number of successfully emitted target paths; Determine a second product value of the second adjustment weight and the number of stall cycles; Generate a first difference between the first product value and the second product value; Determine a third product value of the third adjustment weight and the number of resource conflicts; Generate a second difference between the first difference and the third product value; Use the second difference as the target reward value for evaluating the instruction processing performance.
8. The instruction scheduling method according to any one of claims 2 to 4, characterized in that Adjust the instruction weight according to the target reward value, including: Obtain the actual latency of the target instruction; Detect whether the difference between the actual latency and the estimated latency is greater than a preset threshold; In response to the difference between the actual latency and the estimated latency being greater than the preset threshold, adjust the smoothing factor and adjust the path depth weight or the resource availability weight according to the target reward value; In response to the difference between the actual latency and the estimated latency being less than or equal to the preset threshold, adjust the path depth weight or the resource availability weight according to the target reward value.
9. The instruction scheduling method according to claim 8, wherein Before adjusting the path depth weight or the resource availability weight according to the target reward value, further include: Detect whether the target reward value is less than a first set value; In response to the target reward value being less than the first set value, perform the step of adjusting the path depth weight or the resource availability weight according to the target reward value.
10. The instruction scheduling method according to claim 9, characterized in that, Adjust the path depth weight or the resource availability weight according to the target reward value, including: In response to the target reward value being less than the first set value and greater than or equal to a second set value, adjust the path depth weight.
11. The instruction scheduling method according to claim 10, wherein Adjust the path depth weight or the resource availability weight according to the target reward value, including: In response to the target reward value being less than the second set value, adjust the resource availability weight.
12. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the instruction scheduling method according to any one of claims 1 to 11 are implemented.
13. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the instruction scheduling method according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the instruction scheduling method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Multi-user downlink scheduling method and device, equipment and storage medium
CN117545085A
Dynamic task scheduling method and device, server and storage medium
CN119336466A