Instruction scheduling method, computer program product, electronic equipment and medium
Through methods of collecting instruction information, determining weights, generating priority and adjusting instruction weights, the flexibility of instruction scheduling in the prior art in multi-core and high-load scenarios is solved, and efficient and flexible instruction scheduling is achieved.
Patent Information
- Application Number
- CN202510502889.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art is difficult to achieve flexible instruction adaptive scheduling in multi-core and high-load scenarios, resulting in high power consumption of instruction scheduling, high hardware costs, and difficult to scale.
By collecting instruction information of the command to be scheduled, determining the instruction weight, generating priority information, selecting target instructions for processing, and adjusting the instruction weight according to the processing information to achieve dynamic adaptive scheduling.
It realizes dynamic control and quantification of instruction processing processes, improves the flexibility and efficiency of instruction scheduling, and reduces dependence on complex hardware structures.
Smart Images

Figure CN120010928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to an instruction scheduling method, a computer program product, an electronic device and a medium. Background Art
[0002] In computers, instruction scheduling is one of the factors that affect performance, so it is necessary to optimize the order of instruction issuance, make full use of pipeline and execution unit resources, reduce pipeline stagnation time, and improve instruction throughput. For example, instructions can be scheduled through out-of-order execution, that is, by dynamically tracking instruction dependencies, instructions with ready operands can be scheduled first.
[0003] However, out-of-order execution relies on complex hardware structures, such as reservation stations and reorder buffers, which results in high power consumption and hardware cost for instruction scheduling, and makes it difficult to further expand in multi-core and high-load scenarios. In addition, with the popularity of multi-tasking and multi-threading, instruction scheduling needs to face more complex dynamic loads, but out-of-order scheduling has been difficult to maintain good adaptability.
[0004] In summary, how to flexibly and adaptively schedule instructions is a problem that currently needs to be solved urgently by those skilled in the art. Summary of the invention
[0005] The purpose of the present invention is to provide an instruction scheduling method, which can solve the technical problem of how to flexibly and adaptively schedule instructions to a certain extent. The present invention also provides a computer program product, an electronic device and a computer-readable storage medium.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] In a first aspect, an instruction scheduling method is provided, comprising:
[0008] Collecting instruction information of instructions to be dispatched;
[0009] Determining a current instruction weight corresponding to the instruction information;
[0010] Generate priority information of the instruction to be scheduled according to the instruction information and the instruction weight;
[0011] According to the priority information, a target instruction is selected from the instructions to be scheduled for processing;
[0012] Collecting and processing the processing information of the target instruction;
[0013] generating a target reward value for evaluating instruction processing performance based on the processing information;
[0014] The instruction weight is adjusted according to the target reward value, so as to process the instruction to be scheduled based on the adjusted instruction weight.
[0015] On the other hand, the instruction information of the instruction to be scheduled is collected, including:
[0016] Predict the instruction delay of the instruction to be scheduled and obtain the estimated delay;
[0017] Collect the type information and target path depth of the instructions to be scheduled;
[0018] Collecting availability information of resources corresponding to the instructions to be scheduled;
[0019] Using the type information, the estimated delay, the target path depth and the availability information as instruction information of the instruction to be scheduled;
[0020] Determining the instruction weight currently corresponding to the instruction information includes:
[0021] Determine the instruction weight currently corresponding to the instruction information, wherein the instruction weight includes an instruction type weight, a smoothing factor, a path depth weight, and a resource availability weight.
[0022] On the other hand, priority information of the instruction to be scheduled is generated according to the instruction information and the instruction weight, including:
[0023] According to the positive correlation between the priority and the target path depth, generating path depth information according to the target path depth and the path depth weight;
[0024] Smoothing the path depth information to obtain path priority;
[0025] According to the negative correlation between the priority and the estimated delay, generating a delay priority according to the estimated delay and the smoothing factor;
[0026] According to the positive correlation between the priority and the availability information, generating the availability priority according to the availability information and the resource availability weight;
[0027] According to the positive correlation between the priority and the instruction type weight, the priority information of the instruction to be scheduled is generated according to the instruction type weight, the path priority, the delay priority and the availability priority.
[0028] On the other hand, according to the positive correlation between the priority and the instruction type weight, the priority information of the instruction to be scheduled is generated according to the instruction type weight, the path priority, the delay priority and the availability priority, including:
[0029] According to the instruction type weight, a positive correlation operation is performed on the path priority and the delay priority to generate an initial priority;
[0030] A positive correlation operation is performed on the initial priority and the availability priority to generate priority information of the instruction to be scheduled.
[0031] On the other hand, collecting and processing the processing information of the target instruction includes:
[0032] Collect and process the number of successful transmissions of the target path of the target instruction;
[0033] Collecting the number of pause cycles of processing the target instruction;
[0034] Collect and process the number of resource conflicts of the target instruction;
[0035] The target path transmission success number, the stall cycle number and the resource conflict number are used as processing information for processing the target instruction.
[0036] On the other hand, generating a target reward value for evaluating instruction processing performance according to the processing information includes:
[0037] Obtaining a reward weight set corresponding to the processing information;
[0038] The processing information is weighted based on the reward weight to obtain a target reward value for evaluating instruction processing performance.
[0039] On the other hand, obtaining a reward weight set corresponding to the processing information includes:
[0040] Obtaining a first adjustment weight set for adjusting the effect of path emission on the reward;
[0041] Obtaining a second adjustment weight set for adjusting the effect of the pause on the reward;
[0042] Obtaining a third adjustment weight set for adjusting the impact of resource conflict on rewards;
[0043] The first adjustment weight, the second adjustment weight and the third adjustment weight are used as reward weights corresponding to the processing information.
[0044] On the other hand, performing a weighted operation on the processing information based on the reward weight to obtain a target reward value for evaluating instruction processing performance includes:
[0045] Determine a first product value of the first adjustment weight and the number of successful transmissions of the target path;
[0046] Determine a second product value of the second adjustment weight and the number of pause cycles;
[0047] generating a first difference between the first product value and the second product value;
[0048] Determining a third product value of the third adjustment weight and the number of resource conflicts;
[0049] generating a second difference value between the first difference value and the third product value;
[0050] The second difference is used as a target reward value for evaluating instruction processing performance.
[0051] On the other hand, adjusting the instruction weight according to the target reward value includes:
[0052] Obtaining an actual delay of the target instruction;
[0053] Detecting whether a difference between the actual delay and the estimated delay is greater than a preset threshold;
[0054] In response to a difference between the actual delay and the estimated delay being greater than the preset threshold, adjusting the smoothing factor, and adjusting the path depth weight or the resource availability weight according to the target reward value;
[0055] In response to a difference between the actual delay and the estimated delay being less than or equal to the preset threshold, the path depth weight or the resource availability weight is adjusted according to the target reward value.
[0056] On the other hand, before adjusting the path depth weight or the resource availability weight according to the target reward value, the method further includes:
[0057] Detecting whether the target reward value is less than a first set value;
[0058] In response to the target reward value being less than the first set value, a step of adjusting the path depth weight or the resource availability weight according to the target reward value is performed.
[0059] On the other hand, adjusting the path depth weight or the resource availability weight according to the target reward value includes:
[0060] In response to the target reward value being less than the first setting value and greater than or equal to a second setting value, the path depth weight is adjusted.
[0061] On the other hand, adjusting the path depth weight or the resource availability weight according to the target reward value includes:
[0062] In response to the target reward value being less than the second set value, the resource availability weight is adjusted.
[0063] In a second aspect, a computer program product is provided, comprising a computer program / instruction, which implements the steps of any of the above-mentioned instruction scheduling methods when executed by a processor.
[0064] According to a third aspect, an electronic device is provided, including:
[0065] Memory for storing computer programs;
[0066] A processor is used to implement the steps of any of the above instruction scheduling methods when executing the computer program.
[0067] According to a fourth aspect, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above instruction scheduling methods are implemented.
[0068] The present invention provides an instruction scheduling method, which collects instruction information of instructions to be scheduled; determines the instruction weight currently corresponding to the instruction information; generates priority information of the instructions to be scheduled according to the instruction information and the instruction weight; selects a target instruction from the instructions to be scheduled for processing according to the priority information; collects processing information of the target instruction; generates a target reward value for evaluating instruction processing performance according to the processing information; and adjusts the instruction weight according to the target reward value, so as to process the instruction to be scheduled based on the adjusted instruction weight.
[0069] The beneficial effects of the present invention are: priority information of the instruction to be scheduled can be generated according to the instruction information and the current instruction weight, and the target instruction can be selected for processing according to the priority information, so that the instruction processing process is controlled by the instruction weight; and then the target reward value is generated according to the processing information of the target instruction, reflecting the processing performance of the target instruction from the data level, thereby realizing the quantification of the target instruction processing process; then the instruction weight is adjusted according to the target reward value, which is equivalent to adjusting the instruction weight according to the processing performance of the target instruction from the data level, so that the adjusted instruction weight is adapted to the processing process of the target instruction, and if the target instruction is selected for processing later, it is equivalent to applying the adjusted instruction weight for instruction selection, thereby realizing dynamic adaptive scheduling of instructions, without relying on complex hardware structure, and good flexibility. A computer program product, electronic device and computer-readable storage medium provided by the present invention also solve corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0071] Figure 1 A flowchart of an instruction scheduling method provided by an embodiment of the present invention;
[0072] Figure 2 Schematic diagram of the reinforcement learning model for the processor;
[0073] Figure 3 Enter schematic diagram for processor parameters;
[0074] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0075] Figure 5 Another structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0076] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0077] See also Figure 1 , Figure 1 A flowchart of an instruction scheduling method provided by an embodiment of the present invention.
[0078] An instruction scheduling method provided by an embodiment of the present invention may include the following steps:
[0079] Step S101: Collect instruction information of instructions to be scheduled.
[0080] In actual applications, if instructions are to be scheduled and processed, it is necessary to know the corresponding information of the instructions so that the processing method of the instructions can be determined based on the corresponding information of the instructions. Therefore, the instruction information of the instructions to be scheduled needs to be collected. The number and type of instructions to be scheduled can be flexibly determined according to the application scenario. For example, the instructions to be scheduled can be instructions that need to be processed in the server, or they can be instructions that need to be processed in the processor, etc.
[0081] In a specific application scenario, considering that instruction type, instruction delay, target path depth, and resource status will affect the scheduling of all instructions, in the process of collecting instruction information of the instruction to be scheduled, the instruction delay of the instruction to be scheduled can be predicted to obtain an estimated delay, which can be the delay time when the L1 cache is predicted to be missed. The type information and target path depth of the instruction to be scheduled are collected. The type information can be obtained by analyzing the instruction to be scheduled through the decoding module. For example, the type information can be a load instruction, a storage instruction, an ALU (Arithmetic And Logic unit) instruction, etc. The target path depth can be the critical path depth. This critical path depth refers to the number of instructions on the dependency chain passed from the start instruction that does not depend on other instructions to an instruction on the critical path in the data dependency relationship of the program. As shown in Table 1, the dependency relationship of all instructions in the instruction queue is recorded, including the critical path depth, instruction address, generator instruction address, and consumer instruction address. Collect the availability information of the resources corresponding to the instructions to be scheduled. This availability information refers to the current availability status of the resources required by the instructions, such as the execution unit, and prioritize the instructions that can use the idle resources to avoid resource waste. The type information, estimated delay, target path depth and availability information are used as the instruction information of the instructions to be scheduled.
[0082] Table 1 Critical path depth data structure
[0083]
[0084] It should be noted that the reason why the critical path depth is used as the target path depth is that the critical path is unique during the program execution phase, and the critical path depth reflects the importance of the instruction in the entire program. Instructions with a large critical path depth are located at the end of the dependency chain and have a greater impact on the overall program execution. Therefore, ensuring that the critical path instructions are scheduled first can reduce delays on the critical path and improve processor performance, so the critical path depth can be used as the target path depth.
[0085] Step S102: Determine the instruction weight currently corresponding to the instruction information.
[0086] Step S103: Generate priority information of the instruction to be scheduled according to the instruction information and the instruction weight.
[0087] Step S104: According to the priority information, a target instruction is selected from the instructions to be scheduled for processing.
[0088] In practical applications, when there are multiple instructions to be scheduled, how to select the appropriate instruction for scheduling is a problem that needs to be solved. In the instruction scheduling process, it is considered that the instruction information will affect the instruction processing process. For example, if there is a dependency relationship between two instructions, the latter instruction needs to wait for the previous instruction to be processed before it can be instructed. In other words, if the latter instruction is prioritized for processing at this time, it will only block the instruction processing process. Therefore, the instruction scheduling process can be controlled according to the type of instruction information. Further, the instruction weight corresponding to the current instruction information can be determined, that is, the corresponding instruction weight can be set for each type of instruction information. This instruction weight can be used to characterize the impact of instruction information on instruction scheduling. In this way, only the instruction information and instruction weight need to be calculated to generate the priority information of the instruction to be scheduled after comprehensively considering all instruction information; subsequently, according to the priority information, if the target instruction is selected for processing in the instructions to be scheduled, it is equivalent to comprehensively considering the impact of all instruction information on instruction scheduling to perform instruction scheduling, so as to ensure the rationality of instruction scheduling.
[0089] In specific application scenarios, based on priority information, when selecting target instructions for processing from the instructions to be scheduled, the instruction to be scheduled with the highest priority and not dependent on other instructions can be selected as the target instruction for processing first. Of course, the priority can also be flexibly applied to select the target instruction for processing according to the actual scenario.
[0090] In a specific application scenario, in the process of determining the instruction weight currently corresponding to the instruction information, the instruction weight currently corresponding to the instruction information can be determined, and the instruction weight includes instruction type weight, smoothing factor, path depth weight, and resource availability weight; wherein, the instruction type weight is used to determine the priority of different types of instructions, as shown in Table 2, load instructions, prefetch instructions, floating-point operation instructions, and atomic operation instructions have higher weights, because such instructions have strong data dependence and have a greater impact on pipeline performance, storage instructions, branch instructions, multimedia instructions, and memory management instructions have medium weights, because such instructions are closely related to the scenario, arithmetic operation instructions, logical operation instructions, and displacement instructions have lower weights, because such instruction delay angles are less dependent on other instructions; the smoothing factor is used to smooth the value of the inverse of the emission time to prevent the priority corresponding to the instruction delay from being infinitely large when the predicted delay time is very small, thereby causing the priority to be unreasonably extremely biased towards certain instructions; the path depth weight is used to determine the priority of different target path depths; the resource availability weight is used to indicate the importance of resource availability in the scheduling priority calculation.
[0091] Table 2 Instruction type weight list
[0092]
[0093] In a specific application scenario, in the process of generating priority information of instructions to be scheduled according to instruction information and instruction weight, path depth information can be generated according to the target path depth and path depth weight in accordance with the positive correlation between priority and target path depth; path depth information can be smoothed, such as by smoothing the path depth information through a logarithmic function, to obtain path priority; delay priority can be generated according to the estimated delay and smoothing factor in accordance with the negative correlation between priority and estimated delay; availability priority can be generated according to the availability information and resource availability weight in accordance with the positive correlation between priority and availability information; priority information of instructions to be scheduled can be generated according to instruction type weight, path priority, delay priority and availability priority in accordance with the positive correlation between priority and instruction type weight. In the process of generating priority information of instructions to be scheduled according to instruction type weight, path priority, delay priority and availability priority in accordance with the positive correlation between priority and instruction type weight, path priority and delay priority can be positively correlated according to instruction type weight to generate initial priority; initial priority and availability priority can be positively correlated to generate priority information of instructions to be scheduled.
[0094] In order to understand the generation process of priority information, it is assumed that the instruction type weight is expressed as The smoothing factor is expressed as Indicates that the path depth weight is expressed as Indicates that resource availability weights are expressed as Indicates that the estimated delay is The target path depth is represented by Indicates that availability information is expressed as Indicates that priority information is expressed as Indicates that the generation formula of the priority information of the instruction to be scheduled c can be:
[0095] ;
[0096] in, Indicates that resources are available and instruction c can be issued immediately. Indicates that resources are unavailable and instruction c needs to wait for resources to be released.
[0097] From the implementation process, it can be seen that in the process of generating priority information of instructions to be scheduled, the present invention sets the target path depth, availability information and instruction type weight to positively affect the priority, and sets the estimated delay to negatively affect the priority, which is adapted to the scheduling rule that the larger the value of the target path depth, the better the availability, and the smaller the estimated delay, the higher the instruction scheduling priority. This can ensure the accuracy of priority generation and facilitate the subsequent accurate scheduling of instructions by priority; and in this process, the path depth information can be smoothed to avoid the extreme situation where the path depth value is too large, resulting in a sharp increase in priority, and excessively concentrated scheduling of certain instructions, thereby ensuring the stability of instruction scheduling.
[0098] Step S105: Collect processing information of the processing target instruction.
[0099] Step S106: Generate a target reward value for evaluating instruction processing performance based on the processing information.
[0100] Step S107: adjusting the instruction weight according to the target reward value, so as to process the instruction to be scheduled based on the adjusted instruction weight.
[0101] In actual applications, after the scheduling instruction is determined, the instruction information will no longer change. If the instruction weight setting is not compatible with the instruction processing process, it will lead to unreasonable instruction scheduling, and then lead to unreasonable instruction processing process. This unreasonableness can be reflected in affecting the progress of instruction processing or not meeting the user's instruction processing needs. In order to avoid this situation, it is necessary to collect processing information of the target instruction so that the processing information can be used to reflect the processing instruction situation. Then, the reinforcement learning method is applied to generate a target reward value for evaluating instruction processing performance based on the processing information. Finally, according to the target reward value, the instruction weight is adjusted, so that the instruction weight can be adjusted according to the instruction processing situation. If the scheduling instruction is subsequently scheduled based on the adjusted instruction weight, it is realized that the scheduling instruction is scheduled with reference to the previous round of instruction processing process. If this is repeated, the instruction scheduling process can be adapted to the processing process to ensure the rationality of instruction scheduling.
[0102] In specific application scenarios, whether instructions can be successfully issued, pauses during instruction processing, and resource conflicts with other instructions will all affect the progress of instruction processing. Therefore, in the process of collecting processing information of the target instruction, the number of successful issuance of the target path of the target instruction can be collected. For example, in the instruction issuance stage, check whether the issued instruction belongs to the target path. If it belongs to the target path and is successfully issued, the counter is incremented by one; the number of pause cycles of the target instruction is collected, such as checking the instruction queue in each cycle. If the instruction queue is not empty but no instruction is successfully issued, the counter is incremented by one each time it occurs; the number of resource conflicts of the target instruction is collected, such as checking the instruction queue in each cycle to see if the data is ready but cannot be issued because the required resources are occupied. Each time it occurs, the resource counter is incremented by one; the number of successful issuance of the target path, the number of pause cycles, and the number of resource conflicts are used as processing information to evaluate the performance of the target instruction.
[0103] In a specific application scenario, in the process of generating a target reward value for evaluating instruction processing performance based on processing information, all processing information can be comprehensively considered with the help of weight calculation to obtain the target reward value. For example, the set reward weight corresponding to the processing information can be obtained. Specifically, the first adjustment weight set for adjusting the impact of path emission on the reward is obtained, the second adjustment weight set for adjusting the impact of pause on the reward is obtained, and the third adjustment weight set for adjusting the impact of resource conflict on the reward is obtained. The first adjustment weight, the second adjustment weight and the third adjustment weight are used as the reward weights corresponding to the processing information; the processing information is weighted based on the reward weight to obtain the target reward value for evaluating instruction processing performance.
[0104] In a specific application scenario, in the process of performing weighted operations on processing information based on reward weights to obtain a target reward value for evaluating instruction processing performance, a first product value of a first adjustment weight and the number of successful transmissions of a target path can be determined; a second product value of a second adjustment weight and the number of pause cycles can be determined; a first difference between the first product value and the second product value can be generated; a third product value of a third adjustment weight and the number of resource conflicts can be determined; a second difference between the first difference and the third product value can be generated; and the second difference can be used as the target reward value for evaluating instruction processing performance.
[0105] For ease of understanding, assume that the first adjustment weight is The number of successful launches on the target path is expressed as The second adjustment weight is expressed as Indicates that the number of pause cycles is expressed as Indicates that the third adjustment weight is Indicates the number of resource conflicts. The target reward value is expressed as Represented, the formula for generating the target reward value of instruction c can be:
[0106] .
[0107] In a specific application scenario, in the process of adjusting the instruction weight according to the target reward value, when the gap between the estimated delay and the actual delay is too large, it indicates that the impact of the estimated delay on the priority is inaccurate, so the smoothing factor corresponding to the estimated delay needs to be adjusted, that is, the actual delay of the target instruction needs to be obtained; detect whether the difference between the actual delay and the estimated delay is greater than the preset threshold; in response to the difference between the actual delay and the estimated delay being greater than the preset threshold, the smoothing factor is adjusted. In this process, the actual delay may also cause processing pauses. Therefore, when the target reward value indicates that the pipeline has more pause cycles, the current reliance on the predicted delay is too strong, and the prediction deviation causes scheduling misjudgment. The smoothing factor should be appropriately increased to suppress the role of the predicted delay in the priority calculation, avoid excessive reliance on inaccurate delay predictions, thereby reducing pauses and improving pipeline execution efficiency; in response to the difference between the actual delay and the estimated delay being less than or equal to the preset threshold, the smoothing factor may not be adjusted. Similarly, when the target instruction type with priority processing does not meet user needs, the instruction type weight can be adjusted according to user needs. For example, when the processing priority of a certain type of instruction is low, the instruction type weight of this type of instruction can be increased to improve the priority of the instruction for priority scheduling.
[0108] In specific application scenarios, it is also necessary to adjust the path depth weight or resource availability weight according to the target reward value. In this process, dynamically adjusting the path depth weight can optimize the behavior of the scheduler in different operating scenarios. For example, in delay-sensitive scenarios, increasing the target path depth weight can prioritize the scheduling of instructions on the critical path. In throughput-priority scenarios, reducing the target path depth weight can improve the overall scheduling flexibility. Therefore, when the target reward value indicates that the critical path instruction has a greater impact on the overall processing performance, increase the critical path weight to increase the scheduling priority of the critical path instruction and better reduce the delay of the critical path. When the target reward value indicates that there are more resource conflicts or lower resource utilization, increase the resource weight to strengthen the impact of resource status on instruction scheduling decisions, reduce resource conflicts, and improve resource utilization efficiency. By dynamically adjusting the above instruction weights, the reinforcement learning algorithm can adapt to different operating scenarios in real time, optimize the dynamic scheduling of instructions, and significantly improve the overall performance of the device.
[0109] In a specific application scenario, before adjusting the path depth weight or resource availability weight according to the target reward value, if the target reward value is greater than or equal to the first setting value, it indicates that there is no need to adjust the path depth weight or resource availability weight, so it can be detected whether the target reward value is less than the first setting value; in response to the target reward value being less than the first setting value, the step of adjusting the path depth weight or resource availability weight according to the target reward value is performed. Correspondingly, in response to the target reward value being less than the first setting value and greater than or equal to the second setting value, the path depth weight is adjusted; in response to the target reward value being less than the second setting value, the resource availability weight is adjusted.
[0110] It can be seen from the above implementation process that the instruction scheduling algorithm based on reinforcement learning of the present invention can optimize the scheduling strategy in real time according to the device operation status, such as target path depth, resource utilization, predicted delay, etc., by dynamically adjusting the instruction priority weight, thereby improving the flexibility and efficiency of instruction scheduling. Compared with traditional static rule scheduling, it has stronger adaptability in complex dynamic scenarios, such as multi-task loads, cache misses, resource competition, etc., and can effectively reduce pipeline pauses, optimize the execution order of critical path instructions, reduce resource conflicts, and ultimately improve the overall performance and throughput of the device.
[0111] In actual applications, in the process of adjusting the instruction weight, for any instruction weight, the historical value of the instruction weight can be collected in chronological order, and the required weight value can be selected from the historical value as the candidate weight value. For example, if the weight value needs to be increased, the historical value greater than the current weight value is used as the candidate weight value. Conversely, if the weight value needs to be reduced, the historical value less than the current weight value is used as the candidate weight value; the candidate weight values are sorted in chronological order to obtain the sorting result, and the sliding window size value is determined according to the weight adjustment strength. For example, if the weight adjustment strength is large, the sliding window value is determined to be large, and if the weight adjustment strength is small, the sliding window value is determined to be small; traverse forward from the tail of the sorting result according to the sliding window size value, if the traversed weight value meets the set adjustment margin value, the traversed weight value is used as the adjusted weight value, if the traversed value does not meet the set adjustment margin value, return to execute the step of traversing forward from the tail of the sorting result according to the sliding window size value; if there is no suitable value after the traversal is completed, the adjusted weight value is determined based on the current weight value and the adjustment margin value. In this way, the present invention can use historical weight values as reference weight values to adjust current weight values. The adjustment of weight values can be completed by simply moving the sliding window. The data is stable and there is little noise interference. The set adjustment margin can avoid drastic changes in the adjustment of weight values, prevent overfitting of weight value adjustment, ensure the stability of weight adjustment, and then ensure the stability of instruction scheduling.
[0112] The present invention provides an instruction scheduling method, which collects instruction information of instructions to be scheduled; determines the instruction weight currently corresponding to the instruction information; generates priority information of the instructions to be scheduled according to the instruction information and the instruction weight; selects a target instruction from the instructions to be scheduled for processing according to the priority information; collects processing information of the target instruction; generates a target reward value for evaluating instruction processing performance according to the processing information; and adjusts the instruction weight according to the target reward value, so as to process the instruction to be scheduled based on the adjusted instruction weight. The beneficial effects of the present invention are as follows: priority information of the instruction to be scheduled can be generated according to the instruction information and the current instruction weight, and the target instruction can be selected for processing according to the priority information, so that the instruction processing process is controlled by the instruction weight; and then a target reward value is generated according to the processing information of the target instruction, reflecting the processing performance of the target instruction from the data level, thereby realizing the quantification of the target instruction processing process; and then the instruction weight is adjusted according to the target reward value, which is equivalent to adjusting the instruction weight according to the processing performance of the target instruction from the data level, so that the adjusted instruction weight is compatible with the processing process of the target instruction, and if the target instruction is selected for processing later, it is equivalent to applying the adjusted instruction weight for instruction selection, thereby realizing dynamic adaptive scheduling of instructions, without relying on complex hardware structure, and having good flexibility.
[0113] In order to facilitate understanding of the improved instruction scheduling scheme of the present invention, the instruction scheduling process in the processor is now described. It is assumed that the instructions and related parameters in the processor instruction queue are as shown in Table 3, wherein the reinforcement learning model applied by the processor can be as follows: Figure 2 As shown, the state extraction module collects the operating status of the processor in real time, including instruction type, predicted instruction delay, critical path depth and resource status. The predicted instruction delay refers to the delay time when the L1 cache is predicted to miss; the scheduling execution module selects the instruction with the highest priority for emission; the reward feedback module calculates the reward value based on the number of successful critical path emissions, the number of pipeline pause cycles and the number of resource conflicts; the reinforcement learning algorithm module uses the reward value to dynamically adjust the weight in the priority formula to optimize future scheduling strategies. In addition, decoding, delay cache, critical path depth table, performance monitoring unit and reinforcement learning unit can be added to the processor hardware to assist in instruction scheduling. The decoding module obtains the instruction type by analyzing the instruction. The delay cache is used to record the historical delay data when the cache misses. The critical path depth table is used to record the critical path depth and dependency of the instruction to obtain the critical path depth. The performance monitoring unit is used to count the usage of hardware resources to obtain the resource availability status. The reinforcement learning unit is used to dynamically adjust the weight parameters according to the reward, such as Figure 3 shown.
[0114] Table 3 Example commands and related parameters in the command queue
[0115]
[0116] Assume that the initial critical path depth weight is 2.0, the resource availability weight is 0.5, and the smoothing factor is 1.0. The scheduling process is:
[0117] In the first round of scheduling, instructions I1, I4, and I6 can be scheduled immediately. The priorities of instructions I1, I4, and I6 are:
[0118] ;
[0119] ;
[0120] ;
[0121] The I6 instruction has the highest priority, so the I6 instruction is issued. Since the I6 instruction is a non-critical path, it does not increase the number of instructions successfully issued on the critical path, so the reward calculation for this round is , and assuming no conflicts , , this round of rewards The reward is very low, increasing the critical path depth weight from 2.0 to 2.3, improving the priority of critical path instructions, and giving priority to critical path instructions in the next round;
[0122] During the second scheduling, I6 has been issued, I1 and I4 are waiting to be scheduled, and I2 and I3 cannot be scheduled because the critical path depth weight has been increased. The priority of instructions I1 and I4 is:
[0123] ;
[0124] ;
[0125] The priority of instruction I4 remains unchanged, so issue instruction I3, and the reward calculation for this round is , , , this round of rewards The performance is good and the weights remain the same. But when running it in real life, the actual latency of I3 is 30 which is greater than the predicted value of 15.
[0126] In the fourth round of scheduling, I3, I6, and I1 have been launched, and I2, I4, and I5 can be scheduled. However, because I3 is actually delayed and has not been completed, I5 is not ready. The instruction queue of this round is not empty, but there is no instruction to be launched, and the pipeline is paused. Therefore, the reward calculation for this round is , , , this round of rewards Due to the pause caused by inaccurate delay prediction, the smoothing factor was adjusted from 1.0 to 1.3.
[0127] In the fifth round of scheduling, I3, I6, and I1 have been launched, and I2, I4, and I5 can be scheduled. Instruction I5 calculates the priority:
[0128] ;
[0129] Reward calculation for this round , , , this round of rewards . performs well, the weight remains unchanged, but the critical path ends, clearing the counter, so, , , .
[0130] In the sixth round of scheduling, I5, 3, I6, and I1 have been launched, and I2 and I4 can be scheduled. Assume that I2 is launched and there is a resource conflict. Reward calculation for this round , , , this round of rewards The reward is negative, increasing the resource availability weight from 0.5 to 0.7.
[0131] The seventh round of scheduling, the remaining instruction I4, recalculate the priority:
[0132] ;
[0133] Send command I4. Calculate rewards for this round , , , this round of rewards ,Since there are no subsequent critical instructions and no negative significance, the weight does not need to be adjusted.
[0134] The present invention also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of the instruction scheduling method provided in the embodiment of the present invention. Figure 4 , Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0135] An electronic device provided by an embodiment of the present invention includes a memory 201 and a processor 202. The memory 201 stores a computer program. When the processor 202 executes the computer program, the steps of the instruction scheduling method described in any of the above embodiments are implemented.
[0136] See also Figure 5Another electronic device provided in an embodiment of the present invention may also include: an input port 203 connected to the processor 202, used to transmit commands input from the outside to the processor 202; a display unit 204 connected to the processor 202, used to display the processing results of the processor 202 to the outside; a communication module 205 connected to the processor 202, used to realize communication between the electronic device and the outside. The display unit 204 can be a display panel, a laser scanning display, etc.; the communication method adopted by the communication module 205 includes but is not limited to mobile high-definition link technology (Mobile High-Definition Link, MHL), Universal Serial Bus (Universal Serial Bus, USB), High-Definition Multimedia Interface (High-Definition Multimedia Interface, HDMI), wireless connection: wireless fidelity technology (WIreless Fidelity, WiFi), Bluetooth communication technology, low-power Bluetooth communication technology, and communication technology based on IEEE802.11s.
[0137] An embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the instruction scheduling method described in any of the above embodiments are implemented.
[0138] The computer-readable storage medium involved in the present invention includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.
[0139] A computer program product provided by an embodiment of the present invention includes a computer program / instruction. When the computer program / instruction is executed by a processor, the following steps are implemented:
[0140] Collecting instruction information of instructions to be dispatched;
[0141] Determine the instruction weight currently corresponding to the instruction information;
[0142] Generate priority information of the instructions to be scheduled according to the instruction information and instruction weight;
[0143] According to the priority information, a target instruction is selected from the instructions to be scheduled for processing;
[0144] Collecting processing information of processing target instructions;
[0145] generating a target reward value for evaluating instruction processing performance based on the processing information;
[0146] The instruction weight is adjusted according to the target reward value, so as to process the instruction to be scheduled based on the adjusted instruction weight.
[0147] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: predicting the instruction delay of the instruction to be scheduled to obtain an estimated delay; collecting type information and target path depth of the instruction to be scheduled; collecting availability information of resources corresponding to the instruction to be scheduled; using the type information, estimated delay, target path depth and availability information as instruction information of the instruction to be scheduled; and determining the instruction weight currently corresponding to the instruction information, wherein the instruction weight includes an instruction type weight, a smoothing factor, a path depth weight, and a resource availability weight.
[0148] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: according to the positive correlation between priority and target path depth, path depth information is generated according to the target path depth and path depth weight; the path depth information is smoothed to obtain a path priority; according to the negative correlation between priority and estimated delay, a delay priority is generated according to the estimated delay and a smoothing factor; according to the positive correlation between priority and availability information, an availability priority is generated according to the availability information and resource availability weight; according to the positive correlation between priority and instruction type weight, priority information of instructions to be scheduled is generated according to the instruction type weight, path priority, delay priority and availability priority.
[0149] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: performing positive correlation operations on path priority and delay priority according to instruction type weight to generate an initial priority; performing positive correlation operations on the initial priority and availability priority to generate priority information of the instructions to be scheduled.
[0150] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: collecting the number of target path transmission successes for processing target instructions; collecting the number of pause cycles for processing target instructions; collecting the number of resource conflicts for processing target instructions; and using the number of target path transmission successes, the number of pause cycles, and the number of resource conflicts as processing information for processing target instructions.
[0151] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: obtaining a reward weight set corresponding to processing information; performing a weighted operation on the processing information based on the reward weight to obtain a target reward value for evaluating instruction processing performance.
[0152] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: obtaining a first adjustment weight set for adjusting the impact of path transmission on rewards; obtaining a second adjustment weight set for adjusting the impact of pauses on rewards; obtaining a third adjustment weight set for adjusting the impact of resource conflicts on rewards; and using the first adjustment weight, the second adjustment weight, and the third adjustment weight as reward weights corresponding to processing information.
[0153] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: determining a first product value of a first adjustment weight and the number of successful transmissions of a target path; determining a second product value of a second adjustment weight and the number of pause cycles; generating a first difference between the first product value and the second product value; determining a third product value of a third adjustment weight and the number of resource conflicts; generating a second difference between the first difference and the third product value; and using the second difference as a target reward value for evaluating instruction processing performance.
[0154] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: obtaining an actual delay of a target instruction; detecting whether a difference between the actual delay and the estimated delay is greater than a preset threshold; in response to the difference between the actual delay and the estimated delay being greater than the preset threshold, adjusting a smoothing factor, and adjusting a path depth weight or a resource availability weight according to a target reward value; in response to the difference between the actual delay and the estimated delay being less than or equal to the preset threshold, adjusting a path depth weight or a resource availability weight according to the target reward value.
[0155] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: before adjusting the instruction type weight or the resource availability weight according to the target reward value, detecting whether the target reward value is less than a first set value; in response to the target reward value being less than the first set value, executing the step of adjusting the path depth weight or the resource availability weight according to the target reward value.
[0156] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which implements the following steps when executed by a processor: in response to a target reward value being less than a first set value and greater than or equal to a second set value, adjusting a path depth weight.
[0157] A computer program product provided by an embodiment of the present invention includes a computer program / instruction. When the computer program / instruction is executed by a processor, the following steps are implemented: in response to a target reward value being less than a second set value, adjusting a resource availability weight.
[0158] For the description of the relevant parts of an electronic device, a computer program product, and a computer-readable storage medium provided by an embodiment of the present invention, please refer to the detailed description of the corresponding parts in an instruction scheduling method provided by an embodiment of the present invention, which will not be repeated here. In addition, the parts of the above technical solutions provided by the embodiments of the present invention that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.
[0159] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0160] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An instruction scheduling method, characterized in that: include: Collecting instruction information of instructions to be dispatched; Determining a current instruction weight corresponding to the instruction information; Generate priority information of the instruction to be scheduled according to the instruction information and the instruction weight; According to the priority information, a target instruction is selected from the instructions to be scheduled for processing; Collecting and processing the processing information of the target instruction; generating a target reward value for evaluating instruction processing performance based on the processing information; The instruction weight is adjusted according to the target reward value, so as to process the instruction to be scheduled based on the adjusted instruction weight.
2. The instruction scheduling method according to claim 1, characterized in that: Collect instruction information of the instructions to be scheduled, including: Predict the instruction delay of the instruction to be scheduled and obtain the estimated delay; Collect the type information and target path depth of the instructions to be scheduled; Collecting availability information of resources corresponding to the instructions to be scheduled; Using the type information, the estimated delay, the target path depth and the availability information as instruction information of the instruction to be scheduled; Determining the instruction weight currently corresponding to the instruction information includes: Determine the instruction weight currently corresponding to the instruction information, wherein the instruction weight includes an instruction type weight, a smoothing factor, a path depth weight, and a resource availability weight.
3. The instruction scheduling method according to claim 2, characterized in that: Generate priority information of the instruction to be scheduled according to the instruction information and the instruction weight, including: According to the positive correlation between the priority and the target path depth, generating path depth information according to the target path depth and the path depth weight; Smoothing the path depth information to obtain path priority; According to the negative correlation between the priority and the estimated delay, generating a delay priority according to the estimated delay and the smoothing factor; According to the positive correlation between the priority and the availability information, generating the availability priority according to the availability information and the resource availability weight; According to the positive correlation between the priority and the instruction type weight, the priority information of the instruction to be scheduled is generated according to the instruction type weight, the path priority, the delay priority and the availability priority.
4. The instruction scheduling method according to claim 3, characterized in that: According to the positive correlation between the priority and the instruction type weight, priority information of the instruction to be scheduled is generated according to the instruction type weight, the path priority, the delay priority and the availability priority, including: According to the instruction type weight, a positive correlation operation is performed on the path priority and the delay priority to generate an initial priority; A positive correlation operation is performed on the initial priority and the availability priority to generate priority information of the instruction to be scheduled.
5. The instruction scheduling method according to any one of claims 2 to 4, characterized in that: Collecting and processing the processing information of the target instruction includes: Collect and process the number of successful transmissions of the target path of the target instruction; Collecting the number of pause cycles of processing the target instruction; Collect and process the number of resource conflicts of the target instruction; The target path transmission success number, the stall cycle number and the resource conflict number are used as processing information for processing the target instruction.
6. The instruction scheduling method according to claim 5, characterized in that: Generating a target reward value for evaluating instruction processing performance according to the processing information includes: Obtaining a reward weight set corresponding to the processing information; The processing information is weighted based on the reward weight to obtain a target reward value for evaluating instruction processing performance.
7. The instruction scheduling method according to claim 6, characterized in that: Obtaining a reward weight set corresponding to the processing information, including: Obtaining a first adjustment weight set for adjusting the effect of path emission on the reward; Obtaining a second adjustment weight set for adjusting the effect of the pause on the reward; Obtaining a third adjustment weight set for adjusting the impact of resource conflict on rewards; The first adjustment weight, the second adjustment weight and the third adjustment weight are used as reward weights corresponding to the processing information.
8. The instruction scheduling method according to claim 7, characterized in that: The processing information is weighted based on the reward weight to obtain a target reward value for evaluating instruction processing performance, including: Determine a first product value of the first adjustment weight and the number of successful transmissions of the target path; Determine a second product value of the second adjustment weight and the number of pause cycles; generating a first difference between the first product value and the second product value; Determining a third product value of the third adjustment weight and the number of resource conflicts; generating a second difference value between the first difference value and the third product value; The second difference is used as a target reward value for evaluating instruction processing performance.
9. The instruction scheduling method according to claim 5, characterized in that: According to the target reward value, the instruction weight is adjusted, including: Obtaining an actual delay of the target instruction; Detecting whether a difference between the actual delay and the estimated delay is greater than a preset threshold; In response to a difference between the actual delay and the estimated delay being greater than the preset threshold, adjusting the smoothing factor, and adjusting the path depth weight or the resource availability weight according to the target reward value; In response to a difference between the actual delay and the estimated delay being less than or equal to the preset threshold, the path depth weight or the resource availability weight is adjusted according to the target reward value.
10. The instruction scheduling method according to claim 9, characterized in that: Before adjusting the path depth weight or the resource availability weight according to the target reward value, the method further includes: Detecting whether the target reward value is less than a first set value; In response to the target reward value being less than the first set value, a step of adjusting the path depth weight or the resource availability weight according to the target reward value is performed.
11. The instruction scheduling method according to claim 10, characterized in that: According to the target reward value, adjusting the path depth weight or the resource availability weight includes: In response to the target reward value being less than the first setting value and greater than or equal to a second setting value, the path depth weight is adjusted.
12. The instruction scheduling method according to claim 11, characterized in that: According to the target reward value, adjusting the path depth weight or the resource availability weight includes: In response to the target reward value being less than the second set value, the resource availability weight is adjusted.
13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the instruction scheduling method according to any one of claims 1 to 12 are implemented.
14. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the instruction scheduling method according to any one of claims 1 to 12 when executing the computer program.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the instruction scheduling method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Multi-resource index Kubernetes scheduling method based on delay factor
CN112363827A
Method and device for polling scheduling of processor pipeline instructions
CN116521234A
Multi-user downlink scheduling method and device, equipment and storage medium
CN117545085A
Scheduling automation system application state management method
CN119292745A
Dynamic task scheduling method and device, server and storage medium
CN119336466A
Cited By
Instruction scheduling method and device, electronic equipment, storage medium and computer program product
CN121187653A