An instruction scheduling processing method and device, a storage medium and an electronic device
Patent Information
- Application Number
- CN202211588891.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-12
AI Technical Summary
[0004]本申请实施例提供了一种指令调度处理方法、装置、存储介质及电子装置,以至少解决相关技术中对寄存器分配前的指令调度采用双向调度,存在依赖关系的已调度指令和未调度指令间的启发量联系不明确,导致调度效果不佳的问题
[0017]本申请实施例,将当前待调度的基本块的待调度指令按照数据依赖关系构建有向无环图,所述有向无环图由多个节点构成,每个节点对应一个待调度指令,节点之间的连线为节点需等待的时钟周期;按照自顶而下与自底而上获取当前调度的多个指令;分别根据所述有向无环图中已调度指令和未调度指令确定所述多个指令的目标启发量;根据所述目标启发量对所述多个指令进行调度,可以解决相关技术中对寄存器分配前的指令调度采用双向调度,存在依赖关系的已调度指令和未调度指令间的启发量联系不明确,导致调度效果不佳的问题,通过有向无环图将存在依赖关系的已调度指令和未调度指令联系起来,根据启发量进行指令调度,提高了调度效果。
Smart Images

Figure CN118193060B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and more specifically, to an instruction scheduling processing method, apparatus, storage medium, and electronic device. Background Technology
[0002] Currently, the LLVM compiler uses bidirectional scheduling for instruction scheduling before register allocation. It uses a combinational heuristic scheduling algorithm to schedule code by comprehensively considering register pressure, boundary weight parameters, resource judgment based on deterministic finite automata (DFA), and dependent instructions. However, it does not take into account the influence of scheduled instructions within a basic block on the heuristic of unscheduled instructions, resulting in poor scheduling performance.
[0003] There is no solution yet to address the problem that bidirectional scheduling is used in related technologies for instruction scheduling before register allocation, where the heuristic relationship between scheduled and unscheduled instructions with dependencies is unclear, leading to poor scheduling performance. Summary of the Invention
[0004] This application provides an instruction scheduling processing method, apparatus, storage medium, and electronic device to at least solve the problem in related technologies where instruction scheduling before register allocation uses bidirectional scheduling, resulting in unclear heuristic relationships between scheduled and unscheduled instructions with dependencies, leading to poor scheduling performance.
[0005] According to one embodiment of this application, an instruction scheduling processing method is provided, the method comprising:
[0006] The scheduled instructions of the current basic block to be scheduled are constructed into a directed acyclic graph according to the data dependency relationship. The directed acyclic graph consists of multiple nodes, each node corresponds to a scheduled instruction, and the connection between nodes is the clock cycle that the node needs to wait.
[0007] Obtain multiple instructions for the current schedule from both top-down and bottom-up perspectives;
[0008] The target heuristics of the plurality of instructions are determined based on the scheduled and unscheduled instructions in the directed acyclic graph, respectively.
[0009] The multiple instructions are scheduled according to the target heuristic.
[0010] According to another embodiment of this application, an instruction scheduling processing apparatus is also provided, the apparatus comprising:
[0011] The construction module is used to construct a directed acyclic graph (DAG) of the scheduled instructions of the currently scheduled basic blocks according to the data dependency relationship. The DAG consists of multiple nodes, each node corresponds to a scheduled instruction, and the connection between nodes is the clock cycle that the node needs to wait for.
[0012] The acquisition module is used to acquire multiple currently scheduled instructions in a top-down and bottom-up manner.
[0013] The determination module is used to determine the target heuristic of the plurality of instructions based on the scheduled instructions and unscheduled instructions in the directed acyclic graph, respectively;
[0014] The scheduling module is used to schedule the plurality of instructions according to the target heuristic.
[0015] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0016] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0017] In this embodiment, a directed acyclic graph (DAG) is constructed based on data dependencies among the scheduled instructions of the current basic block. The DAG consists of multiple nodes, each corresponding to a scheduled instruction, with the connections between nodes representing the clock cycles the nodes need to wait. Multiple currently scheduled instructions are obtained from both top-down and bottom-up perspectives. Target heuristics for these instructions are determined based on the scheduled and unscheduled instructions in the DAG. The instructions are then scheduled according to these target heuristics. This approach addresses the problem in related technologies where bidirectional scheduling before register allocation leads to unclear heuristic relationships between scheduled and unscheduled instructions with dependencies, resulting in poor scheduling performance. By connecting the scheduled and unscheduled instructions with dependencies through the DAG and scheduling instructions based on heuristics, the scheduling efficiency is improved. Attached Figure Description
[0018] Figure 1 This is a hardware structure block diagram of a mobile terminal for the instruction scheduling and processing method according to an embodiment of this application;
[0019] Figure 2 This is a flowchart of an instruction scheduling processing method according to an embodiment of this application;
[0020] Figure 3 This is a flowchart of an instruction scheduling processing method according to an optional embodiment of this application;
[0021] Figure 4 This is a schematic diagram of a directed acyclic graph according to this embodiment. Figure 1 ;
[0022] Figure 5 This is a schematic diagram of the height and depth of a directed acyclic graph according to this embodiment;
[0023] Figure 6 This is a flowchart of IHL scheduling according to this embodiment;
[0024] Figure 7 This is a schematic diagram of a directed acyclic graph according to this embodiment. Figure 2 ;
[0025] Figure 8 This is a block diagram of an instruction scheduling processing apparatus according to an embodiment of this application. Detailed Implementation
[0026] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0028] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for the instruction scheduling and processing method according to an embodiment of this application, as shown below. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or programmable logic device, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0029] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the instruction scheduling processing method in this embodiment. The processor 102 executes various functional applications and instruction scheduling processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0030] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0031] This embodiment provides an instruction scheduling and processing method that runs on the aforementioned mobile terminal or network architecture. Figure 2 This is a flowchart of an instruction scheduling processing method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0032] Step S202: Construct a directed acyclic graph (DAG) of the scheduled instructions of the current basic block to be scheduled according to the data dependency relationship. The DAG consists of multiple nodes, each node corresponds to a scheduled instruction, and the connection between nodes is the clock cycle that the node needs to wait for.
[0033] Step S204: Obtain multiple currently scheduled instructions in a top-down and bottom-up manner;
[0034] Step S206: Determine the target heuristic of the plurality of instructions based on the scheduled instructions and unscheduled instructions in the directed acyclic graph, respectively;
[0035] Step S208: Schedule the plurality of instructions according to the target heuristic.
[0036] Through the above steps S202 to S208, the problem of poor scheduling effect caused by the unclear heuristic relationship between scheduled and unscheduled instructions with dependencies when bidirectional scheduling is used for instruction scheduling before register allocation in related technologies can be solved. By using a directed acyclic graph to link scheduled and unscheduled instructions with dependencies, instruction scheduling is performed according to heuristics, thus improving the scheduling effect.
[0037] Figure 3 This is a flowchart of an instruction scheduling processing method according to an optional embodiment of this application, such as... Figure 3 As shown, step S208 above may specifically include:
[0038] Step S302: In the current clock cycle, the plurality of instructions are scheduled by repeatedly executing the following steps: selecting the target scheduling instruction from the plurality of instructions according to the heuristic; if the target scheduling instruction is a top-down obtained instruction, saving the target scheduling instruction to the first queue; if the target scheduling instruction is a bottom-up obtained instruction, saving the target scheduling instruction to the second queue.
[0039] Step S304: If the target instruction among the plurality of instructions is not scheduled in the current clock cycle, the target instruction is scheduled in the next clock cycle, wherein the target instruction is one or more instructions.
[0040] In one embodiment, step S302 above, selecting the target scheduling instruction for the current scheduling from the plurality of instructions based on the heuristic quantity, may specifically include:
[0041] S3021, if the heuristic value of the current instruction is equal to the heuristic value of the pre-selected instruction and the current instruction is a boundary instruction, if the current traversal is top-down, the target scheduling instruction is determined according to the number of child nodes of the multiple instructions; if the current traversal is bottom-up, the target scheduling instruction is determined according to the number of parent nodes of the multiple instructions.
[0042] S3022, if the heuristic value of the current instruction is equal to the heuristic value of the pre-selected instruction, and both the parent and child instructions of the current instruction and the pre-selected instruction have been scheduled, if the unscheduled sequence is not empty, the instruction with the most instructions in the unscheduled sequence is determined as the target scheduling instruction; if the scheduled sequence is not empty, the instruction with the most instructions in the scheduled sequence is determined as the target scheduling instruction; if both the unscheduled sequence and the scheduled sequence are empty, the target scheduling instruction is determined based on the height or depth of the multiple instructions.
[0043] Furthermore, the aforementioned S3021 may specifically include:
[0044] If the current traversal is top-down, and the instruction with the most child nodes is a single instruction, then the instruction with the most child nodes is determined to be the target scheduling instruction; if the instruction with the most child nodes is at least two instructions, then the index values of the at least two instructions are compared, and the instruction with the smallest index value is determined to be the target scheduling instruction.
[0045] If the current traversal is bottom-up, and the instruction with the most parent nodes is a single instruction, then the instruction with the most parent nodes is determined to be the target scheduling instruction; if the instruction with the most parent nodes is at least two instructions, then the index values of the at least two instructions are compared, and the instruction with the largest index value is determined to be the target scheduling instruction.
[0046] In another embodiment, in S3022 above, if both the unscheduled sequence and the scheduled sequence are empty, determining the target scheduling instruction based on the height or depth of the plurality of instructions may specifically include: if the current traversal is top-down, if the instruction with the largest height is a single instruction, determining the instruction with the largest height as the target scheduling instruction; if the instruction with the largest height is at least two instructions, determining the instruction with the smaller depth among the at least two instructions as the target scheduling instruction; if the traversal is bottom-up, if the instruction with the largest depth is a single instruction, determining the instruction with the largest depth as the target scheduling instruction; if the instruction with the largest depth is at least two instructions, determining the instruction with the smallest height among the at least two instructions as the target scheduling instruction.
[0047] In this embodiment, step S206 may specifically include:
[0048] S2061, determine the overlap point of the current instruction in the critical path of the currently scheduled instructions, and further, take the child node of the node where the current instruction is located as the critical node of the scheduled instructions; determine the child node of the node where the scheduled instructions are located; determine the node with the maximum height or maximum depth of the intersection node of the child node of the node where the scheduled instructions are located and the critical node as the overlap point of the current instruction in the critical path of the currently scheduled instructions, wherein the current instruction is each of the plurality of instructions;
[0049] S2062, determine the difference between the height of the current instruction and the overlapping point as a first heuristic value, or determine the difference between the depth of the current instruction and the overlapping point as a first heuristic value;
[0050] S2063, obtain the longest path of the basic block; determine the difference between the longest path and the first heuristic as the second heuristic;
[0051] S2064, determine the proportion of the number of scheduled parent or child instructions of the current instruction to all parent and child instructions of the current instruction as a first weighting coefficient, and determine the proportion of the number of unscheduled parent or child instructions to all parent and child instructions of the current instruction as a second weighting coefficient.
[0052] S2065, determine the height and depth of the scheduled parent or child instruction of the current instruction and the proportion of them to the longest path as a third weighting coefficient, and determine the height and depth of the unscheduled parent or child instruction of the current instruction and the proportion of them to the longest path as a fourth weighting coefficient.
[0053] S2066, the target heuristic of the current instruction is obtained by adjusting the first heuristic and the second heuristic based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient. Further, if the current traversal is top-down, the target heuristic of the current instruction is obtained by adjusting the first heuristic and the second heuristic based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient. Specifically, the target heuristic H can be adjusted in the following way. ′ :
[0054]
[0055] If the current traversal is top-down, the target heuristic of the current instruction is obtained by adjusting the first heuristic and the second heuristic based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient. Specifically, the target heuristic D can be adjusted in the following way. ′ :
[0056]
[0057] in, The first weighting coefficient, This is the second weighting coefficient. The third weighting coefficient, Here, denoted as the fourth weighting coefficient, MaxHeight is the maximum height, HitLevel1 is the height of the overlapping point, MaxDepth is the maximum depth, HitLevel2 is the depth of the overlapping point, D is the depth of the current instruction, and H is the height of the current instruction.
[0058] NIS / NIS+IS refers to the proportion of the sum of the number of parent nodes or child nodes of the current instruction that are not on the critical path to the sum of the total number of parent nodes or child nodes of the current instruction.
[0059] (NHCnt+NDCnt) / MaxLevel refers to the proportion of the sum of the height and depth of all parent or child nodes that are not on the critical path of the current instruction to the longest path of the basic block.
[0060] (IDCnt+IHCnt) / MaxLevel refers to the proportion of the sum of the height and depth of all parent or child nodes of the current instruction on the critical path to the longest path of the basic block.
[0061] IS / NIS+IS refers to the proportion of the sum of the number of parent nodes or child nodes of the current instruction on the critical path to the sum of the number of parent nodes or child nodes of the current instruction.
[0062] In an optional embodiment, the method further includes: adjusting the instructions of the scheduled basic block by one of the following methods: adjusting the target storage instructions of the basic block that are far from parent instructions; adjusting the target loading instructions of the basic block that are far from parent and child instructions; adjusting the distance between instructions without parent and child instructions; adjusting the gaps between instructions due to weight; and adjusting the position of the storage instructions within the basic package.
[0063] This embodiment stores the scheduled instructions in the bidirectional scheduling Top and Bot, uses the invented heuristic calculation function to link the scheduled and unscheduled instructions with dependencies, and then selects the appropriate instruction for scheduling by using an improved method based on heuristics and other information.
[0064] After the basic block scheduling is completed, the instructions that have been scheduled in Top and Bot are merged according to the proposed merging method. Then, the smallest primitive Packet in the basic block is fine-tuned according to the proposed fine-tuning principles and the characteristics of special instructions. Without affecting the functional correctness of the basic block, the register pressure of the current basic block, the characteristics of the processor, and the influence between dependent instructions in the entire basic block are taken into account. This makes the number of smallest primitives in the basic block infinitely close to the longest path of the current basic block, so that the processor can execute more instructions in fewer clock cycles, thereby improving the processor's computing performance.
[0065] This embodiment can be applied to the performance optimization of compilers for multi-issue processors. By scheduling instructions, the number of concurrent instructions executed in each clock cycle of the multi-issue processor reaches the number of concurrent units of the processor. The compiler of the multi-issue processor generates a hardware-executable binary file after a series of optimizations of the input application code. Through the instruction scheduling of this invention, each unit of the multi-issue processor is fully loaded, so as to complete the same operation in fewer clock cycles.
[0066] Instruction scheduling is a process that modifies the arrangement of instructions according to certain rules without compromising correctness. It leverages the parallel capabilities of multi-issue processors to maximize their parallelism, thereby enabling more instruction tasks to be completed in fewer clock cycles.
[0067] During instruction scheduling, the instructions to be scheduled are arranged into a DAG (Directed Acyclic Graph) according to their dependencies. A DAG consists of several nodes, and the connection between two nodes is called a loop. No closed loop can be formed between any nodes in a DAG. Figure 4 This is a schematic diagram of a directed acyclic graph according to this embodiment. Figure 1 ,like Figure 4 As shown, the DAG (Directed Acyclic Graph) consists of 13 nodes, N(0) to N(13), and the number between two nodes represents the number of clock cycles the node needs to wait. By analyzing... Figure 4 The height of the DAG is calculated from bottom to top as Hi = max(D(N(i), N(j))), and the depth of the DAG is calculated from top to bottom as D. i =max(D(N(i), N(j))), Figure 5 This is a schematic diagram of the height and depth of the directed acyclic graph according to this embodiment. The updated diagram is shown in Figure 5.
[0068] Figure 6 This is a flowchart of IHL scheduling according to this embodiment, as follows: Figure 6 As shown, the instruction scheduling in this embodiment includes:
[0069] Step S601, Construct DAG graph: Create DAG graph (Directed Acyclic Graph) according to the data dependencies between instructions in the basic block;
[0070] Step S602, Initialize parameters: Initialize the parameters and variables required by the scheduling subsystem;
[0071] Step S603: Determine whether all instructions have been scheduled. If the determination result is yes, proceed to step S607. If the determination result is no, proceed to step S604.
[0072] Step S604: Through the schedulable instruction acquisition subsystem, schedulable instructions are saved from the unscheduled instructions of the currently scheduled basic block according to the data dependency relationship, in both top-down and bottom-up order.
[0073] Step S605: Calculate scheduling heuristics for each schedulable instruction using the heuristic calculation subsystem.
[0074] Step S606: The instruction selection subsystem selects the best instruction based on the heuristic of the schedulable instructions and stores it in the corresponding queue. Top-down schedulable instructions are stored in Top, and bottom-up schedulable instructions are stored in Bot.
[0075] Step S607: Perform IHS fine-tuning through the IHS fine-tuning subsystem.
[0076] By saving the Top and all scheduled instructions within the Bot, and modifying the heuristic of the scheduling instructions, the current basic block is further transformed based on the characteristics of the target machine after scheduling. This employs a hybrid heuristic semi-global scheduling approach, also known as IHS (Improved Heuristic Scheduler Algorithm). Here, "semi-global" specifically refers to the currently scheduled basic block.
[0077] The overall scheduling process has been improved by adding an IHS fine-tuning subsystem and new calculation formulas. IHS completes scheduling by adding a schedulable instruction acquisition subsystem, an IHS fine-tuning subsystem, and expanding the functionality of the heuristic calculation subsystem and instruction selection subsystem.
[0078] The heuristic computation subsystem is the core of the IHS heuristic computation subsystem and also the core of semi-global scheduling. HitLevel specifically refers to the maximum depth (Bot) or height (Top) of the set of critical path overlap points between the current instruction's successor or predecessor dependent instructions and the scheduled instructions.
[0079] Figure 7 This is a schematic diagram of a directed acyclic graph according to this embodiment. Figure 2 ,like Figure 7 As shown, SU(0) and SU(1) are ready according to the Top-Down scheduling instructions, and SU(11), SU(12), and SU(13) are ready according to the Bot-UP scheduling instructions. The heuristic calculation subsystem calculates a Cost for the above-mentioned ready SUs and determines whether to schedule nodes in the Top-Down or Bot-Up based on certain conditions. HitLevel is calculated based on the Depth of the SU in the Bot-Up and based on the Height of the SU in the Top-Down.
[0080] Taking Top-Down as an example, suppose that only SU(0) is scheduled now, and the child nodes 2, 5, 6, 9, 11, 12, and 13 of SU(0) are taken as the critical nodes of the scheduled instructions. At present, only SU(1) is ready in the Top-Down Schedule, and the HitLevel of SU(1) needs to be calculated. By traversing the child nodes of SU(1), it is found that SU(9) is both a child node of the current node and a critical node. Therefore, the Height of SU(9) is the HitLevel of SU(1). Of course, the specific situation may be much more complicated than this. There may be multiple SUs that are both child nodes and critical nodes. In this case, the largest Height is taken for Top-Down.
[0081] Taking Bot-Up as an example, assuming that only SU(12) is currently scheduled, the parent nodes of SU(12), namely nodes 9, 8, 7, 6, 5, 4, 3, 2, 1, and 0, are designated as critical nodes. Currently, SU(13) and SU(11) are ready in Bot-Up. We search for SUs that are designated as child nodes and are among the critical nodes. As shown in the figure, SU(6) is both a parent node and a critical node. Therefore, the Depth of SU(6) is their common HitLevel.
[0082] Based on the original heuristic value calculation subsystem, and combined with HitLevel, a new calculation formula has been added to the heuristic value calculation subsystem.
[0083] For Top-Down SU, the calculation formula is as follows:
[0084]
[0085] For Bot-Up SU, the calculation formula is as follows:
[0086]
[0087] The above formula combines the distance between the current SU height or depth and the critical node, as well as the weight of its parent or child node in the critical path, to comprehensively adjust the heuristic. Considering the overall impact on the heuristic, its value range is controlled within [-MaxLevel, MaxLevel] by MaxLevel.
[0088] Its calculation process includes:
[0089] Obtain the point of overlap between the current scheduled instruction and the critical path of the currently scheduled instructions, denoted as HitLevel;
[0090] The difference between the height (or depth) of the current scheduling instruction and HitLevel is recorded as the first heuristic.
[0091] Get the longest path of the current basic block, denoted as MaxLevel;
[0092] Obtain the difference between MaxLevel and the first heuristic, and denote it as the second heuristic;
[0093] The proportion of the number of scheduled parent or child instructions of the current scheduling instruction to the total number of all parent and child instructions of the current instruction is denoted as weight coefficient one; the proportion of the number of unscheduled parent or child instructions to the total number of parent and child instructions of the current instruction is denoted as weight coefficient two.
[0094] Get the height and depth of the currently scheduled parent or child instructions and their proportion of MaxLevel, denoted as weight coefficient three; get the height and depth of the currently scheduled parent or child instructions and their proportion of MaxLevel, denoted as weight coefficient four.
[0095] The final heuristic is the value of the first and second heuristics adjusted by weighting coefficients one, two, three, and four.
[0096] The instruction selection subsystem selects the optimal instruction scheduler by comparing the heuristic values of all schedulable instructions. Its selection criteria are as follows:
[0097] If the heuristic value of the current instruction is equal to the heuristic value of the pre-selected instruction and the current instruction is a boundary instruction;
[0098] The current traversal is Top, storing instructions with a large number of child nodes;
[0099] If both have the same number of child nodes, compare their index values and save the instruction with the smaller index value.
[0100] The current traversal is of the Bot type, storing instructions with a large number of parent nodes;
[0101] If both have the same number of parent nodes, compare their index values and save the instruction with the larger index value.
[0102] If the heuristic value of the current instruction is equal to the heuristic value of the preselected instruction, and both the current instruction and the preselected instruction's parent or child instruction have been scheduled.
[0103] If the unscheduled sequence is not empty, then save the instruction with the larger number of unscheduled instructions;
[0104] If the scheduled sequence is not empty, save the instruction with the larger number of scheduled instructions.
[0105] If both are empty;
[0106] If the traversal reaches the top, save the instructions with the larger height;
[0107] If Height is equal, save the instruction with the smaller Depth.
[0108] If Depth is equal, save the instruction with the smaller index value;
[0109] If the traversal is a Bot, save the instructions with the larger Depth;
[0110] If Depth is equal, save the instruction with the smaller Height.
[0111] If Height is equal, save the instruction with the larger index value;
[0112] The IHS fine-tuning subsystem reduces register pressure between instructions by adjusting instructions within scheduled basic blocks. The fine-tuning criteria include the following:
[0113] Adjust the distance between the Store directive and its parent directive if the distance is too large;
[0114] Adjust the Load instruction that has a large distance between it and its parent and child instructions;
[0115] Adjust the distance between instructions without a parent class and their child classes;
[0116] Adjust the stall caused by weighting between instructions;
[0117] Adjust the position of the Store within the base package, moving it to the front of the base package where it can be placed.
[0118] It mainly includes pre-packaging, Store directive adjustment, Load directive adjustment, general directive adjustment, and Store post-adjustment.
[0119] Pre-packaging module: Packs the scheduled basic blocks in sequence according to resources and related constraints to complete the pre-packaging of the entire basic block.
[0120] Store instruction adjustments: These release registers by reducing the lifetime of the registers occupied by the Store instruction, including:
[0121] Iterate through all Stores and find the Preds (parent class) of the Store directives.
[0122] max_index = max(Pred i .index+Latency i );
[0123] Distance=Store.index-max_index;
[0124] If Distance is greater than zero, save the corresponding Store instruction;
[0125] Sort all Stores with the same Distance by their index value;
[0126] Move the Store's position in the pre-packaged order from largest to smallest Distance.
[0127] Load instruction adjustment: This involves freeing up registers by reducing the lifetime of the registers occupied by the Load instruction. The steps include:
[0128] Iterate through all Load instructions and find the Succs (subclasses) of the Load instructions.
[0129] min_index = min(Succ) i .index-Latency i );
[0130] max_index = max(Pred j .index+Latency j );
[0131] If min_index > 0 and max_index < 0, then
[0132] Distance=min_index-LD.index;
[0133] If min_index < 0 and max_index > 0, then
[0134] Distance=max(max_index,has_anti?PIndex.max:(PIndex.max-Cnt))-LD.index;
[0135] If min_index > 0 and max_index > 0, then
[0136] Distance=min_index-max(LD.index, max_index).
[0137] Sort all Loads with the same Distance by index value, and move the Loads in the pre-packaging according to their Distance from largest to smallest.
[0138] General instruction adjustments: First, adjust instructions without parent classes (MOVI, CONST, etc.) so that they do not occupy extra valid space in the pre-packaging; second, add PS_DelayLSU_I to solve the problem of LD-ST-SPU sharing the same register, which leads to more stalls (i.e. nops).
[0139] According to another aspect of the embodiments of this application, an instruction scheduling processing apparatus is also provided. Figure 8 This is a block diagram of an instruction scheduling processing apparatus according to an embodiment of this application, such as... Figure 8 As shown, the device includes:
[0140] Module 82 is used to construct a directed acyclic graph (DAG) of the scheduled instructions of the currently scheduled basic blocks according to the data dependency relationship. The DAG consists of multiple nodes, each node corresponds to a scheduled instruction, and the connection between nodes is the clock cycle that the node needs to wait for.
[0141] The acquisition module 84 is used to acquire multiple currently scheduled instructions in a top-down and bottom-up manner;
[0142] The determination module 86 is used to determine the target heuristic of the plurality of instructions based on the scheduled instructions and unscheduled instructions in the directed acyclic graph, respectively;
[0143] The scheduling module 88 is used to schedule the plurality of instructions according to the target heuristic.
[0144] In one embodiment, the scheduling module 88 includes:
[0145] The repeated scheduling submodule is used to schedule the plurality of instructions by repeatedly executing the following steps during the current clock cycle: selecting the target scheduling instruction from the plurality of instructions according to the heuristic; if the target scheduling instruction is a top-down obtained instruction, saving the target scheduling instruction to a first queue; if the target scheduling instruction is a bottom-up obtained instruction, saving the target scheduling instruction to a second queue.
[0146] The scheduling submodule is used to schedule the target instruction in the next clock cycle if the target instruction among the plurality of instructions is not scheduled in the current clock cycle, wherein the target instruction is one or more instructions.
[0147] In one embodiment, the repeat scheduling submodule includes:
[0148] The first determining unit is configured to, when the heuristic value of the current instruction is equal to the heuristic value of the pre-selected instruction and the current instruction is a boundary instruction, determine the target scheduling instruction based on the number of child nodes of the plurality of instructions if the current traversal is top-down; and determine the target scheduling instruction based on the number of parent nodes of the plurality of instructions if the current traversal is bottom-up.
[0149] The second determining unit is configured to, when the heuristic value of the current instruction is equal to the heuristic value of the pre-selected instruction and both the current instruction and the parent or child instructions of the pre-selected instruction have been scheduled, determine the instruction with the most instructions in the unscheduled sequence as the target scheduling instruction if the unscheduled sequence is not empty; determine the instruction with the most instructions in the scheduled sequence as the target scheduling instruction if the scheduled sequence is not empty; and determine the target scheduling instruction based on the height or depth of the multiple instructions if both the unscheduled sequence and the scheduled sequence are empty.
[0150] In one embodiment, the first determining unit is further configured to: if the current traversal is top-down, and if the instruction with the most child nodes is a single instruction, determine the instruction with the most child nodes as the target scheduling instruction; if the instruction with the most child nodes is at least two instructions, compare the index values of the at least two instructions to determine the instruction with the smallest index value as the target scheduling instruction; if the current traversal is bottom-up, and if the instruction with the most parent nodes is a single instruction, determine the instruction with the most parent nodes as the target scheduling instruction; if the instruction with the most parent nodes is at least two instructions, compare the index values of the at least two instructions to determine the instruction with the largest index value as the target scheduling instruction.
[0151] In one embodiment, the second determining unit is further configured to: if the current traversal is top-down, and if the instruction with the largest height is a single instruction, determine the instruction with the largest height as the target scheduling instruction; if the instruction with the largest height is at least two instructions, determine the instruction with the smallest depth among the at least two instructions as the target scheduling instruction; if the traversal is bottom-up, and if the instruction with the largest depth is a single instruction, determine the instruction with the largest depth as the target scheduling instruction; if the instruction with the largest depth is at least two instructions, determine the instruction with the smallest height among the at least two instructions as the target scheduling instruction.
[0152] In one embodiment, the determining module includes:
[0153] The first determining submodule is used to determine the overlap point of the current instruction in the critical path of the currently scheduled instructions, wherein the current instruction is each of the plurality of instructions;
[0154] The second determining submodule is used to determine the difference between the height of the current instruction and the overlapping point as a first heuristic value, or to determine the difference between the depth of the current instruction and the overlapping point as a first heuristic value;
[0155] The third determining submodule is used to obtain the longest path of the basic block; and to determine the difference between the longest path and the first heuristic as the second heuristic.
[0156] The fourth determining submodule is used to determine the proportion of the number of scheduled parent or child instructions of the current instruction to all parent and child instructions of the current instruction as a first weighting coefficient, and to determine the proportion of the number of unscheduled parent or child instructions to all parent and child instructions of the current instruction as a second weighting coefficient.
[0157] The fifth determining submodule is used to determine the height and depth of the scheduled parent or child instruction of the current instruction and the proportion of them to the longest path as a third weighting coefficient, and to determine the height and depth of the unscheduled parent or child instruction of the current instruction and the proportion of them to the longest path as a fourth weighting coefficient.
[0158] The adjustment submodule is used to adjust the first heuristic and the second heuristic based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient to obtain the target heuristic of the current instruction.
[0159] In one embodiment, the first determining submodule is further configured to identify the child nodes of the node where the current instruction is located as key nodes of the scheduled instructions;
[0160] Identify the child nodes of the node containing the scheduled instruction;
[0161] The node with the maximum height or maximum depth of the intersection node between the child node of the node containing the scheduled instruction and the critical node is determined as the point of overlap of the current instruction in the critical path of the currently scheduled instruction.
[0162] In one embodiment, the adjustment submodule is further configured to, if the current traversal is top-down, adjust the first heuristic and the second heuristic according to the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient to obtain the target heuristic of the current instruction; if the current traversal is top-down, adjust the first heuristic and the second heuristic according to the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient to obtain the target heuristic of the current instruction.
[0163] In one embodiment, the device further includes:
[0164] The adjustment module is used to adjust the instructions of the scheduled basic block in one of the following ways: adjusting the target storage instructions of the basic block that have a large distance from the parent instruction; adjusting the target loading instructions of the basic block that have a large distance from the parent and child instructions; adjusting the distance between instructions without a parent and their child instructions; adjusting the gaps between instructions due to weight; adjusting the position of the storage instructions within the basic package.
[0165] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0166] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0167] Embodiments of this application also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0168] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0169] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0170] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0171] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. An instruction scheduling and processing method, characterized in that, The method includes: The scheduled instructions of the current basic block to be scheduled are constructed into a directed acyclic graph according to the data dependency relationship. The directed acyclic graph consists of multiple nodes, each node corresponds to a scheduled instruction, and the connection between nodes is the clock cycle that the node needs to wait. Obtain multiple instructions for the current schedule from both top-down and bottom-up perspectives; Determining target heuristics for the plurality of instructions based on scheduled and unscheduled instructions in the directed acyclic graph includes: determining the overlap point of the current instruction in the critical path of the currently scheduled instructions, wherein the current instruction is each of the plurality of instructions; determining the difference between the height of the current instruction and the overlap point as a first heuristic, or determining the difference between the depth of the current instruction and the overlap point as a first heuristic; obtaining the longest path of the basic block; determining the difference between the longest path and the first heuristic as a second heuristic; and determining the number of scheduled parent or child instructions of the current instruction relative to the total number of parent and child instructions of the current instruction. The weight of the first heuristic is determined as the first weight coefficient, and the weight of the number of unscheduled parent or child instructions relative to the total number of parent and child instructions of the current instruction is determined as the second weight coefficient; the weight of the height and depth of the scheduled parent or child instructions of the current instruction relative to the longest path is determined as the third weight coefficient, and the weight of the height and depth of the unscheduled parent or child instructions of the current instruction relative to the longest path is determined as the fourth weight coefficient; the target heuristic of the current instruction is obtained by adjusting the first heuristic and the second heuristic based on the first weight coefficient, the second weight coefficient, the third weight coefficient, and the fourth weight coefficient; The multiple instructions are scheduled according to the target heuristic.
2. The method according to claim 1, characterized in that, Scheduling the plurality of instructions based on the heuristic includes: During the current clock cycle, the plurality of instructions are scheduled by repeatedly executing the following steps: selecting the target scheduling instruction from the plurality of instructions based on the heuristic; if the target scheduling instruction is a top-down obtained instruction, saving the target scheduling instruction to a first queue; if the target scheduling instruction is a bottom-up obtained instruction, saving the target scheduling instruction to a second queue. If the target instruction among the plurality of instructions is not scheduled in the current clock cycle, the target instruction shall be scheduled in the next clock cycle, wherein the target instruction is one or more instructions.
3. The method according to claim 2, characterized in that, Selecting the target scheduling instruction for the current scheduling from the plurality of instructions based on the heuristic includes: If the heuristic value of the current instruction is equal to the heuristic value of the pre-selected instruction and the current instruction is a boundary instruction, then if the current traversal is top-down, the target scheduling instruction is determined based on the number of child nodes of the multiple instructions; if the current traversal is bottom-up, the target scheduling instruction is determined based on the number of parent nodes of the multiple instructions. If the heuristic value of the current instruction equals the heuristic value of the pre-selected instruction, and both the current instruction and the pre-selected instruction's parent or child instructions have been scheduled, then if the unscheduled sequence is not empty, the instruction with the most instructions in the unscheduled sequence is determined as the target scheduling instruction; if the scheduled sequence is not empty, the instruction with the most instructions in the scheduled sequence is determined as the target scheduling instruction; if both the unscheduled sequence and the scheduled sequence are empty, the target scheduling instruction is determined based on the height or depth of the multiple instructions.
4. The method according to claim 3, characterized in that, If the current traversal is top-down, the target scheduling instruction is determined based on the number of nodes of the multiple instructions; if the current traversal is bottom-up, the target scheduling instruction is determined based on the number of parent nodes of the multiple instructions, including: If the current traversal is top-down, and the instruction with the most child nodes is a single instruction, then the instruction with the most child nodes is determined to be the target scheduling instruction; if the instruction with the most child nodes is at least two instructions, then the index values of the at least two instructions are compared, and the instruction with the smallest index value is determined to be the target scheduling instruction. If the current traversal is bottom-up, and the instruction with the most parent nodes is a single instruction, then the instruction with the most parent nodes is determined to be the target scheduling instruction; if the instruction with the most parent nodes is at least two instructions, then the index values of the at least two instructions are compared, and the instruction with the largest index value is determined to be the target scheduling instruction.
5. The method according to claim 3, characterized in that, If both the unscheduled sequence and the scheduled sequence are empty, determining the target scheduling instruction based on the height or depth of the plurality of instructions includes: If the current traversal is top-down, and if the instruction with the largest height is a single instruction, then the instruction with the largest height is determined to be the target scheduling instruction; if the instruction with the largest height is at least two instructions, then the instruction with the smaller depth among the at least two instructions is determined to be the target scheduling instruction. If the traversal is bottom-up, and the instruction with the greatest depth is a single instruction, then the instruction with the greatest depth is determined as the target scheduling instruction; if the instruction with the greatest depth is at least two instructions, then the instruction with the smallest height among the at least two instructions is determined as the target scheduling instruction.
6. The method according to claim 1, characterized in that, Determining the overlap points of the current instruction within the critical path of currently scheduled instructions includes: The child nodes of the node where the current instruction is located are designated as key nodes of the scheduled instructions; Identify the child nodes of the node containing the scheduled instruction; The node with the maximum height or maximum depth of the intersection node between the child node of the node containing the scheduled instruction and the critical node is determined as the point of overlap of the current instruction in the critical path of the currently scheduled instruction.
7. The method according to claim 6, characterized in that, The target heuristic for the current instruction is obtained by adjusting the first heuristic and the second heuristic based on the first weighting coefficient, the second weighting coefficient, the third weighting coefficient, and the fourth weighting coefficient, including: If the current traversal is top-down, the first heuristic and the second heuristic are adjusted according to the first weight coefficient, the second weight coefficient, the third weight coefficient, the fourth weight coefficient, and the maximum height to obtain the target heuristic of the current instruction; If the current traversal is from low to high, the first heuristic and the second heuristic are adjusted according to the first weight coefficient, the second weight coefficient, the third weight coefficient, the fourth weight coefficient, and the maximum depth to obtain the target heuristic of the current instruction.
8. The method according to claim 1, characterized in that, The method further includes: Adjust the instructions of the scheduled basic block using one of the following methods: Adjust the target storage instruction in the basic block whose distance from the parent instruction is larger; Adjust the target loading instructions in the basic block that have a large distance from the parent and child class instructions; Adjust the distance between instructions without a parent class and their child classes; Adjust the gaps between instructions caused by weighting; Adjust the location of storage instructions within the basic package.
9. An instruction scheduling and processing device, characterized in that, The device includes: The construction module is used to construct a directed acyclic graph (DAG) of the scheduled instructions of the currently scheduled basic blocks according to the data dependency relationship. The DAG consists of multiple nodes, each node corresponds to a scheduled instruction, and the connection between nodes is the clock cycle that the node needs to wait for. The acquisition module is used to acquire multiple currently scheduled instructions in a top-down and bottom-up manner. The determination module is used to determine the target heuristics of the plurality of instructions based on the scheduled instructions and unscheduled instructions in the directed acyclic graph, including: determining the overlap point of the current instruction in the critical path of the currently scheduled instructions, wherein the current instruction is each of the plurality of instructions; determining the difference between the height of the current instruction and the overlap point as a first heuristic, or determining the difference between the depth of the current instruction and the overlap point as a first heuristic; obtaining the longest path of the basic block; determining the difference between the longest path and the first heuristic as a second heuristic; and determining the number of scheduled parent or child instructions of the current instruction relative to the total number of scheduled parent and child instructions of the current instruction. The weight of instruction class is determined as a first weighting coefficient, and the weight of the number of unscheduled parent or child instructions relative to the total number of parent and child instructions of the current instruction is determined as a second weighting coefficient; the weight of the height and depth of the scheduled parent or child instructions of the current instruction relative to the longest path is determined as a third weighting coefficient, and the weight of the height and depth of the unscheduled parent or child instructions of the current instruction relative to the longest path is determined as a fourth weighting coefficient; the target heuristic of the current instruction is obtained by adjusting the first heuristic and the second heuristic based on the first, second, third, and fourth weighting coefficients. The scheduling module is used to schedule the plurality of instructions according to the target heuristic.
10. A computer-readable storage medium storing a computer program, wherein, The computer program is configured to execute the method described in any one of claims 1 to 8 when it is run.
11. An electronic device comprising a memory and a processor, the memory storing a computer program, the processor being configured to run the computer program to perform the method of any one of claims 1 to 8.