Instruction Scheduling Method, Apparatus, Device, and Storage Medium
By determining and scheduling the duration of memory access instructions in the microarchitecture model, the instruction scheduling performance problems caused by the difference in memory access instructions execution time in the prior art are solved, and more efficient instruction scheduling is achieved.
Patent Information
- Application Number
- CN202111506086.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-10
AI Technical Summary
In the prior art, when performing instruction scheduling, the execution time of some instructions, especially the time consumed by memory access instruction access data, is different from the actual time, resulting in poor instruction scheduling performance.
By determining the target memory access instructions in the instruction set corresponding to the microarchitecture model, and determining the duration of the memory access instructions in different instruction running scenarios, the memory access instructions are scheduled based on these durations.
Improves instruction scheduling performance, high applicability, and can more accurately model and optimize the time consumed by memory access instructions.
Smart Images

Figure CN114201281B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of microarchitecture, and in particular, to an instruction scheduling method, apparatus, device, and storage medium. Background Art
[0002] In the prior art, a compiler usually builds a microarchitecture model based on information such as a hardware pipeline, an instruction execution duration, a bypass function, etc., and schedules each instruction corresponding to the microarchitecture model. However, when performing instruction scheduling in the prior art, the execution duration of some instructions, especially the duration consumed by a memory access instruction to access data, often differs from the actual duration, thereby resulting in poor instruction scheduling performance. Summary of the Invention
[0003] Embodiments of this application provide an instruction scheduling method, apparatus, device, and storage medium, which can schedule memory access instructions based on the duration consumed by the memory access instructions in different instruction running scenarios, and have high applicability.
[0004] In a first aspect, an embodiment of this application provides an instruction scheduling method, which includes:
[0005] Determine at least one target memory access instruction in an instruction set corresponding to a microarchitecture model;
[0006] Determine the duration consumed by each of the above target memory access instructions in multiple instruction running scenarios;
[0007] Based on the duration consumed by each of the above target memory access instructions in multiple above instruction running scenarios, perform instruction scheduling on each of the above target memory access instructions.
[0008] In a second aspect, an embodiment of this application provides an instruction scheduling apparatus, which includes:
[0009] An instruction determination module, configured to determine at least one target memory access instruction in an instruction set corresponding to a microarchitecture model;
[0010] A duration determination module, configured to determine the duration consumed by each of the above target memory access instructions in multiple instruction running scenarios;
[0011] An instruction scheduling module, configured to perform instruction scheduling on each of the above target memory access instructions based on the duration consumed by each of the above target memory access instructions in multiple above instruction running scenarios.
[0012] In a third aspect, an embodiment of this application provides an electronic device, including a processor and a memory, and the processor and the memory are connected to each other;
[0013] The above memory is used to store a computer program;
[0014] The above-mentioned processor is configured to execute the instruction scheduling method provided in the embodiments of the present application when calling the above-mentioned computer program.
[0015] In a fourth aspect, embodiments of the present application provide a computer-readable storage medium storing a computer program, which is executed by a processor to implement the instruction scheduling method provided in the embodiments of the present application.
[0016] In the embodiments of the present application, by determining the duration consumed by a memory access instruction in different instruction execution scenarios to schedule the memory access instruction, the instruction scheduling performance can be improved, and the applicability is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 is a flowchart of the instruction scheduling method provided in the embodiments of the present application;
[0019] Figure 2 is a schematic structural diagram of the instruction scheduling device provided in the embodiments of the present application;
[0020] Figure 3 is a schematic structural diagram of the electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0022] See Figure 1 , Figure 1 is a flowchart of the instruction scheduling method provided in the embodiments of the present application. As Figure 1 shown, the instruction scheduling method provided in the embodiments of the present application may include the following steps:
[0023] Step S11: Determine at least one target memory access instruction in the instruction set corresponding to the microarchitecture model.
[0024] In some feasible embodiments, the instruction set corresponding to the microarchitecture model is a set of commands indicating the hardware to perform certain arithmetic and processing functions.
[0025] Among them, the target memory access instruction is a memory access instruction in the instruction set whose execution time for corresponding operations is not fixed.
[0026] For example, for any memory access instruction in the instruction set, if the execution time of the memory access instruction is inconsistent in different instruction running scenarios, it can be determined that the memory access instruction is the target memory access instruction in the instruction set.
[0027] Among them, for a memory access instruction, the instruction running scenarios corresponding to the memory access instruction include but are not limited to accessing data in the case of cache miss or accessing data in the case of cache hit, which can be specifically determined based on the requirements of the actual application scenario and are not limited here.
[0028] Specifically, when determining at least one target memory access instruction from the instruction set corresponding to the microarchitecture model, the execution time of each memory access instruction in the instruction set in different instruction running scenarios can be determined by means of hardware simulation or actual testing based on the microarchitecture model, and then the memory access instructions with inconsistent execution times in different instruction running scenarios are determined as the target memory access instructions in the instruction set.
[0029] Optionally, when determining at least one target memory access instruction from the instruction set corresponding to the microarchitecture model, statement judgment can be performed on each instruction in the instruction set to determine the loop statements therein (for convenience of description, hereinafter referred to as the first instructions).
[0030] After determining the first instructions from the instruction set, at least one target memory access instruction can be determined from the first instructions. That is, the loop statements are determined from the instruction set corresponding to the microarchitecture model, and at least one target memory access instruction is determined from the loop statements.
[0031] Furthermore, for each determined first instruction, the first instruction can be loop-unrolled, and the first instruction is unrolled from a loop statement into multiple linear instructions. That is, the loop instruction is unrolled into multiple independent linear instructions.
[0032] Furthermore, all memory access instructions are determined from all the linear instructions corresponding to each first instruction, and the determined memory access instructions are determined as the target memory access instructions in the instruction set. That is, all the loop statements in the instruction set are unrolled to obtain multiple independent linear instructions, and the memory access instructions in each linear instruction are determined as the memory access instructions in the instruction set.
[0033] Optionally, among all the linear instructions corresponding to the first instruction in the instruction set, there may be memory access instructions that only correspond to one instruction execution scenario, such as memory access instructions for accessing data from a fixed storage space. The time consumed for such memory access instructions to access data is a fixed time.
[0034] Based on this, after determining all the linear instructions corresponding to each first instruction, it is also possible to determine, through hardware simulation or actual testing based on a microarchitecture model, memory access instructions with inconsistent time consumption in different instruction execution scenarios from all the linear instructions, and determine the memory access instructions with inconsistent time consumption in different instruction execution scenarios as the target memory access instructions in the instruction set.
[0035] Optionally, after determining the linear instructions corresponding to all the first instructions in the instruction set, it is possible to first determine the instructions that consume a fixed time for executing the corresponding operations among all the linear instructions (hereinafter referred to as the second instructions for convenience of description), and then determine at least one target memory access instruction from the other linear instructions except the second instructions.
[0036] Among them, the second instructions among all the linear instructions can be determined through hardware simulation or actual testing based on a microarchitecture model. Since the time consumed for the other linear instructions except the second instructions among all the linear instructions to execute the corresponding operations is not fixed, it is possible to determine memory access instructions from the other linear instructions except the second instructions, and determine the memory access instructions except the second instructions as the target memory access instructions in the instruction set.
[0037] In some feasible implementation manners, after determining the target memory access instructions from the loop statements (the first instructions) in the instruction set, the non-loop statements in the instruction set may also include memory access instructions with inconsistent time consumption in different instruction execution scenarios.
[0038] Based on this, after determining the target memory access instructions in the loop statements in the instruction set, it is possible to determine the other memory access instructions in the instruction set except the first instructions (loop statements). Further, determine the target memory access instructions with inconsistent time consumption in different instruction execution scenarios from the other memory access instructions, such as determining the target memory access instructions with inconsistent time consumption in different instruction execution scenarios in the non-loop statements in the instruction set through hardware simulation or actual testing based on a microarchitecture model.
[0039] Optionally, the memory access instructions in the non-loop statements can also be directly determined as the target memory access instructions in the instruction set.
[0040] Step S12: Determine the time consumed by each target memory access instruction in multiple instruction execution scenarios.
[0041] In some feasible embodiments, after determining the target memory access instructions in the instruction set, the duration consumed by each target memory access instruction in different instruction execution scenarios can be determined. For example, for each target memory access instruction, the duration consumed for accessing data in the case of cache miss and the duration consumed for accessing data in the case of cache hit can be determined.
[0042] Among them, if all the memory access instructions in the instruction set are determined as target memory access instructions, then after determining the target memory access instructions, the duration consumed by each memory access instruction in different instruction execution scenarios can be determined by means of hardware emulation or actual testing based on a microarchitecture model.
[0043] Among them, if when determining the target memory access instructions, first the first instruction (loop statement) in the instruction set is loop-unrolled, and then the memory access instructions in the linear instructions obtained after loop-unrolling the first instruction are determined as target memory access instructions, then after determining the target memory access instructions, the duration consumed by each memory access instruction in different instruction execution scenarios can be determined by means of hardware emulation or actual testing based on a microarchitecture model.
[0044] Among them, if first the duration consumed by the memory access instructions in the instruction set in different instruction execution scenarios is determined by means of hardware emulation or actual testing based on a microarchitecture model, and then the memory access instructions with inconsistent durations consumed in different instruction execution scenarios are determined as the target memory access instructions in the instruction set, then the duration consumed by the target memory access instructions in different instruction execution scenarios can be directly obtained.
[0045] Step S13: Perform instruction scheduling on each target memory access instruction based on the duration consumed by each target memory access instruction in multiple instruction execution scenarios.
[0046] In some feasible embodiments, after determining the duration consumed by the target memory access instructions in the instruction set in multiple instruction execution scenarios, instruction scheduling can be performed on each target memory access instruction based on the duration consumed by each target memory access instruction in multiple instruction execution scenarios.
[0047] Specifically, for each target memory access instruction, the scheduling priority of the target memory access instruction in different instruction execution scenarios can be determined based on the duration consumed by the target memory access instruction in multiple instruction execution scenarios.
[0048] If it can be determined that the scheduling priority of accessing data by the target memory access instruction in the case of cache hit is higher than that in the case of cache miss, then when scheduling the target memory access instruction, the target memory access instruction is preferentially scheduled in the case of cache hit.
[0049] Optionally, based on the hardware pipeline, bypass information, etc. of the microarchitecture model and the time consumed by each target memory access instruction in different instruction execution scenarios, etc., the execution order of each target memory access instruction can be reordered so that more target memory access instructions can be executed in less time.
[0050] Optionally, the loop statements in the instruction set can be replaced with linear instructions, and based on the hardware pipeline, bypass information, etc. of the microarchitecture model and the time consumed by each target memory access instruction in different instruction execution scenarios, etc., the target memory access instructions in the linear instructions and / or the target memory access instructions in other statements in the instruction set except the loop statements are rescheduled, so as to run a larger number of memory access instructions in less time while reducing the time error caused by the different time consumed by each target memory access instruction in different instruction execution scenarios during the statement loop process, and improving the instruction scheduling performance.
[0051] In some feasible implementation manners, after determining the target memory access instructions in the instruction set based on any of the above feasible implementation manners, it is also possible to determine the time consumed by other instructions (hereinafter referred to as the third instructions) in the instruction set except the target memory access instructions when performing corresponding operations.
[0052] Among them, the third instructions in the instruction set may include instructions with a fixed time consumption when performing corresponding operations, or may include instructions with inconsistent time consumption in different instruction execution scenarios, which can be specifically determined based on the actual application scenario requirements and are not limited here.
[0053] That is to say, for each instruction in the instruction set including the target memory access instruction, if there are multiple instruction execution scenarios for the instruction, then determine the time consumed by the instruction in each instruction execution scenario. If there is only one instruction execution scenario for the instruction, that is, the time consumed by the instruction when performing the corresponding operation is a fixed time, then only the corresponding fixed time of the instruction needs to be determined.
[0054] Furthermore, based on the time consumed by each target memory access instruction in the first instruction set in each instruction execution scenario and the time consumed by each third instruction when performing the corresponding operation, instruction scheduling is performed on each of the target memory access instructions and each of the third instructions. That is, based on the time consumed by each instruction in the instruction set in each instruction execution scenario, all the instructions in the instruction set are scheduled.
[0055] Specifically, for the third instruction and the target memory access instructions, the scheduling priorities of the third instruction and the target memory access instructions can be determined based on the time consumed by the third instruction when performing the corresponding operation and the time consumed by each target memory access instruction in each instruction running scenario.
[0056] For example, for any target memory access instruction, it can be determined that the scheduling priority of accessing data in the case of cache hit is higher than that of a certain third instruction, and the scheduling priority of accessing data in the case of cache miss is lower than that of the third instruction. Therefore, when scheduling the target memory access instruction and the third instruction, the target memory access instruction is preferentially scheduled in the case of cache hit.
[0057] Optionally, based on information such as the hardware pipeline and bypass information of the microarchitecture model, the time consumed by each target memory access instruction in different instruction running scenarios, and the time consumed by each third instruction when performing the corresponding operation, etc., the running order of each target memory access instruction and each third instruction can be reordered to run more instructions in less time.
[0058] Optionally, the loop statements in the instruction set can be replaced with linear instructions, and based on information such as the hardware pipeline and bypass information of the microarchitecture model, the time consumed by each third instruction when performing the corresponding operation, and the time consumed by each target memory access instruction in different instruction running scenarios, etc., the target memory access instructions and the third instructions in the linear instructions can be rescheduled to run a larger number of instructions in less time while improving the instruction scheduling performance.
[0059] For example, for the microarchitecture model, based on the instruction set corresponding to the microarchitecture model, the time consumed by each target memory access instruction in different instruction running scenarios, and the time consumed by each third instruction in the instruction set when performing the corresponding operation, the execution order of the instructions corresponding to the microarchitecture model can be adjusted. For example, the target memory access instructions and / or the third instructions with shorter consumed time can be concentratedly arranged without affecting the performance of the microarchitecture model, so that the microarchitecture model can run more instructions per unit time.
[0060] In some feasible implementation manners, the data stored in the cache corresponding to the microarchitecture model can also be determined in real time, and the first memory access instruction corresponding to the data in the cache can be determined. That is, the first memory access instruction can access data in the case of cache hit.
[0061] Further, the determined target memory access instructions in the instruction set are determined as the first memory access instruction and the second memory access instruction. Based on the time consumed for accessing data in the case of cache hit by the first memory access instruction, the time consumed for accessing data in the case of cache miss by the second memory access instruction, and the time consumed for the third instruction to execute the corresponding operation, instruction scheduling is performed on each target memory access instruction and the third instruction, so that the scheduling of each target memory access instruction and the third instruction is more in line with the actual operation requirements of the microarchitecture model, and the applicability is high.
[0062] In the embodiment of the present application, by determining the time consumed by the target memory access instructions in the instruction set in each instruction operation scenario, when scheduling the instructions in the instruction set, the consumption time of the target memory access instructions can be more accurately modeled, achieving a better instruction scheduling effect.
[0063] See Figure 2 , Figure 2 which is a schematic structural diagram of the instruction scheduling device provided by the embodiment of the present application. The instruction scheduling device provided by the embodiment of the present application includes:
[0064] An instruction determination module 21, configured to determine at least one target memory access instruction in the instruction set corresponding to the microarchitecture model;
[0065] A duration determination module 22, configured to determine the time consumed by each of the above target memory access instructions in multiple instruction operation scenarios;
[0066] An instruction scheduling module 23, configured to perform instruction scheduling on each of the above target memory access instructions based on the time consumed by each of the above target memory access instructions in multiple above instruction operation scenarios.
[0067] In some feasible implementation manners, the above instruction operation scenarios include accessing data in the case of cache miss and accessing data in the case of cache hit.
[0068] In some feasible implementation manners, when determining at least one target memory access instruction in the instruction set corresponding to the microarchitecture model, the above instruction determination module 21 is configured to:
[0069] Perform statement judgment on each instruction in the instruction set corresponding to the microarchitecture model to determine the first instruction in the above instruction set, and the first instruction is a loop statement;
[0070] Determine at least one target memory access instruction from the above first instruction.
[0071] In some feasible implementation manners, when determining at least one target memory access instruction from the above first instruction, the above instruction determination module 21 is configured to:
[0072] The above first instruction is loop-unrolled to obtain a plurality of linear instructions corresponding to the above first instruction;
[0073] At least one target memory access instruction is determined from each of the above linear instructions.
[0074] In some feasible embodiments, when determining at least one target memory access instruction from each of the above linear instructions, the above instruction determination module 21 is configured to:
[0075] Determine a second instruction in each of the above linear instructions that consumes a fixed duration for performing the corresponding operation;
[0076] At least one target memory access instruction is determined from the other linear instructions except the above second instruction.
[0077] In some feasible embodiments, the above instruction determination module 21 is further configured to:
[0078] Determine at least one target memory access instruction from the other instructions in the above instruction set except the above first instruction.
[0079] In some feasible embodiments, the above instruction scheduling module 23 is further configured to:
[0080] Determine the duration consumed by a third instruction in the above instruction set for performing the corresponding operation except each of the above target memory access instructions;
[0081] Based on the duration consumed by each of the above target memory access instructions in each of the above instruction running scenarios and the duration consumed by each of the above third instructions for performing the corresponding operation, instruction scheduling is performed on each of the above target memory access instructions and each of the above third instructions.
[0082] In a specific implementation, the above instruction scheduling device may execute the implementation manners provided in the above steps through its built-in respective functional modules. Specifically, reference may be made to the implementation manners provided in the above respective steps, which will not be elaborated herein. Figure 1 In a specific implementation, the above instruction scheduling device may execute the implementation manners provided in the above steps through its built-in respective functional modules. Specifically, reference may be made to the implementation manners provided in the above respective steps, which will not be elaborated herein.
[0083] See Figure 3 , Figure 3 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 3As shown in the figure, the electronic device 300 in this embodiment may include: a processor 301, a network interface 304, and a memory 305. In addition, the above-mentioned electronic device 300 may further include: a user interface 303 and at least one communication bus 302. Among them, the communication bus 302 is used to realize the connection and communication between these components. Among them, the user interface 303 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface. The network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 305 may be a high-speed RAM memory or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As Figure 3 shown, the memory 305, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0084] In Figure 3 the electronic device 300 shown in the figure, the network interface 304 can provide network communication functions; while the user interface 303 is mainly used to provide an input interface for users; and the processor 301 can be used to call the device control application program stored in the memory 305 to achieve:
[0085] Determine at least one target memory access instruction in the instruction set corresponding to the microarchitecture model;
[0086] Determine the duration consumed by each of the above target memory access instructions in multiple instruction execution scenarios;
[0087] Based on the duration consumed by each of the above target memory access instructions in multiple above-mentioned instruction execution scenarios, perform instruction scheduling on each of the above target memory access instructions.
[0088] In some feasible embodiments, the above-mentioned instruction execution scenarios include accessing data in the case of cache misses and accessing data in the case of cache hits.
[0089] In some feasible embodiments, when determining at least one target memory access instruction in the instruction set corresponding to the microarchitecture model, the above-mentioned processor 301 is used to:
[0090] Perform statement judgment on each instruction in the instruction set corresponding to the microarchitecture model to determine the first instruction in the above-mentioned instruction set, and the first instruction is a loop statement;
[0091] Determine at least one target memory access instruction from the above first instruction.
[0092] In some feasible embodiments, when determining at least one target memory access instruction from the above first instruction, the processor 301 is configured to:
[0093] Perform loop unrolling on the above first instruction to obtain a plurality of linear instructions corresponding to the above first instruction;
[0094] Determine at least one target memory access instruction from each of the above linear instructions.
[0095] In some feasible embodiments, when determining at least one target memory access instruction from each of the above linear instructions, the processor 301 is configured to:
[0096] Determine a second instruction in each of the above linear instructions that consumes a fixed duration for performing the corresponding operation;
[0097] Determine at least one target memory access instruction from other linear instructions except the above second instruction.
[0098] In some feasible embodiments, the processor 301 is further configured to:
[0099] Determine at least one target memory access instruction from other instructions in the above instruction set except the above first instruction.
[0100] In some feasible embodiments, the processor 301 is further configured to:
[0101] Determine the duration consumed by a third instruction in the above instruction set except each of the above target memory access instructions when performing the corresponding operation;
[0102] Based on the duration consumed by each of the above target memory access instructions in each of the above instruction running scenarios and the duration consumed by each of the above third instructions when performing the corresponding operation, perform instruction scheduling on each of the above target memory access instructions and each of the above third instructions.
[0103] It should be understood that in some possible embodiments, the above-mentioned processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0104] In specific implementation, the above-mentioned electronic device 300 may execute the implementation manners provided in each of the above Figure 1 steps through its built-in functional modules. For specific details, refer to the implementation manners provided in each of the above steps, which will not be elaborated here.
[0105] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program that is executed by a processor to implement Figure 1 the methods provided in each of the above steps. For specific details, refer to the implementation manners provided in each of the above steps, which will not be elaborated here.
[0106] The above computer-readable storage medium may be an internal storage unit of the instruction scheduling device or the electronic device provided in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. The above computer-readable storage medium may further include magnetic disks, optical disks, read-only memory (ROM), or random access memory (RAM), etc. Further, the computer-readable storage medium may include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.
[0107] An embodiment of the present application provides a computer program product, which includes a computer program or computer instructions. When the above computer program or computer instructions are executed by a processor, the voice playback method provided by the embodiment of the present application is executed. Figure 1 The methods provided by each step in the above are executed.
[0108] Terms such as "first" and "second" in the claims, the description and the drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or electronic device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or electronic devices. The mention of "embodiment" in this document means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The display of this phrase at various positions in the description does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the description and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0109] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0110] The above-disclosed are only the preferred embodiments of the present application, and thus cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. An instruction scheduling method, characterized in that, the method includes: Performing statement judgment on each instruction in the instruction set corresponding to the microarchitecture model to determine the first instruction in the instruction set, where the first instruction is a loop statement; Performing loop unrolling on the first instruction to obtain multiple linear instructions corresponding to the first instruction; determining a second instruction that consumes a fixed duration for performing a corresponding operation among each of the linear instructions; determining at least one target memory access instruction from other linear instructions except the second instruction; Determining the duration consumed by each of the target memory access instructions in multiple instruction execution scenarios; wherein, the multiple instruction execution scenarios at least include accessing data in case of cache miss and accessing data in case of cache hit; Based on the duration consumed by each of the target memory access instructions in multiple instruction execution scenarios, performing instruction scheduling on each of the target memory access instructions.
2. The method according to claim 1, characterized in that, the instruction execution scenarios include accessing data in case of cache miss and accessing data in case of cache hit.
3. The method according to claim 1, characterized in that, the method further includes: Determining at least one target memory access instruction from other instructions in the instruction set except the first instruction.
4. The method according to claim 1, characterized in that, the method further includes: Determining the duration consumed by a third instruction in the instruction set except each of the target memory access instructions when performing a corresponding operation; Based on the duration consumed by each of the target memory access instructions in each of the instruction execution scenarios and the duration consumed by each of the third instructions when performing a corresponding operation, performing instruction scheduling on each of the target memory access instructions and each of the third instructions.
5. An instruction scheduling device, characterized in that, the device includes: An instruction determination module, configured to perform statement judgment on each instruction in the instruction set corresponding to the microarchitecture model to determine the first instruction in the instruction set, where the first instruction is a loop statement; The instruction determination module is configured to perform loop unrolling on the first instruction to obtain multiple linear instructions corresponding to the first instruction; determine a second instruction that consumes a fixed duration for performing a corresponding operation among each of the linear instructions; determine at least one target memory access instruction from other linear instructions except the second instruction; A duration determination module, configured to determine the duration consumed by each of the target memory access instructions in multiple instruction execution scenarios; wherein, the multiple instruction execution scenarios at least include accessing data in case of cache miss and accessing data in case of cache hit; An instruction scheduling module, configured to perform instruction scheduling on each of the target memory access instructions based on the duration consumed by each of the target memory access instructions in multiple instruction execution scenarios.
6. An electronic device, characterized in that, including a processor and a memory, the processor and the memory are connected to each other; The memory is used to store a computer program; The processor is configured to execute the method according to any one of claims 1 to 4 when calling the computer program.
7. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method for measuring latencies by randomly selected sampling of the instructions while the instructions are executed
US6092180A