Instruction rearrangement method, electronic equipment and computer storage medium

By performing dependency analysis and determining multiple reordering orders in the instructions to be executed by the processor, the problem of poor efficiency and effectiveness of manual reordering instructions in the prior art is solved, and more efficient instruction reordering is achieved.

CN119938147APending Publication Date: 2025-05-06T-HEAD (SHANGHAI) SEMICON CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411790830.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, instruction rearrangement is mostly done manually, resulting in poor efficiency and effect, and is time-consuming and easy to make mistakes.

Method used

By obtaining the instruction fragments to be rearranged in the instructions to be executed by the processor, performing dependency analysis, determining multiple reorder orders, and running these orders through the processor to obtain performance data, and finally determining the target reorder order.

Benefits of technology

It improves the efficiency of the processor running instructions, improves the efficiency and effect of instruction reordering, and avoids the time-consuming and easy-to-missing problems of manual reordering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938147A_ABST
    Figure CN119938147A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an instruction rearrangement method, electronic equipment and a computer storage medium. The instruction rearrangement method comprises the steps that an instruction fragment to be rearranged is obtained from an instruction to be executed by a processor, and the instruction fragment comprises at least one loop body; performing dependency relationship analysis on the instructions in the instruction fragments, determining various rearrangement sequences of the instruction fragments, and keeping the sequence of the instructions with the dependency relationship unchanged; respectively running the instruction fragments according to the multiple rearrangement sequences through the processor, and obtaining performance data respectively corresponding to the multiple rearrangement sequences; and according to the performance data corresponding to the multiple rearrangement sequences, determining a target rearrangement sequence in the multiple rearrangement sequences. According to the instruction rearrangement method and the instruction rearrangement device, the instruction fragment comprising at least one loop body is subjected to dependency relationship analysis to perform instruction rearrangement, the rearranged instruction is actually operated through the processor, the target rearrangement sequence with better performance is determined, and the efficiency and the effect of instruction rearrangement are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to an instruction rearrangement method, an electronic device, and a computer storage medium. Background Art

[0002] In the process of computer processor executing instructions, compiler is a general software used to translate one language (usually high-level language) into another language (usually low-level language). Among them, high-level language is convenient for people to write, read, communicate and maintain. Low-level language is a computer language that the processor can directly interpret and run. The instruction set architecture defines how the processor executes and operates instructions, including basic data types, instruction sets, registers, addressing modes, storage systems, interrupts, exception handling and other aspects.

[0003] Before executing instructions, the processor can rearrange the order of instructions to improve the efficiency of the processor in executing instructions. However, currently instruction rearrangement is mostly done manually. Although manual instruction rearrangement can improve performance, the rearrangement process is time-consuming and error-prone, resulting in poor efficiency and effect of instruction rearrangement. Summary of the invention

[0004] In view of this, embodiments of the present application provide an instruction re-arrangement method, an electronic device, and a computer storage medium to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of an embodiment of the present application, there is provided an instruction re-arrangement method, comprising: obtaining instruction fragments to be re-arranged from instructions to be executed by a processor, wherein the instruction fragments include at least one loop body; performing dependency analysis on the instructions in the instruction fragments, and determining a plurality of re-arrangement orders of the instruction fragments according to the dependencies between the instructions in the instruction fragments; wherein the order of the instructions with dependencies in the instruction fragments and the re-arrangement order remains unchanged; executing the instruction fragments respectively according to the plurality of re-arrangement orders of the instruction fragments by the processor, and obtaining performance data corresponding to the plurality of re-arrangement orders of the instruction fragments; and determining a target re-arrangement order among the plurality of re-arrangement orders according to the performance data corresponding to the plurality of re-arrangement orders of the instruction fragments.

[0006] According to a second aspect of an embodiment of the present application, an instruction re-arrangement device is provided, comprising: an acquisition module, used to acquire instruction fragments to be re-arranged in instructions to be executed by a processor, wherein the instruction fragments include at least one loop body; a re-arrangement module, used to perform dependency analysis on the instructions in the instruction fragments, and determine multiple re-arrangement orders of the instruction fragments according to the dependencies between the instructions in the instruction fragments; wherein the order of instructions with dependencies in the instruction fragments and the re-arrangement order remains unchanged; a performance module, used to respectively run the instruction fragments according to the multiple re-arrangement orders of the instruction fragments through the processor, and obtain performance data corresponding to the multiple re-arrangement orders of the instruction fragments; and a determination module, used to determine a target re-arrangement order among the multiple re-arrangement orders according to the performance data corresponding to the multiple re-arrangement orders of the instruction fragments.

[0007] According to the third aspect of an embodiment of the present application, there is provided an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.

[0008] According to a fourth aspect of an embodiment of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0009] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer instructions, wherein the computer instructions instruct a computing device to perform operations corresponding to the method described in the first aspect.

[0010] According to an instruction reordering method provided by an embodiment of the present application, an instruction fragment to be reordered including at least one loop body is obtained from the instructions to be executed by the processor, and the instructions in the instruction fragment are subjected to dependency analysis and reordered to obtain a plurality of reordering orders of the instruction fragment. Since the loop body is executed a large number of times, reordering the instructions in the loop body can greatly improve the operating efficiency of the processor. Moreover, the order between the instructions with dependencies remains unchanged, thereby ensuring the correctness of the execution of the instructions. By using the processor to respectively run the instruction fragments according to a plurality of reordering orders of the instruction fragments, and obtaining the performance data corresponding to the plurality of reordering orders, the actual performance data of the processor executing the reordered instructions can be obtained, so that the target reordering order determined in the plurality of reordering orders is more suitable for the processor, thereby improving the efficiency of the processor in running instructions, and improving the efficiency and effect of instruction reordering. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0012] Figure 1 A flow chart of the instruction re-arrangement method provided in an embodiment of the present application;

[0013] Figure 2 Another flow chart of the instruction re-arrangement method provided in the embodiment of the present application;

[0014] Figure 3 A structural diagram of a directed acyclic graph provided in an embodiment of the present application;

[0015] Figure 4 A schematic diagram of determining the order of re-arranging instructions according to a directed acyclic graph provided in an embodiment of the present application;

[0016] Figure 5 A schematic diagram of determining the order of re-arranging instructions according to a directed acyclic graph provided in an embodiment of the present application;

[0017] Figure 6 A schematic diagram of determining the order of re-arranging instructions according to a directed acyclic graph provided in an embodiment of the present application;

[0018] Figure 7 A schematic diagram of determining the order of re-arranging instructions according to a directed acyclic graph provided in an embodiment of the present application;

[0019] Figure 8 A schematic diagram of the structure of an instruction re-arrangement device provided in an embodiment of the present application;

[0020] Fig. 9 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the embodiments of the present application should fall within the scope of protection of the embodiments of the present application.

[0022] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.

[0023] Figure 1 A flowchart of the instruction re-arrangement method provided in the embodiment of the present application. The instruction re-arrangement method provided in the present embodiment can be executed by an instruction re-arrangement device or an electronic device. Figure 1 As shown, the instruction reordering method provided in this embodiment may include:

[0024] S101. Obtain instruction fragments to be reordered from instructions to be executed by a processor, where the instruction fragments include at least one loop body.

[0025] The instructions to be executed by the processor refer to instructions that have been processed by the compiler and can be understood and executed by the processor. Optionally, the instructions to be executed by the processor may include assembly instructions or machine instructions.

[0026] The instruction fragments to be rearranged include at least one loop body. The loop body, also known as the loop structure, is a computer processing process in a programming language that repeatedly executes certain codes. The loop body is executed according to the loop condition and the number of loops.

[0027] Since the loop body needs to be executed multiple times and the execution time is long, the instruction fragments to be rearranged including at least one loop body are rearranged to optimize the instruction execution order and improve the efficiency of the processor in executing instructions.

[0028] Optionally, the loop body may include a sub-loop body and / or a called function.

[0029] Optionally, the instruction fragments to be re-arranged also include other types of instructions besides the loop body.

[0030] It should be noted that this embodiment does not limit the number of instruction fragments to be reordered, the total number of instructions in the instruction fragments, the number of loop bodies in the instruction fragments, the number of instructions in the loop body, the number of sub-loop bodies in the loop body, and the number of called functions in the loop body.

[0031] For example, there are three instruction fragments, represented as instruction fragments 1 to 3. Instruction fragment 1 includes loop body 1 and 5 other instructions, loop body 1 does not include sub-loop body and called function, and loop body 1 includes 10 instructions. Instruction fragment 2 includes loop body 2, and loop body 2 includes sub-loop body 21. Instruction fragment 3 includes loop body 31, loop body 32 and called function 33, loop body 31 does not include sub-loop body and called function, and loop body 32 includes sub-loop body 321 and called function 322.

[0032] In the subsequent steps of this embodiment, the processor is used to run the instruction fragments after instruction reordering to compare the reordering performance. In order to improve the efficiency of performance verification, the number of loop bodies included in the instruction fragments to be reordered can be small. For example, when the instruction fragments to be reordered include a loop body, a better reordering order can be obtained for the loop body.

[0033] Optionally, the loop body included in the instruction fragment to be reordered may satisfy at least one of the following conditions:

[0034] The number of instructions in the loop body is greater than a first preset number;

[0035] The loop body includes sub-loop bodies, and the number of the sub-loop bodies is greater than a second preset number;

[0036] The loop body includes called functions, and the number of the called functions is greater than a third preset number.

[0037] In this embodiment, the values ​​of the first preset number, the second preset number and the third preset number are not limited.

[0038] It can be understood that the more instructions there are in the loop body, the more sub-loop bodies included in the loop body, and the more called functions in the loop body, the longer the execution time of the loop body will be, and the greater the efficiency of the processor in executing instructions will be improved by reordering the instructions for the loop body.

[0039] S102, analyzing the dependencies of the instructions in the instruction fragments, and determining multiple rearrangement orders of the instruction fragments according to the dependencies between the instructions in the instruction fragments, wherein the order of the instructions with dependencies in the instruction fragments and the rearrangement order remains unchanged.

[0040] The so-called dependency between instructions means that the execution of one instruction needs to depend on another instruction, and the order of the dependent instructions needs to remain unchanged, otherwise it will affect the instruction semantics.

[0041] Give an example.

[0042] Instruction 1: add x0,x1,x2. It means adding the value in register x1 and the value in register x2 and writing them into register x0.

[0043] Instruction 2: sub x3,x0,x4. It means subtracting the value in register x0 from the value in register x4 and writing the result to register x3.

[0044] The processor needs to execute instruction 1 first, read the values ​​in registers x1 and x2, complete the addition operation, and write the result to register x0. Then, execute instruction 2, read the values ​​in registers x0 and x4, complete the subtraction operation, and write the result to register x3.

[0045] In the instruction fragment, that is, before the instruction reordering, instruction 1 is located before instruction 2. In the process of instruction reordering, in order to ensure that the instruction semantics do not change, in the reordering sequence, instruction 1 needs to still be located before instruction 2, and the order between instruction 1 and instruction 2 remains unchanged. In the reordering sequence, the order of other instructions before instruction 1, between instruction 1 and instruction 2, and after instruction 2 is not restricted to remain unchanged.

[0046] S103: Run the instruction fragments respectively according to the multiple rearrangement orders of the instruction fragments through the processor, and obtain performance data corresponding to the multiple rearrangement orders of the instruction fragments.

[0047] Even under the same instruction set architecture, processors of different models may have different performance when running the same instructions due to differences in pipeline depth, resource components, and instruction latency. Pipeline depth, also known as pipeline stages, is the process of decomposing the execution process of instructions into multiple independent stages through pipeline technology, with each stage using different hardware resources, thereby achieving parallel processing of instructions and improving the processor's computing speed. Resource components include the hardware in the processor, including but not limited to the arithmetic unit, controller, and registers.

[0048] By using a preset model processor and running the instruction fragments respectively according to multiple rearrangement orders of the instruction fragments, the actual performance data of the multiple rearrangement orders can be obtained for the preset model processor, thereby optimizing the instruction reordering for the preset model processor, obtaining an instruction reordering result that better matches the performance of the preset model processor, and improving the operating efficiency of the preset model processor.

[0049] S104. Determine a target rearrangement sequence among the multiple rearrangement sequences according to the performance data corresponding to the multiple rearrangement sequences of the instruction fragments.

[0050] The target re-arrangement order is a re-arrangement order with better performance among the performance data corresponding to the various re-arrangement orders of the instruction fragments. The performance data corresponding to the various re-arrangement orders can be sorted from best to worst. In one implementation, the re-arrangement order corresponding to the first ranked order is determined as the target re-arrangement order. In another implementation, any one of the P re-arrangement orders ranked in front is determined as the target re-arrangement order, where P is a positive integer. In yet another implementation, a performance threshold value is preset, and any one of the at least one re-arrangement order ranked in front and having performance data better than the performance threshold value is determined as the target re-arrangement order.

[0051] It can be seen that the instruction reordering method provided in this embodiment obtains instruction fragments to be reordered including at least one loop body from the instructions to be executed by the processor, performs dependency analysis and reorders the instructions in the instruction fragments, and obtains multiple reordering orders of the instruction fragments. Since the loop body is executed many times, reordering the instructions in the loop body can greatly improve the operating efficiency of the processor. Moreover, in the process of instruction reordering, the order between instructions with dependencies remains unchanged, ensuring that the program semantics is not changed and the correctness of instruction execution is guaranteed. By using a preset model processor to run the instruction fragments according to multiple reordering orders of the instruction fragments, and obtaining the performance data corresponding to the multiple reordering orders of the instruction fragments, the actual reordering performance test can be performed on the preset model processor, and then, according to the performance data corresponding to the multiple reordering orders of the instruction fragments, the target reordering order is determined in the multiple reordering orders. The instruction re-arrangement method provided in this embodiment does not require manual re-ordering of instructions. It re-orders the instructions in the loop body according to the dependencies between the instructions and obtains performance data through actual operation of the processor, thereby ensuring that the program semantics and execution correctness are not changed, improving the efficiency of the processor in running instructions, and improving the efficiency and effect of instruction re-ordering.

[0052] In a feasible implementation, in S101, obtaining instruction fragments to be reordered from instructions to be executed by the processor includes:

[0053] Get the instruction file to be executed by the processor;

[0054] Generate an abstract syntax tree based on the instruction file;

[0055] Recognize the loop statement structure in the abstract syntax tree to obtain at least one loop body;

[0056] The instruction file is segmented according to at least one loop body to obtain instruction segments to be rearranged.

[0057] Among them, the Abstract Syntax Tree (AST) is an abstract representation of the grammatical structure of the source code, which represents the grammatical structure of the programming language in a tree form. The basic unit of AST is the node, and each node on the tree represents a structure in the source code, such as an expression, statement, declaration, etc. The node includes a parent node and zero, one or more child nodes. The node has a type attribute to indicate the type of node, such as function declaration, variable declaration, binary expression, etc. The node also has position information, indicating the position of the node in the source code. The structure of AST is hierarchical and can reflect the nested structure of the code. For example, the statements inside a function are child nodes of the function node.

[0058] In this implementation, an abstract syntax tree is generated according to an instruction file to be executed by a processor. By analyzing the abstract syntax tree, usually traversing from the root node of the abstract syntax tree, recursively accessing the child nodes, and identifying the loop statement structure, a loop body in the instruction file can be obtained. The number of loop bodies is at least one, assuming that there are M loop bodies, where M is a positive integer. The instructions in the instruction file are fragmented according to the M loop bodies to obtain at least one instruction fragment to be rearranged. Among them, the instruction fragment to be rearranged includes at least one loop body among the M loop bodies.

[0059] By generating an abstract syntax tree for the instruction file to be executed by the processor, identifying the loop body and segmenting it, the instructions in the instruction file are segmented and reordered. Subsequently, the instruction sequence is optimized for each instruction fragment and the actual reordering performance is obtained. Since the number of instructions in the instruction fragment is small, the efficiency and effect of instruction reordering are improved.

[0060] Optionally, the instruction file includes an assembly instruction file or a binary executable file.

[0061] In a feasible implementation, in S102, dependency analysis is performed on the instructions in the instruction fragments, and multiple rearrangement orders of the instruction fragments are determined according to the dependency between the instructions in the instruction fragments, including:

[0062] Determine whether the loop body includes a sub-loop body and / or a called function;

[0063] If the loop body includes a sub-loop body and / or a called function, instructions in the sub-loop body and / or the called function are rearranged to obtain a plurality of rearrangement sequences corresponding to the sub-loop body and / or the called function;

[0064] According to the multiple rearrangement sequences respectively corresponding to the sub-loop bodies and / or the called functions, multiple rearrangement sequences of the instruction fragments are determined.

[0065] Specifically, the loop body may also include a sub-loop body or a called function, and the sub-loop body or the called function usually includes multiple instructions. In the case where the loop body includes a sub-loop body or a called function, it is necessary to reorder the multiple instructions in the sub-loop body, reorder the multiple instructions in the called function, and continuously recurse the sub-loop body and the called function in the loop body to complete the reordering of all instructions in the loop.

[0066] By reordering the instructions of the sub-loop bodies and called functions included in the loop body, a more comprehensive sorting result of the instruction reordering of the loop body is obtained, which provides more options for the subsequent selection of a reordering order with better performance for a preset model of processor, and improves the effect of instruction reordering.

[0067] Optionally, instructions in the sub-loop body and / or the called function are rearranged to obtain multiple rearrangement sequences corresponding to the sub-loop body and / or the called function, including:

[0068] Perform dependency analysis on the instructions in the sub-loop body, and determine multiple rearrangement orders of the sub-loop body according to the dependency between the instructions in the sub-loop body. And / or,

[0069] Dependency analysis is performed on the instructions in the called function, and multiple rearrangement orders of the called function are determined according to the dependencies between the instructions in the called function.

[0070] In this implementation, instructions of the sub-loop body and the called function are reordered respectively.

[0071] Next, the dependencies between instructions are explained.

[0072] Dependencies include: data dependency, hardware resource dependency, function call dependency, and instruction type dependency.

[0073] Data dependency: used to indicate that the input operand and / or output operand of the first instruction has a sequence dependency with the input operand and / or output operand of the second instruction.

[0074] Optionally, the data dependency includes read after write (RAW) data dependency, write after write (WAW) data dependency and read after read (RAW) data dependency.

[0075] The following example illustrates this.

[0076] For read-after-write (RAW) data dependencies.

[0077] Instruction 1: add x0,x1,x2. It means adding the value in register x1 and the value in register x2 and writing them into register x0.

[0078] Instruction 2: sub x3,x0,x4. It means subtracting the value in register x0 from the value in register x4 and writing the result to register x3.

[0079] The processor writes register x0 when executing instruction 1 and reads register x0 when executing instruction 2. The processor writes the register first and then reads the register. If the execution order of instruction 1 and instruction 2 is changed, when the processor executes instruction 2, it will get an incorrect value when reading register x0.

[0080] For write-after-write (WAW) data dependencies.

[0081] Instruction 1: add x0,x1,x2. It means adding the value in register x1 and the value in register x2 and writing them into register x0.

[0082] Instruction 2: sub x0,x3,x4. It means subtracting the value in register x3 from the value in register x4 and writing the result to register x0.

[0083] The processor writes register x0 when executing instruction 1, and also writes register x0 when executing instruction 2. There are two operations of writing registers. If the execution order of instruction 1 and instruction 2 is changed, the wrong value will be written to register x0.

[0084] For read-after-write (RAW) data dependencies.

[0085] Instruction 1: add x0,x1,x2. It means adding the value in register x1 and the value in register x2 and writing them into register x0.

[0086] Instruction 2: sub x1,x3,x4. It means subtracting the value in register x3 from the value in register x4 and writing the result to register x1.

[0087] When the processor executes instruction 1, it reads the value in register x1 as an addend, and when it executes instruction 2, it subtracts the value in register x3 from the value in register x4 and writes the result to register x1. If the execution order of instructions 1 and 2 is changed, the value in register x1 will change, and an incorrect value will be read from register x1 when instruction 1 is executed.

[0088] Hardware resource dependency: used to indicate that the first instruction and the second instruction have a sequential dependency on the processing of hardware resources.

[0089] Optionally, the hardware resource dependency includes a status register dependency. The status register dependency is used to indicate that the first instruction and the second instruction have a sequential dependency on the reading or modification of the status register.

[0090] The status register is also called the condition code register or CPSR, which stands for current program status register. The status register can be operated bit by bit. The status register usually stores two types of information: one is various status information or condition codes that reflect the result of the current instruction execution, such as whether there is a carry (CF bit or C bit), whether there is an overflow (OV bit), whether the result is positive or negative (SF bit), whether the result is zero (ZF bit), parity flag (P bit), etc.; the other is to store control information, such as enabling interrupts (IF bit), tracking flag (TF bit), etc., which is also called the program status word register (PSW) or flag register (FR).

[0091] Some instructions modify and read the status register, and there are also order dependencies between these instructions. For example, instructions include but are not limited to: adds, subs, adcs, cmp, condition branch.

[0092] The adds instruction is used to perform the addition operation of a signed number. In the addition operation, if a carry occurs, the C bit (carry flag) of the status register will be updated.

[0093] The subs instruction is used to perform a subtraction operation and store the result in the destination register. If a borrow operation occurs, the C bit (carry flag) in the status register is set to 0; if no borrow operation occurs, it is set to 1.

[0094] The adcs instruction is used to perform addition of the values ​​in two registers, and consider the C bit (carry flag) in the status register during the addition process, write the result to the destination register, and update the status flag (NZCV flag) at the same time.

[0095] The cmp instruction is used to compare the sizes of two operands and set the flag bits in the status register according to the comparison results. The basic format of the CMP instruction is: cmp source, destination, where source and destination can be registers, memory addresses, or immediate values. The flag bits include the zero flag (ZF), the sign flag (SF), or the carry flag (CF). For the ZF bit, if the source and destination are equal, the ZF bit is set to 1, otherwise it is set to 0. For the SF bit, it indicates the positive or negative of the signed number. If the source is less than the destination, the SF bit is set to 1, otherwise it is 0. For the CF bit, it indicates the overflow of the unsigned number. If the source is greater than the destination, the CF bit is set to 1, otherwise it is set to 0.

[0096] The condition branch instruction is a conditional branch instruction that is used to change the execution flow of a program when a preset condition is met.

[0097] Function call dependency: used to indicate that during the function call process, there is a sequence dependency between the function's input parameters, return values, and function call instructions.

[0098] Specifically, before a function call instruction, an instruction is required to pass parameters. After the function call, there may be other instructions that use the return value of the function. There is also a sequence dependency between these instructions.

[0099] Instruction type dependency: used to indicate that there is a sequence dependency between the execution of the first type of instructions and the second type of instructions.

[0100] Optionally, the instruction type dependency includes a memory access instruction dependency and a barrier instruction dependency.

[0101] The memory access instruction dependency is used to indicate that there is a sequence dependency between the read instruction and the store instruction.

[0102] The barrier instruction dependency relationship is used to indicate that there is a sequence dependency between the barrier instruction and the instructions before the barrier instruction.

[0103] Among them, the read instruction is also called the load instruction, which is used to read data from the register. The store instruction is used to write data into the register. The barrier instruction is used to control the order of memory operations in the processor to ensure data consistency and thread safety in concurrent programming. The barrier instruction ensures the order of memory operations by preventing the reordering of some instructions, thereby avoiding errors caused by out-of-order execution.

[0104] The following example illustrates the memory access instruction dependency relationship, wherein the order of read instructions is allowed to be swapped, but the order of read instructions and storage instructions is not allowed to be swapped, and the order of storage instructions is not allowed to be swapped.

[0105] Example 1:

[0106] Instruction 1: ldr x0,[x9]. This means loading the value at memory address x9 into register x0.

[0107] Instruction 2: str x1,[x9]. This means storing the value in register x1 to memory address x9.

[0108] If the order of instruction 1 and instruction 2 is changed, and instruction 2 is placed before instruction 1, then the processor will read an incorrect value when executing instruction 1.

[0109] Example 2:

[0110] Instruction 1: str x0,[x9]. This means storing the value in register x0 to memory address x9.

[0111] Instruction 2: str x1,[x9]. This means storing the value in register x1 to memory address x9.

[0112] The following describes various ways of determining the rearrangement order of the instruction fragments.

[0113] In a feasible implementation, in S102, multiple re-arrangement orders of the instruction fragments are determined according to the dependency relationship between the instructions in the instruction fragments, including:

[0114] Constructing a directed acyclic graph according to the dependency relationship between instructions in the instruction fragment;

[0115] A topological sort is performed based on the directed acyclic graph to determine multiple rearrangement orders of the instruction fragments.

[0116] Wherein, the directed acyclic graph includes nodes and edges between nodes, nodes represent instructions, edges between nodes represent dependencies between instructions, and the direction of the edges between nodes is from the instruction executed earlier to the instruction executed later. For example, referring to the example of read-after-write (RAW) data dependency in the above data dependency, instruction 1 is: add x0, x1, x2, instruction 2 is sub x3, x0, x4, instruction 1 and instruction 2 have a dependency relationship, instruction 1 is executed before instruction 2, then, in the directed acyclic graph, instruction 1 and instruction 2 are two nodes in the graph, and the direction of the edge between the two nodes is from the node of instruction 1 to the node of instruction 2.

[0117] Among them, topological sort is a sorting method for directed acyclic graphs, which can traverse all nodes in the directed acyclic graph. Moreover, when traversing each node, its predecessor node has been traversed. The traversal order of the nodes is called the topological sequence.

[0118] In this implementation, a directed acyclic graph is constructed according to the dependencies between instructions in the instruction fragments, and multiple rearrangement orders of the instruction fragments are determined by topologically sorting the nodes in the directed acyclic graph. The implementation is efficient and can obtain more rearrangement orders, and finally determine a better rearrangement order for a preset model processor, thereby improving the instruction rearrangement effect.

[0119] Optional, Figure 2 Another flow chart of the instruction reordering method provided in the embodiment of the present application is as follows: Figure 2 As shown, a topological sorting is performed based on the directed acyclic graph to determine multiple rearrangement orders of the instruction fragments, including:

[0120] S201. Obtain the in-degree of a node in a directed acyclic graph, where the in-degree is the sum of the number of edges pointing to the node.

[0121] S202. Select a target node from the nodes in the directed acyclic graph, where the in-degree of the target node is 0.

[0122] S203, deleting the target node and the edge with the target node as the endpoint in the directed acyclic graph, and updating the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph.

[0123] S204, looping through the steps of selecting and deleting target nodes in the directed acyclic graph until the number of nodes in the directed acyclic graph is zero, and obtaining a rearrangement order of the instruction fragments according to the order of the selected target nodes.

[0124] The following example is used to illustrate this. Assume that the instruction fragment includes the following instructions:

[0125] Instruction 1: add x1,x1,#1

[0126] Instruction 2: add x2,x2,#1

[0127] Instruction 3: add x0,x1,x2

[0128] Instruction 4: add x3,x4,x0

[0129] Instruction 5: add x1,x5,x6

[0130] Through the dependency analysis between instructions, the following five pairs of instructions with dependency are obtained: <1,5>, <1,3>, <2,3>, <3,4>, <3,5>. The instructions in front of the brackets need to be executed first, and the instructions behind need to be executed later.

[0131] Construct a directed acyclic graph based on the dependencies between instructions in the instruction fragment, see Figure 3 .exist Figure 3 In the example, the node number is represented by the instruction number, and instructions 1 to 5 correspond to nodes 1 to 5 respectively. For the instruction pair <1,5> with a dependency relationship, the edge between node 1 and node 5 points from node 1 to node 5. For the instruction pair <1,3> with a dependency relationship, the edge between node 1 and node 3 points from node 1 to node 3. The edge pointing principle of other instruction pairs with dependency relationships is the same.

[0132] In an example of determining the reordering order of instruction fragments, see Figure 4 .

[0133] Initially, the directed acyclic graph includes 5 nodes, namely, node 1 to node 5. The in-degree of node 1 is 0, the in-degree of node 2 is 0, the in-degree of node 3 is 2, the in-degree of node 4 is 1, and the in-degree of node 5 is 2. For the first time, one of the nodes with in-degree 0 is selected as the target node, specifically node 2. Then, in the directed acyclic graph, node 2 and the edge with node 2 as the endpoint are deleted, that is, node 2 and the edge between node 2 and node 3 are deleted.

[0134] After selecting the target node 2 for the first time, update the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph. At this time, the directed acyclic graph includes 4 nodes, including: node 1, node 3, node 4 and node 5. The in-degree of node 1 is 0, the in-degree of node 3 is 0, the in-degree of node 4 is 1, and the in-degree of node 5 is 2. Select a node as the target node for the second time among the nodes with in-degree 0, specifically select node 1. Then, delete node 1 and the edge with node 1 as the endpoint in the directed acyclic graph, that is, delete node 1, the edge between node 1 and node 3, and the edge between node 1 and node 5.

[0135] After selecting the target node 1 for the second time, the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph are updated. At this time, the directed acyclic graph includes 3 nodes, including: node 3, node 4 and node 5. The in-degree of node 3 is 0, the in-degree of node 4 is 1, and the in-degree of node 5 is 1. Select a node as the target node for the third time among the nodes with in-degree 0, specifically select node 3. Then, delete node 3 and the edge with node 3 as the endpoint in the directed acyclic graph, that is, delete node 3, the edge between node 3 and node 4, and the edge between node 3 and node 5.

[0136] After selecting the target node 3 for the third time, the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph are updated. At this time, the directed acyclic graph includes 2 nodes, including: node 4 and node 5. The in-degree of node 4 is 0, and the in-degree of node 5 is 0. Select a node as the target node for the fourth time among the nodes with in-degree 0, specifically select node 4. Then, delete node 4 and the edge with node 4 as the endpoint in the directed acyclic graph, that is, delete node 4.

[0137] After selecting the target node 4 for the fourth time, the in-degrees of the directed acyclic graph and the remaining nodes in the directed acyclic graph are updated. At this time, the directed acyclic graph includes 1 node, which is node 5. The last selected target node is node 5. The number of nodes in the directed acyclic graph is 0, and all nodes have been traversed.

[0138] The order of the selected target nodes is: node 2, node 1, node 3, node 4, node 5. Then, the rearrangement order of the instruction fragments obtained according to the order of the selected target nodes is: instruction 2, instruction 1, instruction 3, instruction 4, instruction 5.

[0139] Figure 5 to Figure 7 Based on Figure 3 The selection process is similar to Figure 4 The principles shown are similar and will not be described in detail. Figure 5 The determined rearrangement order of the instruction fragments is: instruction 2, instruction 1, instruction 3, instruction 5, instruction 4. Figure 6 The determined rearrangement order of the instruction fragments is: instruction 1, instruction 2, instruction 3, instruction 4, instruction 5. Figure 7 The determined rearrangement order of the instruction fragments is: instruction 1, instruction 2, instruction 3, instruction 5, instruction 4.

[0140] It can be seen that by topologically sorting the nodes in the directed acyclic graph and determining multiple rearrangement orders of the instruction fragments, the implementation method is efficient and can obtain a comprehensive rearrangement order.

[0141] Optionally, in a feasible implementation of this embodiment, in S103, the processor executes the instruction fragments respectively according to the multiple rearrangement orders of the instruction fragments to obtain the performance data corresponding to the multiple rearrangement orders of the instruction fragments, including:

[0142] According to the various rearrangement sequences of the instruction fragments, the multiple instructions in the instruction fragments are respectively filled into the loop body to be tested, so as to obtain the loop bodies to be tested corresponding to the various rearrangement sequences of the instruction fragments.

[0143] The processor runs a plurality of loop bodies to be tested respectively and reaches a preset number of loop runs, thereby obtaining performance data corresponding to a plurality of rearrangement sequences of the instruction fragments.

[0144] It should be noted that this embodiment does not limit the value of the preset number of cyclic operations, for example, 4096 times.

[0145] For example. Assume that there are three kinds of rearrangement orders of the instruction fragment, represented as reordering 1, reordering 2, and reordering 3. According to reordering 1, multiple instructions in the instruction fragment are filled into the loop body to be tested to obtain loop body 1 to be tested. Assume that the processor model is Yitian 710 chip, use Yitian 710 chip to execute test loop body 1, and loop it 4096 times to obtain performance data 1 corresponding to reordering 1. Similarly, according to reordering 1, loop body 2 to be tested is obtained, and Yitian 710 chip is used to loop test loop body 2 4096 times to obtain performance data 2 corresponding to reordering 2; according to reordering 3, loop body 3 to be tested is obtained, and Yitian 710 chip is used to loop test loop body 3 4096 times to obtain performance data 3 corresponding to reordering 3.

[0146] In this implementation, a loop body to be tested is constructed according to the rearranged order of instruction fragments, and a preset model processor is used to execute the loop a certain number of times, so as to obtain the actual performance data running on the preset model processor. The differences between different models of processors are taken into consideration, and the ordering of instructions in the instruction fragments is optimized for the instruction fragments and the preset model processor, thereby improving the operating efficiency of the preset model processor and the efficiency and effect of instruction reordering.

[0147] The performance data may reflect the efficiency of the processor in executing instructions. Optionally, the performance data includes but is not limited to the number of clock cycles of the processor running the instruction fragments according to the rearranged order of the instruction fragments.

[0148] Reference Figure 8 The present application also provides an instruction re-arrangement device for executing the instruction re-arrangement method provided in the embodiment of the present application. The technical principle and technical effect are similar and will not be described in detail. Figure 8 As shown, the instruction re-arrangement device provided in this embodiment includes:

[0149] An acquisition module 801 is used to acquire instruction fragments to be rearranged from instructions to be executed by a processor, wherein the instruction fragments include at least one loop body;

[0150] A rearrangement module 802 is used to perform dependency analysis on the instructions in the instruction fragments, and determine multiple rearrangement orders of the instruction fragments according to the dependency between the instructions in the instruction fragments; wherein the order of the instructions with dependency in the instruction fragments and the rearrangement order remains unchanged;

[0151] A performance module 803 is used to respectively execute the instruction fragments according to the multiple rearrangement orders of the instruction fragments through the processor to obtain performance data corresponding to the multiple rearrangement orders of the instruction fragments;

[0152] The determination module 804 determines a target rearrangement sequence from among the multiple rearrangement sequences of the instruction fragments according to the performance data respectively corresponding to the multiple rearrangement sequences of the instruction fragments.

[0153] Optionally, the rearrangement module 802 is used to:

[0154] Determine whether the loop body includes a sub-loop body and / or a called function;

[0155] If the loop body includes the sub-loop body and / or the called function, instructions in the sub-loop body and / or the called function are respectively rearranged to obtain a plurality of rearrangement sequences corresponding to the sub-loop body and / or the called function;

[0156] According to the multiple rearrangement sequences respectively corresponding to the sub-loop body and / or the called function, multiple rearrangement sequences of the instruction fragments are determined.

[0157] Optionally, the rearrangement module 802 is used to:

[0158] Performing dependency analysis on the instructions in the sub-loop body, and determining multiple rearrangement orders of the sub-loop body according to the dependency between the instructions in the sub-loop body; and / or,

[0159] Dependency analysis is performed on the instructions in the called function, and multiple rearrangement orders of the called function are determined according to the dependencies between the instructions in the called function.

[0160] Optionally, the rearrangement module 802 is used to:

[0161] Constructing a directed acyclic graph according to the dependency relationship between the instructions in the instruction fragments; wherein the directed acyclic graph includes nodes and edges between the nodes, the nodes represent instructions, the edges between the nodes represent the dependency relationship between the instructions, and the direction of the edges between the nodes is from the previously executed instruction to the later executed instruction;

[0162] A topological sorting is performed according to the directed acyclic graph to determine a plurality of rearrangement orders of the instruction fragments.

[0163] Optionally, the rearrangement module 802 is used to:

[0164] Obtaining the in-degree of the node in the directed acyclic graph, where the in-degree is the sum of the number of edges pointing to the node;

[0165] Selecting a target node from the nodes in the directed acyclic graph, wherein the in-degree of the target node is 0;

[0166] Deleting the target node and the edge with the target node as the endpoint in the directed acyclic graph, and updating the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph;

[0167] The steps of selecting and deleting target nodes from the nodes in the directed acyclic graph are executed cyclically until the number of nodes in the directed acyclic graph is zero, and the rearrangement order of the instruction fragments is obtained according to the order of the selected target nodes.

[0168] Optionally, the performance module 803 is used to:

[0169] According to the multiple rearrangement orders of the instruction fragments, the multiple instructions in the instruction fragments are respectively filled into the loop body to be tested, so as to obtain the loop bodies to be tested corresponding to the multiple rearrangement orders of the instruction fragments;

[0170] The processor executes the plurality of loop bodies to be tested respectively and reaches a preset number of loop execution times, thereby obtaining performance data corresponding to the plurality of rearrangement sequences of the instruction fragments.

[0171] Optionally, the performance data includes the number of clock cycles of executing the instruction fragments by the processor in the rearranged order of the instruction fragments.

[0172] Optionally, the dependency relationship includes: data dependency relationship, hardware resource dependency relationship, function call dependency relationship and instruction type dependency relationship;

[0173] The data dependency relationship is used to indicate that an input operand and / or an output operand of a first instruction is sequentially dependent on an input operand and / or an output operand of a second instruction;

[0174] The hardware resource dependency is used to indicate that the first instruction and the second instruction have a sequential dependency on the processing of the hardware resources;

[0175] The function call dependency is used to indicate that during the function call process, there is a sequence dependency between the function's input parameters, return values ​​and the function call instruction;

[0176] The instruction type dependency is used to indicate that there is a sequence dependency between the execution of the first type of instructions and the second type of instructions.

[0177] Optionally, the instruction type dependency includes a memory access instruction dependency and a barrier instruction dependency;

[0178] The memory access instruction dependency is used to indicate that there is a sequence dependency between the read instruction and the store instruction;

[0179] The barrier instruction dependency relationship is used to indicate that a barrier instruction has a sequence dependency with instructions before the barrier instruction.

[0180] Optionally, the hardware resource dependency includes a status register dependency;

[0181] The status register dependency is used to indicate that the first instruction and the second instruction have a sequence dependency on reading or modifying the status register.

[0182] Optionally, the acquisition module 801 is used to:

[0183] Obtaining an instruction file to be executed by the processor;

[0184] Generate an abstract syntax tree according to the instruction file;

[0185] Identify a loop statement structure in the abstract syntax tree to obtain at least one loop body;

[0186] The instruction file is segmented according to the at least one loop body to obtain the instruction fragments to be rearranged.

[0187] Optionally, the instruction file includes an assembly instruction file or a binary executable file.

[0188] Optionally, the loop body satisfies at least one of the following conditions:

[0189] The number of the plurality of instructions in the loop body is greater than a first preset number;

[0190] The loop body includes sub-loop bodies, and the number of the sub-loop bodies is greater than a second preset number;

[0191] The loop body includes called functions, and the number of the called functions is greater than a third preset number.

[0192] Reference Fig. 9 , shows a schematic diagram of the structure of an electronic device according to an embodiment of the present application, and the specific embodiment of the present application does not limit the name and specific implementation of the electronic device. For example, the electronic device may also be referred to as a terminal device, and the electronic device may include a mobile device, a tablet computer, a laptop computer, a desktop computer, a wearable computer, a game console, a media player, a vehicle entertainment system, and / or any other suitable type of device.

[0193] like Fig. 9As shown, the electronic device may include: a processor (processor) 502 , a communication interface (Communications Interface) 504 , a memory (memory) 506 , and a communication bus 508 .

[0194] in:

[0195] The processor 502 , the communication interface 504 , and the memory 506 communicate with each other via a communication bus 508 .

[0196] The communication interface 504 is used to communicate with other electronic devices or servers.

[0197] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above method embodiment.

[0198] Specifically, the program 510 may include program codes, which include computer operation instructions.

[0199] The processor 502 may be a CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0200] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0201] The program 510 may include multiple computer instructions. Specifically, the program 510 may enable the processor 502 to execute operations corresponding to the method described in any of the aforementioned method embodiments through the multiple computer instructions.

[0202] The specific implementation of each step in program 510 can refer to the corresponding description of the corresponding steps and units in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, which will not be repeated here.

[0203] The present application also provides a computer storage medium on which a computer program is stored, and when the program is executed by a processor, the method described in any of the above-mentioned multiple method embodiments is implemented. The computer storage medium includes but is not limited to: a compact disc read-only memory (CD-ROM), a random access memory (RAM), a floppy disk, a hard disk or a magneto-optical disk, etc.

[0204] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to any one of the above-mentioned multiple method embodiments.

[0205] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0206] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or implemented as a computer code originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded through a network and stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., a random access memory (RAM), a read-only memory (ROM), a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.

[0207] Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.

[0208] The above implementation methods are only used to illustrate the embodiments of the present application, and are not limitations on the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The scope of patent protection of the embodiments of the present application should be limited by the claims.

[0209] The embodiment of the present application also provides A1. an instruction reordering method, comprising:

[0210] Obtaining instruction fragments to be reordered from instructions to be executed by a processor, wherein the instruction fragments include at least one loop body;

[0211] Performing dependency analysis on the instructions in the instruction fragments, and determining multiple rearrangement orders of the instruction fragments according to the dependency between the instructions in the instruction fragments; wherein the order of the instructions with dependency in the instruction fragments and the rearrangement order remains unchanged;

[0212] The processor executes the instruction fragments respectively according to the multiple rearrangement orders of the instruction fragments to obtain performance data corresponding to the multiple rearrangement orders of the instruction fragments;

[0213] A target rearrangement sequence is determined among the multiple rearrangement sequences according to the performance data respectively corresponding to the multiple rearrangement sequences of the instruction fragments.

[0214] A2. The method according to A1, wherein the performing dependency analysis on the instructions in the instruction fragments and determining multiple rearrangement orders of the instruction fragments according to the dependencies between the instructions in the instruction fragments comprises:

[0215] Determine whether the loop body includes a sub-loop body and / or a called function;

[0216] If the loop body includes the sub-loop body and / or the called function, instructions in the sub-loop body and / or the called function are respectively rearranged to obtain a plurality of rearrangement sequences corresponding to the sub-loop body and / or the called function;

[0217] According to the multiple rearrangement sequences respectively corresponding to the sub-loop body and / or the called function, multiple rearrangement sequences of the instruction fragments are determined.

[0218] A3. The method according to A2, wherein the instructions in the sub-loop body and / or the called function are respectively rearranged to obtain a plurality of rearrangement sequences corresponding to the sub-loop body and / or the called function, including:

[0219] Performing dependency analysis on the instructions in the sub-loop body, and determining multiple rearrangement orders of the sub-loop body according to the dependency between the instructions in the sub-loop body; and / or,

[0220] Dependency analysis is performed on the instructions in the called function, and multiple rearrangement orders of the called function are determined according to the dependencies between the instructions in the called function.

[0221] A4. The method according to A1, wherein the determining of the multiple rearrangement orders of the instruction fragments according to the dependency relationship between the instructions in the instruction fragments comprises:

[0222] Constructing a directed acyclic graph according to the dependency relationship between the instructions in the instruction fragments; wherein the directed acyclic graph includes nodes and edges between the nodes, the nodes represent instructions, the edges between the nodes represent the dependency relationship between the instructions, and the direction of the edges between the nodes is from the previously executed instruction to the later executed instruction;

[0223] A topological sorting is performed according to the directed acyclic graph to determine a plurality of rearrangement orders of the instruction fragments.

[0224] A5. The method according to A4, wherein the topological sorting is performed according to the directed acyclic graph to determine the multiple rearrangement orders of the instruction fragments, including:

[0225] Obtaining the in-degree of the node in the directed acyclic graph, where the in-degree is the sum of the number of edges pointing to the node;

[0226] Selecting a target node from the nodes in the directed acyclic graph, wherein the in-degree of the target node is 0;

[0227] Deleting the target node and the edge with the target node as the endpoint in the directed acyclic graph, and updating the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph;

[0228] The steps of selecting and deleting target nodes from the nodes in the directed acyclic graph are executed cyclically until the number of nodes in the directed acyclic graph is zero, and the rearrangement order of the instruction fragments is obtained according to the order of the selected target nodes.

[0229] A6. The method according to A1, wherein the step of executing the instruction fragments respectively according to the multiple rearrangement orders of the instruction fragments by the processor to obtain performance data corresponding to the multiple rearrangement orders of the instruction fragments respectively comprises:

[0230] According to the multiple rearrangement orders of the instruction fragments, the multiple instructions in the instruction fragments are respectively filled into the loop body to be tested, so as to obtain the loop bodies to be tested corresponding to the multiple rearrangement orders of the instruction fragments;

[0231] The processor executes the plurality of loop bodies to be tested respectively and reaches a preset number of loop execution times, thereby obtaining performance data corresponding to the plurality of rearrangement sequences of the instruction fragments.

[0232] A7. A method according to any one of A1-A6, wherein the performance data includes the number of clock cycles of the processor executing the instruction fragments in the rearranged order of the instruction fragments.

[0233] A8. The method according to any one of A1-A6, wherein the dependency relationship includes: data dependency relationship, hardware resource dependency relationship, function call dependency relationship and instruction type dependency relationship;

[0234] The data dependency relationship is used to indicate that an input operand and / or an output operand of a first instruction is sequentially dependent on an input operand and / or an output operand of a second instruction;

[0235] The hardware resource dependency is used to indicate that the first instruction and the second instruction have a sequential dependency on the processing of the hardware resources;

[0236] The function call dependency is used to indicate that during the function call process, there is a sequence dependency between the function's input parameters, return values ​​and the function call instruction;

[0237] The instruction type dependency is used to indicate that there is a sequence dependency between the execution of the first type of instructions and the second type of instructions.

[0238] A9. The method according to A8, wherein the instruction type dependency includes a memory access instruction dependency and a barrier instruction dependency;

[0239] The memory access instruction dependency is used to indicate that there is a sequence dependency between the read instruction and the store instruction;

[0240] The barrier instruction dependency relationship is used to indicate that a barrier instruction has a sequence dependency with instructions before the barrier instruction.

[0241] A10. The method according to A8, wherein the hardware resource dependency includes a state register dependency;

[0242] The status register dependency is used to indicate that the first instruction and the second instruction have a sequence dependency on reading or modifying the status register.

[0243] A11. The method according to any one of A1-A6, wherein obtaining the instruction fragments to be reordered from the instructions to be executed by the processor comprises:

[0244] Obtaining an instruction file to be executed by the processor;

[0245] Generate an abstract syntax tree according to the instruction file;

[0246] Identify a loop statement structure in the abstract syntax tree to obtain at least one loop body;

[0247] The instruction file is segmented according to the at least one loop body to obtain the instruction fragments to be rearranged.

[0248] A12. The method according to A11, wherein the instruction file comprises an assembly instruction file or a binary executable file.

[0249] A13. A method according to any one of A1-A6, wherein the loop body satisfies at least one of the following conditions: the number of instructions in the loop body is greater than a first preset number;

[0250] The loop body includes sub-loop bodies, and the number of the sub-loop bodies is greater than a second preset number;

[0251] The loop body includes called functions, and the number of the called functions is greater than a third preset number.

Claims

1. An instruction reordering method, comprising: Obtaining instruction fragments to be reordered from instructions to be executed by a processor, wherein the instruction fragments include at least one loop body; Performing dependency analysis on the instructions in the instruction fragments, and determining multiple rearrangement orders of the instruction fragments according to the dependency between the instructions in the instruction fragments; wherein the order of the instructions with dependency in the instruction fragments and the rearrangement order remains unchanged; The processor executes the instruction fragments respectively according to the multiple rearrangement orders of the instruction fragments to obtain performance data corresponding to the multiple rearrangement orders of the instruction fragments; A target rearrangement sequence is determined among the multiple rearrangement sequences according to the performance data respectively corresponding to the multiple rearrangement sequences of the instruction fragments.

2. The method according to claim 1, wherein: The performing dependency analysis on the instructions in the instruction fragments and determining multiple rearrangement orders of the instruction fragments according to the dependency relationships between the instructions in the instruction fragments includes: Determine whether the loop body includes a sub-loop body and / or a called function; If the loop body includes the sub-loop body and / or the called function, instructions in the sub-loop body and / or the called function are respectively rearranged to obtain a plurality of rearrangement sequences corresponding to the sub-loop body and / or the called function; According to the multiple rearrangement sequences respectively corresponding to the sub-loop body and / or the called function, multiple rearrangement sequences of the instruction fragments are determined.

3. The method according to claim 2, wherein: Instructions in the sub-loop body and / or the called function are respectively rearranged to obtain a plurality of rearrangement sequences corresponding to the sub-loop body and / or the called function, including: Performing dependency analysis on the instructions in the sub-loop body, and determining multiple rearrangement orders of the sub-loop body according to the dependency between the instructions in the sub-loop body; and / or, Dependency analysis is performed on the instructions in the called function, and multiple rearrangement orders of the called function are determined according to the dependencies between the instructions in the called function.

4. The method according to claim 1, wherein: The determining of the multiple rearrangement orders of the instruction fragments according to the dependency relationship between the instructions in the instruction fragments includes: Constructing a directed acyclic graph according to the dependency relationship between the instructions in the instruction fragments; wherein the directed acyclic graph includes nodes and edges between the nodes, the nodes represent instructions, the edges between the nodes represent the dependency relationship between the instructions, and the direction of the edges between the nodes is from the previously executed instruction to the later executed instruction; A topological sorting is performed according to the directed acyclic graph to determine a plurality of rearrangement orders of the instruction fragments.

5. The method according to claim 4, wherein: The topological sorting is performed according to the directed acyclic graph to determine multiple rearrangement orders of the instruction fragments, including: Obtaining the in-degree of the node in the directed acyclic graph, where the in-degree is the sum of the number of edges pointing to the node; Selecting a target node from the nodes in the directed acyclic graph, wherein the in-degree of the target node is 0; Deleting the target node and the edge with the target node as the endpoint in the directed acyclic graph, and updating the in-degree of the directed acyclic graph and the remaining nodes in the directed acyclic graph; The steps of selecting and deleting target nodes from the nodes in the directed acyclic graph are executed cyclically until the number of nodes in the directed acyclic graph is zero, and the rearrangement order of the instruction fragments is obtained according to the order of the selected target nodes.

6. The method according to claim 1, wherein: The step of respectively running the instruction fragments according to the multiple rearrangement orders of the instruction fragments by the processor to obtain performance data corresponding to the multiple rearrangement orders of the instruction fragments includes: According to the multiple rearrangement orders of the instruction fragments, the multiple instructions in the instruction fragments are respectively filled into the loop body to be tested, so as to obtain the loop bodies to be tested corresponding to the multiple rearrangement orders of the instruction fragments; The processor executes the plurality of loop bodies to be tested respectively and reaches a preset number of loop execution times, thereby obtaining performance data corresponding to the plurality of rearrangement sequences of the instruction fragments.

7. The method according to any one of claims 1 to 6, wherein: The dependencies include: data dependency, hardware resource dependency, function call dependency and instruction type dependency; The data dependency relationship is used to indicate that an input operand and / or an output operand of a first instruction is sequentially dependent on an input operand and / or an output operand of a second instruction; The hardware resource dependency is used to indicate that the first instruction and the second instruction have a sequential dependency on the processing of the hardware resources; The function call dependency is used to indicate that during the function call process, there is a sequence dependency between the function's input parameters, return values ​​and the function call instruction; The instruction type dependency is used to indicate that there is a sequence dependency between the execution of the first type of instructions and the second type of instructions.

8. The method according to claim 7, wherein: The instruction type dependency relationship includes a memory access instruction dependency relationship and a barrier instruction dependency relationship; The memory access instruction dependency is used to indicate that there is a sequence dependency between the read instruction and the store instruction; The barrier instruction dependency relationship is used to indicate that a barrier instruction has a sequence dependency with instructions before the barrier instruction.

9. The method according to claim 7, wherein: The hardware resource dependency includes a status register dependency; The status register dependency is used to indicate that the first instruction and the second instruction have a sequence dependency on reading or modifying the status register.

10. An electronic device, comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 9.

11. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. A computer program product, comprising computer instructions, wherein the computer instructions instruct a computing device to perform operations corresponding to the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Method for synchronizing multiple hardware engines and computer readable storage medium

    CN121833054A