A GPU Compiler Optimization Method
By identifying and optimizing FMA instructions in the GPU compiler and deleting temporary instructions using optimization templates and rules, the problem of FMA instruction sequence redundancy in the prior art is solved, and the execution efficiency and resource utilization of the processor are improved.
Patent Information
- Application Number
- CN202211563572.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-07
AI Technical Summary
The existing compilation optimization technology fails to effectively simplify and delete FMA-related instruction sequences, resulting in a large amount of redundancy in the generated instructions, limiting the execution speed of the processor.
A GPU compiler optimization method is designed to traverse the basic blocks in the program, identify FMA instructions, and use optimization templates and corresponding rules to match and optimize instructions, delete temporary instructions, and reduce redundant calculations.
The optimized instruction sequence reduces program run time and resource usage, and the generated assembly code is more streamlined.
Smart Images

Figure CN115964048B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method for optimizing a GPU compiler. Background Art
[0002] FMA (Fused Multiply Add) is one of the most common arithmetic operators in CPUs and GPUs. A large number of such operations are included in actual application programs, and these operations are often dependent on each other before and after, making it difficult to execute them in parallel in the processor. Mainstream compilers such as the GCC compiler (GNU Compiler Collection), the LLVM compiler, and the JVM compiler (Java Virtual Machine) perform optimizations such as common sub-expression deletion, expression simplification, code enhancement, and redundancy deletion on multiplication and addition expressions during the optimization phase, without considering simplifying and optimizing FMA-related instructions separately. Since not all backend processors support FMA instructions, relevant optimizations in the middle end of the compiler rarely consider FMA instructions.
[0003] In addition, there is also a lack of separate optimization for multiply-add expressions in backend compilation optimization, such as expression constant simplification. This also results in many redundant instructions not being eliminated. In GPUs, FMA is one of the most frequently used instructions. A large amount of FMA redundancy will limit the execution speed of the processor.
[0004] In view of this, overcoming the defects of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that existing compilation optimization technologies do not perform further expression simplification and expression deletion optimizations on FMA-related instruction sequences, resulting in a large amount of redundancy still existing in the finally generated instructions.
[0006] The present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for optimizing a GPU compiler, including:
[0008] Traverse the basic blocks of each function in the program, traverse the instructions in the basic blocks, and find FMA instructions;
[0009] If the number of instructions in the basic block where the FMA instruction is located is less than or equal to a preset value, then use all the instructions in the basic block as a template matching window; otherwise, use the FMA instruction, the instruction before the FMA instruction, and the two instructions after the FMA instruction together as a template matching window;
[0010] According to the dependency relationships of the instructions in the template matching window, check whether there is at least one temporary instruction in the template matching window; wherein, the temporary instruction is an instruction that is not relied on by any instruction outside the template matching window.
[0011] Use each optimization template to match the first window which is the template matching window with at least one temporary instruction. If the first window successfully matches the corresponding optimization template, optimize the instructions in the first window according to the optimization rules corresponding to the optimization template.
[0012] Wherein, each optimization template includes at least a temporary instruction and a first instruction, and the first instruction depends on the temporary instruction; the optimization rules corresponding to each optimization template replace the temporary variable in the first instruction with other parameters and delete the temporary instruction.
[0013] Preferably, the step of using each optimization template to match the first window, if the first window successfully matches the corresponding optimization template, then optimizing the instructions in the first window according to the optimization rules corresponding to the optimization template specifically includes:
[0014] If there is a temporary instruction and a second FMA instruction before the first FMA instruction in the first window, and the temporary instruction is an ADD instruction, then perform the following process:
[0015] If one of the multiplication variables in the second FMA instruction participates in the addition operation in the temporary instruction as a first variable, and the other multiplication variable participates in the multiplication operation in the first FMA instruction together with the temporary variable, and the addition variable in the first FMA instruction is the same as the addition variable in the second FMA instruction;
[0016] Then replace the temporary variable in the first FMA instruction with a second variable, replace the addition variable in the first FMA instruction with the result variable of the second FMA instruction, and delete the temporary instruction; wherein, the temporary variable is the result variable of the temporary instruction, and the second variable is the other addition variable in the temporary instruction.
[0017] Preferably, the step of using each optimization template to match the first window, if the first window successfully matches the corresponding optimization template, then optimizing the instructions in the first window according to the optimization rules corresponding to the optimization template further includes:
[0018] If there is a temporary instruction and a second FMA instruction before the first FMA instruction in the first window, and the temporary instruction is an FMA instruction, then perform the following process:
[0019] If the two multiplication variables in the second FMA instruction are the same as the two multiplication variables in the first FMA instruction, the addition variable in the first FMA instruction is a temporary variable, and the addition variable in the temporary variable is the same as the addition variable in the second FMA instruction;
[0020] Then replace the two multiplication variables in the first FMA instruction with the two multiplication variables in the temporary instruction, replace the addition variable in the first FMA instruction with the result variable of the second FMA instruction, and delete the temporary instruction; wherein, the temporary variable is the result variable of the temporary instruction.
[0021] Preferably, when using each optimization template to match the first window, if the first window successfully matches the corresponding optimization template, then according to the optimization rules corresponding to the optimization template, optimize each instruction in the first window, and further include:
[0022] If there is a temporary instruction and a second MUL instruction before the first MUL instruction in the first window, and the temporary instruction is an FMA instruction, then perform the following process:
[0023] If one multiplication variable of the second MUL instruction participates in the multiplication operation in the temporary instruction together with the first variable and the second variable, and the other multiplication variable participates in the multiplication operation in the first MUL instruction together with the third variable and the temporary variable, and the addition variable in the temporary instruction is 1 or the first variable;
[0024] Then change the first MUL instruction to a first FMA instruction, and use the result variable of the second MUL instruction and the second variable as the two multiplication variables of the first FMA instruction;
[0025] When the addition parameter in the temporary instruction is 1, use the third variable as the addition variable of the first FMA instruction; when the addition parameter in the temporary instruction is the first variable, then use the result variable of the second MUL instruction as the addition variable of the first FMA instruction, and delete the temporary instruction.
[0026] In a second aspect, the present invention further provides a GPU compiler optimization method, using the GPU compiler optimization method described in the first aspect for compiler optimization, and further optimizing the instructions in the first window that contain constant parameters.
[0027] Preferably, the optimization of the instructions in the first window that contain constant parameters specifically includes:
[0028] If the temporary instruction in the first window is an FMA instruction, there is a first constant participating in the multiplication operation in the FMA instruction, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first MUL instruction using a temporary variable and a third constant participating in the multiplication operation,
[0029] change the first MUL instruction to a first FMA instruction, use a first value and the multiplication variable in the temporary instruction to participate in the multiplication operation in the first FMA instruction, use a second value to participate in the addition operation in the first FMA instruction, and delete the temporary instruction; wherein, the first value is the product of the first constant and the third constant, and the second value is the product of the second constant and the third constant.
[0030] Preferably, the optimization of the instruction containing constant parameters in the first window further includes:
[0031] If the temporary instruction in the first window is an FMA instruction, there is a first constant and a first variable participating in the multiplication operation together in the FMA instruction, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first ADD instruction using a temporary variable to participate in the addition operation, then perform the following process:
[0032] If in the first ADD instruction, the temporary variable and the third constant participate in the addition operation together, change the first ADD instruction to a first FMA instruction, the first FMA instruction uses the first constant and the first variable in the temporary instruction to participate in the multiplication operation together, uses the first value to participate in the addition operation, and deletes the temporary instruction; wherein, the first value is the sum of the second constant and the third constant;
[0033] If in the first ADD instruction, the temporary variable and the first variable participate in the addition operation together, change the first ADD instruction to a second FMA instruction, the second FMA instruction uses the second value and the first variable in the temporary instruction to participate in the multiplication operation together, uses the second variable in the temporary instruction to participate in the addition operation, and deletes the temporary instruction; wherein, the second value is the value obtained by adding 1 to the first constant.
[0034] Preferably, the optimization of the instruction containing constant parameters in the first window further includes:
[0035] If the temporary instruction in the first window is an FMA instruction, there is a first constant and a first variable participating in the multiplication operation together in the FMA instruction, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first FMA instruction using a temporary variable to participate in the addition operation,
[0036] If in the first FMA instruction, the third constant participates in a multiplication operation together with a temporary variable and the fourth constant participates in an addition operation, then replace the two multiplication parameters in the first FMA instruction with a first value and the first variable respectively, replace the addition parameter in the first FMA instruction with a second value, and delete the temporary instruction; wherein, the first value is the product of a first constant and the third constant, and the second value is the sum obtained by multiplying the second constant by the third constant and then adding the fourth constant.
[0037] If after the temporary instruction, there is a second FMA instruction that uses the temporary variable to participate in a multiplication operation, and in the second FMA instruction, the third constant participates in a multiplication operation together with the first variable and the temporary variable participates in an addition operation, then replace the second constant in the second FMA instruction with a third value, replace the addition parameter in the second FMA instruction with the second constant, and delete the temporary instruction; wherein, the third value is the sum of the first constant and the second constant.
[0038] Preferably, optimizing the instruction containing constant parameters in the first window further includes:
[0039] If in the first FMA instruction of the first window, a temporary variable and a first constant are used to participate in a multiplication operation, a second constant is used to participate in an addition operation, and the temporary command corresponding to the temporary variable is a MUL command or an ADD command,
[0040] If the temporary command is a MUL command, and in the temporary command, a third constant and the first variable are used to participate in a multiplication operation, then replace the two multiplication parameters in the first FMA instruction with a first value and the first variable respectively, and delete the temporary instruction; wherein, the first value is the product of the third constant and the first constant;
[0041] If the temporary command is an ADD command, and in the temporary command, a third constant and the first variable are used to participate in an addition operation, then replace the temporary variable in the first FMA instruction with the first variable, and replace the second constant in the first FMA instruction with a second value, and delete the temporary instruction; wherein, the second value is the sum obtained by multiplying the first constant by the third constant and then adding the second constant.
[0042] Preferably, the template matching optimization is further extended, including:
[0043] Before performing the FMA instruction template matching optimization, check whether the instructions in the basic block contain division operations. If so, make the second operand its reciprocal and change the division to a multiplication operation. Then it can be matched with the defined multiply-add instruction template;
[0044] Further expand the above instruction template, replace the plus sign in the instruction template with a minus sign correspondingly, increase the number of templates, and expand the scope of template matching and optimization;
[0045] After performing the aforementioned template matching optimization, for some processor backends such as CPUs, DSPs, and ASICs that do not support FMA instructions, further convert the optimized FMA instruction into a multiplication instruction and an addition instruction.
[0046] The present invention optimizes the assembly code by setting an optimization template to identify FMA instructions and optimize them according to corresponding optimization rules to delete temporary instructions, thereby removing redundant instructions, reducing the time required for program operation, and reducing the occupation of some redundant resources.
[0047] The above preferred methods provided by the present invention specifically include the following steps in actual implementation:
[0048] 1. Read the program to be compiled, perform lexical, syntactic, and semantic analysis after preprocessing, convert it into an intermediate representation of an abstract syntax tree and perform mid-end optimization, convert it into a back-end intermediate representation, and perform back-end related optimization; in the back-end optimization, traverse all functions in the program, traverse all basic blocks in the function, traverse all instructions in the basic block, and find FMA instructions;
[0049] 2. If the number of instructions in the basic block where the FMA instruction is located is less than or equal to 4, then use all the instructions in the basic block as the template matching window; otherwise, use the FMA instruction, the instruction before the FMA instruction, and the two instructions after the FMA instruction together as the template matching window;
[0050] 3. According to the dependency relationship of each instruction in the template matching window, check whether there is at least one temporary instruction in the template matching window; wherein, the temporary instruction is an instruction not depended on by any instruction outside the template matching window;
[0051] 4. Use the template matching window with at least one temporary instruction as the first window, match the first window with each optimization template, and if the first window successfully matches the corresponding optimization template, then optimize each instruction in the first window according to the optimization rule corresponding to the optimization template; wherein, each optimization template includes at least a temporary instruction and a first instruction, and the first instruction depends on the temporary instruction; the optimization rule corresponding to each optimization template replaces the temporary variable in the first instruction with other parameters and deletes the temporary instruction.
[0052] 5. After the optimization of the FMA instruction is completed, update the instruction list, update the instruction dependency graph and the register list; continue with other back-end optimizations.
[0053] Among them, the goal of FMA instruction optimization is to reduce the number of instructions without affecting the semantics of the program.
[0054] The present invention designs 16 FMA instruction optimization templates and optimization rules:
[0055] For ease of description, the instructions are represented here in the most concise expression form. For example, each of the following expressions can be equivalently represented by one instruction, and eliminating the expression reduces the corresponding instruction.
[0056] w = x × y + z can be represented as fma w, x, y, z
[0057] w = x + y can be represented as add w, x, y
[0058] w = x × y can be represented as mul w, x, y
[0059] In the expressions described in the present invention, x, y, z, k, w, r, d represent variables, temp represents a temporary variable, and A, B, C, D represent constants.
[0060] The instruction optimization templates designed by the present invention are divided into 2 categories, a total of 16:
[0061] The first category is optimization templates based on common sub-expression deletion. The following 7 instruction templates all contain implicit common sub-expressions:
[0062] pattern 1:
[0063] 1) w = x × y + z; 2) temp = x + k; 3) r = y × temp + z
[0064] Among them, r = x × y + z + y × k, so x * y + z is a common sub-expression, and the expression can be optimized to:
[0065] 1) w = x × y + z; 2) r = y × k + w
[0066] pattern 2:
[0067] 1) w = x × y + z; 2) temp = y + z; 3) r = x × temp + z
[0068] Among them, r = x × y + z + x × z, so x × y + z is a common sub-expression, and the expression can be optimized to:
[0069] 1) w = x × y + z; 2) r = x × z + w
[0070] pattern 3:
[0071] 1) w = x × y + z; 2) temp = d × k + z; 3) r = x × y + temp
[0072] Among them, r = x × y + z + d × k, so x × y + z is a common sub-expression, and the expression can be optimized to:
[0073] 1) w = x × y + z; 2) r = d × k + w
[0074] pattern 4:
[0075] 1) w = x × y; 2) temp = x × z + x; 3) r = y × temp
[0076] Among them, r = x × y × z + x × y, so x × y is a common sub-expression, and the expression can be optimized to:
[0077] 1) w = x × y; 2) r = z × w + w
[0078] pattern 5:
[0079] 1) w = x × y; 2) temp = x × z + 1; 3) r = y × temp
[0080] Among them, r = x × y × z + y, so x × y is a common sub-expression, and the expression can be optimized to:
[0081] 1) w = x × y; 2) r = z × w + y
[0082] pattern 6:
[0083] 1) w = x × y; 2) temp = x × x + 1; 3) r = y × temp
[0084] Among them, r = x × y × x + y, so x × y is a common sub-expression, and the expression can be optimized to:
[0085] 1) w = x × y; 2) r = x × w + y
[0086] pattern 7:
[0087] 1) w = x × y; 2) temp = x × x + x; 3) r = y × temp
[0088] Among them, r = x × y × x + x × y, so x × y is a common sub-expression, and the expression can be optimized to:
[0089] 1) w = x × y; 2) r = x × w + w
[0090] The second type is the optimization template based on expression simplification. The following 9 instruction templates all imply expressions that can be simplified by constants.
[0091] pattern 8:
[0092] 1) temp = A × x + B; 2) y = C × temp
[0093] Both of these expressions contain variables, and current compilers will correspondingly generate two instructions. However, A, B, and C are immediate constants, and their values are known at compile time. The values of A × C and B × C can be calculated by the compiler at compile time and directly stored in registers, without the need for the target processor to calculate again. Therefore, the expression can be optimized to one instruction:
[0094] 1) y = AC × x + BC
[0095] pattern 9:
[0096] 1) temp = A × x + B; 2) y = temp + C
[0097] Similarly, the expression can be optimized to:
[0098] 1) y = A × x + (B + C)
[0099] pattern 10:
[0100] 1) temp = A × x + B; 2) y = temp + x
[0101] Similarly, the expression can be optimized to:
[0102] 1) y = (A + 1) × x + B
[0103] pattern 11:
[0104] 1) temp = A × x + B; 2) y = C × temp + D
[0105] Similarly, the expression can be optimized to:
[0106] 1) y = AC × x + (BC + D)
[0107] pattern 12:
[0108] 1) temp = A × x + B; 2) y = C × x + temp
[0109] Similarly, the expression can be optimized to:
[0110] 1) y = (A + C) × x + B
[0111] pattern 13:
[0112] 1) temp = A × x; 2) y = B × temp + C
[0113] Similarly, the expression can be optimized to:
[0114] 1) y = AB × x + C
[0115] pattern 14:
[0116] 1) temp = x + A; 2) y = B × temp + C
[0117] Similarly, the expression can be optimized to:
[0118] 1) y = B × x + (AB + C)
[0119] pattern 15:
[0120] 1) temp = A × x; 2) y = B × x + temp
[0121] Similarly, the expression can be optimized to:
[0122] 1) y = (A + B) × x
[0123] pattern 16:
[0124] 1) temp = x + A; 2) y = B × x + temp
[0125] Similarly, the expression can be optimized to:
[0126] 1) y = (B + 1) × x + A
[0127] If the compiler detects that the instructions in the program match one of the above 16 templates, it will perform FMA instruction optimization, thus reducing one multiply-add, multiplication, or addition instruction.
[0128] It should be noted that the present invention is not limited to the above 16 templates. In fact, according to the optimization idea provided by the present invention, more optimization templates can be easily obtained. Examples are not given here for further elaboration.
[0129] The present invention also provides an extension of the above compilation optimization method, which is not limited to only optimizing the multiply-add instruction sequence and is not limited to the GPU processor. The method includes:
[0130] 1. Before performing FMA instruction template matching optimization, check whether the instructions in the basic block contain a division operation. If so, make the second operand its reciprocal and change the division to a multiplication operation. For example, w = x ÷ y becomes d = 1 ÷ y; w = x × d. Then it can be matched with the defined multiply-add instruction templates.
[0131] 2. Expand the above instruction template, replace the plus sign in the instruction template with a minus sign accordingly, increase the number of templates, and expand the scope of template matching and optimization.
[0132] 3. Perform the aforementioned template matching optimization.
[0133] 4. For some processor backends such as CPUs, DSPs, and ASICs that do not support FMA instructions, further convert the optimized FMA instructions into a multiplication instruction and an addition instruction, and the overall number of instructions is still reduced.
[0134] The present invention designs multiple instruction optimization templates. Through corresponding optimization rules, redundant instructions implicit in the code are deleted, so that the generated assembly instructions are as concise as possible, and the redundant resource occupation is reduced, and the running time of the executable program is reduced. Brief Description of the Drawings
[0135] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required to be used in the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0136] Figure 1 is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0137] Figure 2 is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0138] Figure 3 is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0139] Figure 4 is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0140] Figure 5 is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0141] Figure 6 is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0142] Figure 7 is a schematic diagram of a GPU compiler optimization method provided by an embodiment of the present invention;
[0143] Figure 8It is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention;
[0144] Figure 9 It is a schematic flowchart of a GPU compiler optimization method provided by an embodiment of the present invention. Specific embodiments
[0145] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0146] It should be noted here that the restrictive descriptions such as "first" and "second" in this embodiment do not refer to a specific object, nor do they refer to a specific order meaning. They are only used to distinguish the corresponding restricted objects from the same category, and are added for the convenience of describing two or more different objects in the same category. They should not be interpreted as having a further restrictive meaning. And in this embodiment, since there are multiple optimization templates, and each template involves multiple variables, for the convenience of description, the restrictive descriptions such as "first" and "second" are all relative to a single template. In different templates, the same naming may be used. For example, "the first variable" appears in both the first template and the second template. At this time, it should be understood as different objects and should not be interpreted as the same described object.
[0147] FMA is one of the most common arithmetic operators in CPUs and GPUs. It has been found in actual tests that there is often further room for optimization in this type of instruction sequence.
[0148] For example, when there are three consecutive expressions: 1) w = x × y; 2) temp = x × z + x; 3) r = y × temp, and temp is a temporary variable that is only used in expression 3), the expressions can be optimized to 1) w = x × y 2) r = w × z + w. Among them, x × y is an implicit common sub-expression. However, the existing optimization techniques cannot identify this common sub-expression, and thus will not perform the above optimization. After actual tests, the GCC, LLVM, and JVM compilers still cannot optimize this type of expression when using the highest optimization options. The generated X86 assembly code still contains three multiplication instructions, and the second multiplication instruction corresponding to the temporary instruction is redundant.
[0149] At present, the expressions optimized by mainstream compilers are relatively obvious and easy to replace, and they only optimize explicit common sub-expressions. For example, when the above consecutive expressions are slightly modified to 4) w = x × y; 5) temp = x × y + x; 6) r = y × temp, since there is an explicit common sub-expression x × y in 4) and 5), existing compilation optimization techniques can optimize these three instructions into two instructions: 1) w = x × y; 2) r = w × y + w, and the generated X86 assembly code only contains two multiplication instructions, that is, optimization will only be performed when there are exactly the same sub-expressions in two expressions. This optimization method has certain limitations, and there may still be a large amount of redundancy in the finally generated executable program.
[0150] It should also be emphasized that in this embodiment, descriptions such as "constant", "variable", and "parameter" are involved many times, which are all commonly used terms in the art. Among them, those that participate in the operations in the corresponding instructions are all called parameters, the parameters with determined values are called constants, and the parameters whose values depend on the operation results of a certain operation process, that is, the parameters that may change, are called variables. For example, there is an FMA instruction represented by the intermediate representation as fma w, x, y, 3, which is converted into a mathematical expression as w = x × y + 3. Among them, w, x, y, and 3 are all parameters of this instruction, the values of x and y are calculated by a certain instruction before this instruction, that is, x and y are variables, the value of w depends on the operation result of this instruction, and w is also a variable, while 3 is a constant.
[0151] In different types of instructions, in order to distinguish different variables or parameters, prefixes such as "add", "multiply", and "result" are used to distinguish and describe them. For example, in the ADD instruction, there is only an addition operation, which involves a result parameter and two add parameters. When one of the add parameters is a variable, it can also be described as an add variable. In the FMA instruction, there is a multiplication operation and an addition operation, which involves a result parameter, two multiply parameters, and one add parameter. Similarly, when the multiply parameter is a variable, it can also be called a multiply variable, and when the add parameter is a variable, it can also be called an add variable. Among them, since the result parameters in the instructions are all variables, they are called result variables in the embodiment.
[0152] In addition, the technical features involved in each embodiment of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0153] An embodiment of the present invention provides a GPU compiler optimization method, as Figure 1 shown, including:
[0154] In step 201, traverse all functions in the program, traverse the basic blocks in the function, traverse all instructions in the basic block, and find FMA instructions.
[0155] In step 202, if the number of instructions in the basic block where the FMA instruction is located is less than or equal to a preset value, all the instructions in the basic block are used as the template matching window. In an embodiment of the present invention, the preset value is taken as 4 for example; otherwise, the FMA instruction, the instruction before the FMA instruction, and the two instructions after the FMA instruction are used as the template matching window together.
[0156] In step 203, according to the dependency relationships of the instructions in the template matching window, it is checked whether there is at least one temporary instruction in the template matching window; wherein, the temporary instruction is an instruction that is not relied on by any instruction outside the template matching window.
[0157] In step 204, the template matching window with at least one temporary instruction is used as the first window, and each optimization template is used to match the first window. If the first window successfully matches the corresponding optimization template, then according to the optimization rules corresponding to the optimization template, the instructions in the first window are optimized; wherein, each optimization template includes at least a temporary instruction and a first instruction, and the first instruction depends on the temporary instruction; the optimization rules corresponding to each optimization template replace the temporary variable in the first instruction with other parameters and delete the temporary instruction.
[0158] Replacing the temporary variable in the first instruction with other parameters needs to be determined according to specific optimization rules, which usually replace it with the corresponding parameters in the temporary instruction and other instructions. In some cases, the instruction type of the first instruction may also be changed.
[0159] In subsequent embodiments, the result variable of the temporary instruction is referred to as a temporary variable.
[0160] In this embodiment, by setting optimization templates and corresponding optimization rules, temporary instructions are deleted, thereby removing redundant calculations and optimizing the assembly code. The purpose of the optimization is mainly to reduce the number of instructions without affecting the semantics of the program, so as to reduce the time required for the program to run and reduce the occupation of some redundant resources.
[0161] In actual use, there are various situations for the optimization of the FMA instruction, and each situation may correspond to different optimization rules, so there are multiple optimization templates. As an optional implementation manner, this embodiment provides 16 optional optimization templates here to specifically match and optimize the situations of the FMA instruction optimization. The corresponding optimization rules for each optimization template will be specifically described below.
[0162] When using the optimized template one and the optimized template two for matching, each optimized template is used to match the first window. If the first window successfully matches the corresponding optimized template, then according to the optimization rules corresponding to the optimized template, each instruction in the first window is optimized, as Figure 2 shown, specifically including:
[0163] In step 301, if there are a temporary instruction and a second FMA instruction before the first FMA instruction in the first window, and the temporary instruction is an ADD instruction, then the following process is executed:
[0164] In step 302, if one of the multiplication variables in the second FMA instruction participates in the addition operation in the temporary instruction as the first variable, and the other multiplication variable participates in the multiplication operation in the first FMA instruction together with the temporary variable, and the addition variable in the first FMA instruction is the same as the addition variable in the second FMA instruction.
[0165] In step 303, then replace the temporary variable in the first FMA instruction with the second variable, replace the addition variable in the first FMA instruction with the result variable of the second FMA instruction, and delete the temporary instruction; where the temporary variable is the result variable of the temporary instruction, and the second variable is the other addition variable in the temporary instruction (that is, the two addition variables in the temporary instruction are the first variable and the second variable respectively).
[0166] The optimized template one can be expressed in the form of a mathematical formula as:
[0167] Second FMA instruction: w = x × y + z.
[0168] Temporary instruction: temp = x + k.
[0169] First FMA instruction: r = y × temp + z.
[0170] Where x is the first variable and k is the second variable. Intuitively, there is no explicit common sub-expression in the first FMA instruction and the second FMA instruction. However, by substituting and simplifying the temporary variable temp in the first FMA instruction using the temporary instruction, we get: r = y × (x + k) + z = (x × y + z) + y × k. It can be seen that there is an implicit common sub-expression x × y + z in the first FMA instruction and the second FMA instruction, and the prior art cannot optimize this implicit common sub-expression.
[0171] In this embodiment, by replacing the temporary variable temp in the first FMA instruction with the second variable k and replacing the addition variable z in the first FMA instruction with the result variable w of the second FMA instruction, the first FMA instruction does not need to depend on the temporary instruction, and then the temporary instruction is deleted, and the above three instructions are optimized into the following two FMA instructions:
[0172] Second FMA instruction: w = x × y + z.
[0173] First FMA instruction: r = y × k + w.
[0174] It should be noted that each of the above mathematical formulas represents the corresponding intermediate representation instruction. For example, the instruction represented by r = y × k + w is fma r, y, k, w. Here, it is only for the convenience of intuitively showing the operation relationship in each instruction that the mathematical formula is used for representation. In actual use, the actual matching and optimization of each instruction are carried out.
[0175] There are also template two derived from optimization template one and the corresponding optimization rules:
[0176] Second FMA instruction: w = x × y + z.
[0177] Temporary instruction: temp = y + z.
[0178] First FMA instruction: r = x × temp + z.
[0179] Its final optimization result is:
[0180] Second FMA instruction: w = x × y + z.
[0181] First FMA instruction: r = x × z + w.
[0182] When using optimization template three for matching, each optimization template is used to match the first window. If the first window matches the corresponding optimization template successfully, then according to the optimization rules corresponding to the optimization template, each instruction in the first window is optimized. As Figure 3 shown, it further includes:
[0183] In step 401, if there are a temporary instruction and a second FMA instruction before the first FMA instruction in the first window, and the temporary instruction is an FMA instruction, then the following process is executed:
[0184] In step 402, if the two multiplication variables in the second FMA instruction are the same as the two multiplication variables in the first FMA instruction, the addition variable in the first FMA instruction is a temporary variable, and the addition variable in the temporary variable is the same as the addition variable in the second FMA instruction.
[0185] In step 403, replace the two multiplication variables in the first FMA instruction with the two multiplication variables in the temporary instruction, replace the addition variable in the first FMA instruction with the result variable of the second FMA instruction, and delete the temporary instruction; wherein, the temporary variable is the result variable of the temporary instruction.
[0186] The optimization template three can be expressed in the form of a mathematical formula as:
[0187] Second FMA instruction: w = x × y + z.
[0188] Temporary instruction: temp = k1 × k2 + z.
[0189] First FMA instruction: r = x × y + temp.
[0190] Among them, intuitively, there is no explicit common sub-expression in the first FMA instruction and the second FMA instruction. Substituting and simplifying the temporary variable temp in the first FMA instruction using the temporary instruction gives: r = x × y + (k1 × k2 + z) = x × y + z + k1 × k2. It can be seen that there is an implicit common sub-expression x × y + z in the first FMA instruction and the second FMA instruction, and the prior art cannot optimize this implicit common sub-expression.
[0191] In this embodiment, by replacing the two multiplication variables (i.e., x and y) in the first FMA instruction with the two multiplication variables (i.e., k1 and k2) in the temporary instruction, and replacing the addition variable temp in the first FMA instruction with the result variable w of the second FMA instruction, the first FMA instruction does not need to depend on the temporary instruction, and then the temporary instruction is deleted. The above three instructions are optimized into the following two FMA instructions:
[0192] Second FMA instruction: w = x × y + z.
[0193] First FMA instruction: r = k1 × k2 + w.
[0194] When using template four, template five, template six, and template seven for matching, use each optimization template to match the first window. If the first window matches the corresponding optimization template successfully, then optimize each instruction in the first window according to the optimization rules corresponding to the optimization template, as Figure 4 shown, further including:
[0195] In step 501, if there is a temporary instruction and a second MUL instruction before the first MUL instruction in the first window, and the temporary instruction is an FMA instruction, then perform the following process:
[0196] In step 502, if one multiplication variable of the second MUL instruction participates in the multiplication operation in the temporary instruction together with the first variable and the second variable, and the other multiplication variable participates in the multiplication operation in the first MUL instruction together with the third variable and the temporary variable, and the addition variable in the temporary instruction is 1 or the first variable.
[0197] In step 503, the first MUL instruction is changed to a first FMA instruction, and the result variable of the second MUL instruction and the second variable are used as the two multiplication variables of the first FMA instruction.
[0198] In step 504, when the addition parameter in the temporary instruction is 1, the third variable is used as the addition variable of the first FMA instruction; when the addition parameter in the temporary instruction is the first variable, the result variable of the second MUL instruction is used as the addition variable of the first FMA instruction, and the temporary instruction is deleted.
[0199] The optimization template four can be expressed in the form of a mathematical formula as follows:
[0200] Second MUL instruction: w = x × y.
[0201] Temporary instruction: temp = x × z + x.
[0202] First MUL instruction: r = y × temp.
[0203] Where x is the first variable, z is the second variable, and y is the third variable. Intuitively, there are no explicit common sub-expressions in the first FMA instruction and the second FMA instruction. Substituting and simplifying the temporary variable temp in the first FMA instruction using the temporary instruction gives: r = y × (x × z + x) = (x × y) × z + x × y. It can be seen that there is an implicit common sub-expression x × y in the first FMA instruction and the second FMA instruction, and the prior art cannot optimize this implicit common sub-expression.
[0204] In this embodiment, by changing the first MUL instruction to a first FMA instruction, using the result variable w of the second MUL instruction and the second variable z as the two multiplication variables of the first FMA instruction, and since the addition parameter in the temporary instruction is the first variable x, according to the optimization rule, the result variable w of the second MUL instruction and the second variable z are used as the two multiplication variables of the first FMA instruction, and the result variable w of the second MUL instruction is used as the addition variable of the first FMA instruction, and the temporary instruction is deleted, optimizing the above three instructions into the following two FMA instructions:
[0205] Second FMA instruction: w = x × y.
[0206] First FMA instruction: r = z × w + w.
[0207] When the add parameter in the temporary instruction is 1, there is an optimization template five, which can be expressed in the form of a mathematical formula as:
[0208] Second MUL instruction: w = x × y.
[0209] Temporary instruction: temp = x × z + 1.
[0210] First MUL instruction: r = y × temp.
[0211] Where x is the first variable, z is the second variable, and y is the third variable. Intuitively, there is no explicit common sub-expression in the first FMA instruction and the second FMA instruction. By substituting and simplifying the temporary variable temp in the first FMA instruction using the temporary instruction, we get: r = y × (x × z + 1) = (x × y) × z + y. It can be seen that there is an implicit common sub-expression x × y in the first FMA instruction and the second FMA instruction, and the prior art cannot optimize this implicit common sub-expression.
[0212] In this embodiment, by changing the first MUL instruction to the first FMA instruction, using the result variable w of the second MUL instruction and the second variable z as the two multiplication variables of the first FMA instruction, and since the add parameter in the temporary instruction is 1, according to the optimization rule, using the result variable w of the second MUL instruction and the second variable z as the two multiplication variables of the first FMA instruction, using the third variable y of the second MUL instruction as the addition variable of the first FMA instruction, and deleting the temporary instruction, the above three instructions are optimized into the following two FMA instructions:
[0213] Second FMA instruction: w = x × y.
[0214] First FMA instruction: r = z × w + y.
[0215] Based on optimization template five, optimization template six and the corresponding optimization rules are derived as:
[0216] Second MUL instruction: w = x × y.
[0217] Temporary instruction: temp = x × x + 1.
[0218] First MUL instruction: r = y × temp.
[0219] The final optimization result is:
[0220] Second FMA instruction: w = x × y.
[0221] First FMA instruction: r = x × w + y.
[0222] Based on Optimization Template Four, derivative Optimization Template Seven and the corresponding optimization rules are obtained as follows:
[0223] Second MUL instruction: w = x × y.
[0224] Temporary instruction: temp = x × x + x.
[0225] First MUL instruction: r = y × temp.
[0226] The final optimization result is:
[0227] Second FMA instruction: w = x × y.
[0228] First FMA instruction: r = x × w + w.
[0229] In actual use, there is also a type of SUB instruction in the operation of instructions. When the SUB instruction is related to the FMA instruction, there may also be instruction redundancy. For example, there are the following three instructions:
[0230] Second FMA instruction: w = x × y + z.
[0231] Temporary instruction: temp = y - k.
[0232] First FMA instruction: r = x × temp + z.
[0233] Among them, the temporary instruction is a SUB instruction. Substituting and simplifying the temporary variable temp in the first FMA instruction with the temporary instruction, we get: r = x × (y - k) + z = (x × y + z) - k × x. It can be seen that the first FMA instruction and the second FMA instruction contain an implicit common sub-expression x × y + z, that is, there is instruction redundancy. To solve this problem, the present embodiment also provides the following preferred implementation manners, which specifically include:
[0234] When using each optimization template to match the first window, if the first window matches the corresponding optimization template successfully, then according to the optimization rules corresponding to the optimization template, optimize each instruction in the first window, as Figure 5 shown, it also includes:
[0235] In step 601, when a SUB command appears in the first window, convert the SUB command into an ADD command, and let the ADD command participate in the optimization template matching. The specific method of converting the SUB command into an ADD command is to convert the subtraction parameter in the SUB command into the corresponding negative value of the addition parameter, so as to convert the subtraction operation into an addition operation, thereby obtaining the ADD command.
[0236] In step 602, if the first window matches the corresponding optimization template successfully, then according to the optimization rules corresponding to the optimization template, optimize each instruction in the first window.
[0237] For example, for the above-mentioned temporary instruction temp = y - k, it can be regarded as the addition operation of two parameters y and -k, and thus converted into an ADD instruction to participate in the optimization template matching. If the above three instructions match template two, then optimize according to the corresponding optimization rules. Finally, the following is obtained:
[0238] Second FMA instruction: w = x × y + z.
[0239] First FMA instruction: r = -k × x + w.
[0240] Thus, redundant temporary instructions are removed.
[0241] In actual use, the compiler needs to perform some conversions on the program to be compiled before template matching optimization. Incorporating the above embodiments, the following complete implementation steps are provided:
[0242] 1) Read the program to be compiled, perform lexical, syntactic, and semantic analysis after preprocessing, convert it into an intermediate representation of an abstract syntax tree and perform preliminary optimization, convert it into a back-end intermediate representation, and perform back-end related optimization.
[0243] 2) In the back-end optimization, traverse all functions in the program, traverse all basic blocks in the function, traverse all FMA instructions in the basic block, and check whether the FMA instruction and the instructions before and after it match the defined instruction templates (i.e., each optimization template).
[0244] 3) If it matches the instruction template, replace the several matched instructions with the instructions after optimization of the instruction template according to the corresponding optimization rules; otherwise, continue to traverse the next instruction.
[0245] 4) After the FMA instruction optimization is completed, update the instruction list, update the instruction dependency relationship and the register list.
[0246] 5) Continue with other back-end optimizations, and finally generate assembly code and encode it into an executable file for output.
[0247] In actual use, not only may common subexpressions lead to redundant instructions, but when there are constants involved in temporary instructions or instructions dependent on temporary instructions, there may also be problems of redundant instructions. To solve this problem, an embodiment of the present invention also provides a GPU compiler optimization method.
[0248] While using the GPU compiler optimization method described in the above embodiments for compiler optimization, the method also optimizes the instructions containing constant parameters in the first window.
[0249] The optimization of the instructions containing constant parameters in the first window specifically includes:
[0250] Using the template matching window with at least one temporary instruction as the first window, and using each simplification template to match the first window. If the first window successfully matches the corresponding simplification template, then according to the simplification rules corresponding to the simplification template, optimize each instruction in the first window.
[0251] This embodiment also provides 9 simplification templates for matching and simplification. The following will specifically elaborate on these simplification templates and the corresponding simplification rules.
[0252] When using the first simplification template for matching, the optimization of the instructions containing constant parameters in the first window is as Figure 6 shown, and specifically includes:
[0253] In step 701, if the temporary instruction in the first window is an FMA instruction, there is a first constant participating in the multiplication operation in the FMA instruction, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first MUL instruction using a temporary variable and a third constant participating in the multiplication operation:
[0254] In step 702, change the first MUL instruction to a first FMA instruction, let the first value participate in the multiplication operation in the first FMA instruction together with the multiplication variable in the temporary instruction, let the second value participate in the addition operation in the first FMA instruction, and delete the temporary instruction; where the first value is the product of the first constant and the third constant, and the second value is the product of the second constant and the third constant.
[0255] The first simplification template can be expressed in the form of a mathematical formula as:
[0256] Temporary instruction: temp = A × x + B.
[0257] First MUL instruction: y = C × temp.
[0258] Where A is the first constant, B is the second constant, C is the third constant. Substitute the temporary variable temp in the first FMA instruction with the temporary instruction and simplify to get: y = C × (A × x + B) = A × C × x + B × C. Since A, B, and C are all constants, the values of A × C and B × C can be calculated in advance by the compiler at compile time and directly store the operation results in the register, so it can be optimized.
[0259] Change the first MUL instruction to the first FMA instruction. Let the first value (i.e., A×C) participate in the multiplication operation in the first FMA instruction together with the multiplication variable in the temporary instruction, let the second value (i.e., B×C) participate in the addition operation in the first FMA instruction, and delete the temporary instruction. Optimize the above two instructions into one FMA instruction y = AC×x + BC.
[0260] When using Simplification Template Two and Simplification Template Three for matching, the optimization of the instructions containing constant parameters in the first window further includes:
[0261] If the temporary instruction in the first window is an FMA instruction, in the FMA instruction, there is a first constant and a first variable participating in the multiplication operation together, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first ADD instruction using a temporary variable to participate in the addition operation, then the following process is executed:
[0262] If in the first ADD instruction, the temporary variable and a third constant participate in the addition operation together, then change the first ADD instruction to the first FMA instruction. The first FMA instruction uses the first constant and the first variable in the temporary instruction to participate in the multiplication operation together, uses the first value to participate in the addition operation, and deletes the temporary instruction; where the first value is the sum of the second constant and the third constant.
[0263] If in the first ADD instruction, the temporary variable and the first variable participate in the addition operation together, then change the first ADD instruction to the second FMA instruction. The second FMA instruction uses the second value and the first variable in the temporary instruction to participate in the multiplication operation together, uses the second variable in the temporary instruction to participate in the addition operation, and deletes the temporary instruction; where the second value is the value obtained by adding 1 to the first constant.
[0264] The Simplification Template Two can be expressed in the form of a mathematical formula as:
[0265] Temporary instruction: temp = A×x + B.
[0266] First ADD instruction: y = temp + C.
[0267] Where A is the first constant, B is the second constant, C is the third constant, and x is the first variable. Substitute and simplify the temporary variable temp in the first ADD instruction using the temporary instruction to get: y = (A×x + B) + C = A×x + (B + C). Since both B and C are constants, the value of B + C can be calculated in advance by the compiler at compile time, so it can be optimized.
[0268] Change the first ADD instruction to the first FMA instruction. Use the first constant A in the temporary instruction and the first variable x to participate in the multiplication operation. Use the first value (i.e., B + C) to participate in the addition operation and delete the temporary instruction. Optimize the above two instructions into one first FMA instruction y = A × x + (B + C).
[0269] The simplification template three can be expressed in the form of a mathematical formula as follows:
[0270] Temporary instruction: temp = A × x + B.
[0271] First ADD instruction: y = temp + x.
[0272] Among them, A is the first constant, B is the second constant, and x is the first variable. Substitute and simplify the temporary variable temp in the first ADD instruction using the temporary instruction to get: y = (A × x + B) + x = (A + 1) × x + B. Since A is a constant, the value of A + 1 can be calculated in advance by the compiler at compile time, so it can be optimized.
[0273] Change the first ADD instruction to the second FMA instruction. Use the second value (i.e., A + 1) and the first variable x in the temporary instruction to participate in the multiplication operation. Use the second variable B in the temporary instruction to participate in the addition operation, and delete the temporary instruction. Optimize the above two instructions into one second FMA instruction y = (A + 1) × x + B.
[0274] Among them, when using the simplification template four and the simplification template five for matching, optimizing the instructions containing constant parameters in the first window also includes:
[0275] If the temporary instruction in the first window is an FMA instruction, there is a first constant and a first variable participating in the multiplication operation in the FMA instruction, there is a second constant participating in the addition operation, and there is a first FMA instruction using the temporary variable to participate in the addition operation after the temporary instruction:
[0276] If in the first FMA instruction, a third constant and the temporary variable participate in the multiplication operation, and a fourth constant participates in the addition operation, then replace the two multiplication parameters in the first FMA instruction with the first value and the first variable respectively, replace the addition parameter in the first FMA instruction with the second value, and delete the temporary instruction; among them, the first value is the product of the first constant and the third constant, and the second value is the sum obtained by multiplying the second constant by the third constant and then adding the fourth constant.
[0277] If, after the temporary instruction, there is a second FMA instruction that uses a temporary variable in a multiplication operation, and in the second FMA instruction, a third constant and a first variable participate in the multiplication operation while the temporary variable participates in the addition operation, then replace the second constant in the second FMA instruction with a third value, replace the addition parameter in the second FMA instruction with the second constant, and delete the temporary instruction; where the third value is the sum of the first constant and the second constant.
[0278] The simplification template four can be expressed in the form of a mathematical formula as:
[0279] Temporary instruction: temp = A × x + B.
[0280] First FMA instruction: y = C × temp + D.
[0281] Where A is the first constant, B is the second constant, C is the third constant, D is the third constant, and x is the first variable. Substitute and simplify the temporary variable temp in the first FMA instruction using the temporary instruction to get: y = C × (A × x + B) + D = (A × C) × x + (B × C + D). Since A, B, C, and D are all constants, the values of A × C and B × C + D can be calculated in advance by the compiler at compile time, so it can be optimized.
[0282] Replace the two multiplication parameters in the first FMA instruction with the first value and the first variable respectively, replace the addition parameter in the first FMA instruction with the second value, and delete the temporary instruction. Optimize the above two instructions into one first FMA instruction: y = AC × x + (BC + D).
[0283] The simplification template five can be expressed in the form of a mathematical formula as:
[0284] Temporary instruction: temp = A × x + B.
[0285] First FMA instruction: y = C × x + temp.
[0286] Where A is the first constant, B is the second constant, C is the third constant, and x is the first variable. Substitute and simplify the temporary variable temp in the first FMA instruction using the temporary instruction to get: y = C × x + (A × x + B) = (A + C) × x + B. Since A, B, C, and D are all constants, the value of A + C can be calculated in advance by the compiler at compile time, so it can be optimized.
[0287] Replace the two multiplication parameters in the first FMA instruction with a first value and the first variable respectively, replace the addition parameter in the first FMA instruction with a second value, and delete the temporary instruction. Optimize the above two instructions into a first FMA instruction: y = (A + C) × x + B.
[0288] When using Simplification Template Six and Simplification Template Seven for matching, the optimization of the instruction containing constant parameters in the first window further includes:
[0289] If in the first FMA instruction of the first window, a temporary variable and a first constant are used for multiplication operation, a second constant is used for addition operation, and the temporary command corresponding to the temporary variable is a MUL command or an ADD command:
[0290] If the temporary command is a MUL command, and in the temporary command, a third constant and a first variable are used for multiplication operation, then replace the two multiplication parameters in the first FMA instruction with a first value and the first variable respectively, and delete the temporary instruction; wherein, the first value is the product of the third constant and the first constant.
[0291] If the temporary command is an ADD command, and in the temporary command, a third constant and a first variable are used for addition operation, then replace the temporary variable in the first FMA instruction with the first variable, and replace the second constant in the first FMA instruction with a second value, and delete the temporary instruction; wherein, the second value is the sum of the product of the first constant and the third constant plus the second constant.
[0292] The Simplification Template Six can be expressed in the form of a mathematical formula as:
[0293] Temporary instruction: temp = A × x.
[0294] First FMA instruction: y = temp × B + C.
[0295] Wherein, B is the first constant, C is the second constant, A is the third constant, x is the first variable. Substitute and simplify the temporary variable temp in the first FMA instruction using the temporary instruction to get: y = (A × x) × B + C = (A × B) × x + C. Since both A and B are constants, the value of A × B can be calculated in advance by the compiler at compile time, so it can be optimized.
[0296] Replace the two multiplication parameters in the first FMA instruction with a first value (i.e., A × B) and the first variable respectively, and optimize the first FMA instruction to: y = AB × x + C.
[0297] Furthermore, delete the temporary instruction, and optimize the above two instructions into one FMA instruction to remove the redundant calculations of the instructions, which can be expressed in the form of a mathematical formula as follows:
[0298] The first FMA instruction: y = AB × x + C.
[0299] The simplification template seven can be expressed in the form of a mathematical formula as follows:
[0300] Temporary instruction: temp = x + A.
[0301] The first FMA instruction: y = B × temp + C.
[0302] Wherein, B is the first constant, C is the second constant, A is the third constant, and x is the first variable. Substitute and simplify the temporary variable temp in the first FMA instruction using the temporary instruction to obtain: y = B × (x + A) + C = B × x + (A × B + C). Since A, B, and C are all constants, the value of A × B + C can be calculated in advance by the compiler at compile time, so it can be optimized.
[0303] Replace the two multiplication parameters in the first FMA instruction with the first value (i.e., A × B + C) and the first variable x respectively, and delete the temporary instruction. Optimize the above two instructions into one first FMA instruction y = B × x + (AB + C).
[0304] When using the simplification template eight and the simplification template nine for matching, the optimization of the instruction containing constant parameters in the first window further includes:
[0305] If in the first FMA instruction of the first window, the first constant and the first variable are used for multiplication operations, the temporary variable is used for addition operations, and the temporary command corresponding to the temporary variable is a MUL command or an ADD command:
[0306] If the temporary command is a MUL command, and in the temporary command, the second constant and the first variable are used for multiplication operations, then replace the first FMA instruction with the first MUL instruction, use the first variable and the first value as the multiplication parameters in the first MUL instruction respectively, and delete the temporary instruction; wherein, the first value is the sum of the first constant and the second constant.
[0307] When the temporary command is an ADD command, and in the temporary command, the second constant and the first variable are used for addition operations, then replace the first constant in the first FMA instruction with the first value, and replace the temporary variable in the first FMA instruction with the second constant; wherein, the first value is the value obtained by adding 1 to the first constant, and delete the temporary instruction.
[0308] The simplified template eight can be expressed in the form of a mathematical formula as follows:
[0309] Temporary instruction: temp = A × x.
[0310] First FMA instruction: y = B × x + temp.
[0311] Where B is the first constant, A is the second constant, and x is the first variable. Substituting and simplifying the temporary variable temp in the first FMA instruction using the temporary instruction gives: y = B × x + (A × x) = (A + B) × x. Since both A and B are constants, the value of A + B can be calculated in advance by the compiler at compile time, so it can be optimized.
[0312] Replace the first FMA instruction with a first MUL instruction, use the first variable B and the first value (i.e., A + B) as the multiplication parameters in the first MUL instruction respectively, and delete the temporary instruction. Optimize the above two instructions into one first MUL instruction y = (A + B) × x.
[0313] The simplified template nine can be expressed in the form of a mathematical formula as follows:
[0314] Temporary instruction: temp = x + A.
[0315] First FMA instruction: y = B × x + temp.
[0316] Where B is the first constant, A is the second constant, and x is the first variable. Substituting and simplifying the temporary variable temp in the first FMA instruction using the temporary instruction gives: y = B × x + (x + A) = (B + 1) × x + A. Since B is a constant, the value of B + 1 can be calculated in advance by the compiler at compile time, so it can be optimized.
[0317] Replace the first constant B in the first FMA instruction with the first value (i.e., B + 1), replace the temporary variable in the first FMA instruction with the second constant A, optimize the first FMA instruction to: y = (B + 1) × x + A, and delete the temporary instruction. Optimize the above two instructions into one first MUL instruction: y = (B + 1) × x + A.
[0318] As Figure 7 shown, the program to be compiled generates an intermediate representation in the middle end after preprocessing, lexical analysis, syntax analysis, and semantic analysis by the compiler front end. The intermediate representation undergoes control flow analysis, data flow analysis, dependency analysis, middle-end optimization, and intermediate representation conversion in the middle end to generate a back-end intermediate representation. The back-end intermediate representation undergoes back-end optimization, instruction selection, instruction scheduling, register allocation, code generation, and instruction encoding to generate the target file. The target file can be executed on the GPU hardware.
[0319] The compilation optimization method of the present invention is an optimization process in the backend optimization of the above compilation process, which is called a pass in the compiler. After being optimized by the FMA instruction, the backend intermediate representation continues to be transmitted to the next optimization pass.
[0320] The following Figure 8 and embodiments are used to make a detailed description of the specific implementation method of the FMA instruction optimization of the present invention, which specifically includes:
[0321] In step 801, traverse all functions in the program; enter step 802, and when all functions in the program are traversed, enter step 811.
[0322] In step 802, delete unreachable basic blocks in each function, merge consecutive basic blocks without branch jumps; update the instruction dependency graph; enter step 803.
[0323] In step 803, traverse all basic blocks in each function and enter step 804.
[0324] In step 804, determine whether all basic blocks in the corresponding function have been traversed; each time a basic block is accessed, enter step 805; if all basic blocks in the function have been traversed, enter step 801 to continue accessing the next function.
[0325] In step 805, traverse all FMA instructions in each basic block; specifically: check whether there is an FMA-type instruction in the corresponding basic block, if not, continue to traverse the next basic block; if there is an FMA instruction, enter step 806.
[0326] In step 806, check whether one of the source registers in the FMA instruction is of variable type (that is, determine whether there is a variable in the FMA instruction. If there is no variable, there is no dependency on other instructions and no optimization is required). If not, enter step 805 to continue traversing the next FMA instruction; otherwise, enter step 807.
[0327] In step 807, determine whether the number of instructions in the current basic block is less than or equal to 4. If it is less than or equal to 4, enter step 808, otherwise, enter step 809.
[0328] In step 808, if the number of instructions in the current basic block is less than or equal to 4, take out all the instructions as the template matching window; enter step 810.
[0329] In step 809, otherwise, take out the FMA instruction, the instruction before the FMA instruction, and the two instructions after the FMA instruction. These 4 instructions are used as the template matching window; enter step 810.
[0330] In step 810, template matching is performed on the template matching window for code optimization. After that, in step 805, the next FMA instruction is traversed.
[0331] In step 811, after all instructions have been traversed, if the FMA instruction has been optimized, the instruction list is updated, the dependency graph is updated, and the register list is updated.
[0332] In the said step 810, the template matching on the template matching window for code optimization is as Figure 9 shown, and specifically includes:
[0333] In step 901, the dependency graphs of these fetched instructions are checked respectively. If an instruction is not truly dependent on other instructions outside the current window (i.e., the target register is not read later), then the instruction is a temporary instruction, and the attribute of the target register is marked as temp.
[0334] In step 902, it is checked whether there is a target register with the temp attribute in the instruction window. If not, no code optimization is performed. Otherwise, in step 903, template matching is performed.
[0335] In step 903, the instruction types, source registers, and target registers of these instructions are respectively matched in sequence with 16 set instruction templates (including 7 optimization templates in Embodiment 1 and 9 simplification templates in Embodiment 2).
[0336] In step 904, if 3 of the instructions match one of the optimization templates described in Embodiment 1, the original first instruction is retained, the original second instruction is deleted, and the instruction type and source register of the original third instruction are modified according to the aforementioned template rules.
[0337] In step 905, if 2 of the instructions match one of the simplification templates in Embodiment 2, the original first instruction is deleted, the value of the source register in the second instruction is calculated according to the corresponding optimization rules and simplification rules of the aforementioned templates, and the instruction type and source register are modified.
[0338] Among them, the above steps 901 to 905 are the core of FMA instruction optimization. The process of instruction template matching will be described in detail below in combination with a part of the code in a specific embodiment.
[0339] For the following program:
[0340] in float3 a
[10000] ;
[0341] out float2 b
[10000] ;
[0342] int i;
[0343] float temp;
[0344] for(i = 0; i < 10000; i++)
[0345] {
[0346] b[i].x = a[i].x * a[i].y;
[0347] temp = a[i].x * a[i].z + a[i].x;
[0348] b[i].y = temp * a[i].y;
[0349] }
[0350] The expressions in the loop generate the following intermediate representation instructions, which are in a basic block:
[0351] mul r19,r1,r0;
[0352] fma r21,r2,r0,r0;
[0353] mul r22,r21,r1
[0354] According to step 808 or step 809, take out these instructions as the window for template matching.
[0355] According to step 901, check the dependency graph of these three instructions. Among them, the first and the third instructions are dependent on instructions in other basic blocks (because variable b will be output), that is, r19 and r22 will be read later. r21 will not be read by subsequent instructions, mark its attribute as temp.
[0356] According to step 902, it is checked that the instruction window contains the target register with the attribute of temp.
[0357] According to step 903, first check that the instruction containing temp is {fma temp,r2,r0,r0}, which matches the template instructions containing temp in pattern3 and pattern 4. Then check that the instruction before temp is {mul r19,r1,r0}, which matches the template instruction before temp in pattern 4. Finally, check that the instruction after temp is {mulr22,temp,r1}, which matches the template instruction after temp in pattern 4. Output the mark that the instruction window matches pattern4.
[0358] According to step 904 , the first instruction remains unchanged, the second instruction is deleted, and the third instruction is modified to {fma r22, r19, r2, r19}.
[0359] For step 905, the instruction matches templates 8 to 16. The compiler calculates the simplified constant coefficients and puts them into registers before performing instruction replacement. The process is similar to the above and will not be described in detail here.
[0360] As you can see, after the above program is optimized with FMA instructions, the instructions in the loop become:
[0361] mul r19,r1,r0;
[0362] fma r22,r19,r2,r19
[0363] One multiplication instruction is reduced in the core loop. Actual tests show that the execution time of the optimized code is reduced by about 33%.
[0364] This optimization method is particularly suitable for GPU applications that require intensive rendering, such as 3D games, film and television animation, 4K video encoding and decoding, engineering modeling and simulation, etc. Their underlying code contains a large number of float-type FMA operations.
[0365] An embodiment of the present invention further provides an extension of the above compiler optimization method, so that it is not limited to the optimization of multiply-add instruction sequences or GPU processors. The extension method includes:
[0366] 1. In step 805 of the above embodiment, before performing the FMA instruction template matching optimization, the instruction in the basic block is checked to see if it contains a division operation. If so, the second operand is converted to its reciprocal, converting the division into a multiplication operation. For example, w = x ÷ y becomes r = 1 ÷ y; w = x × r. This can then be matched against the defined multiply-add instruction template.
[0367] 2. Continue to expand the above instruction templates, replace the plus signs in the instruction templates with minus signs accordingly, increase the number of templates, expand the scope of template matching and optimization, and reduce related redundant instructions as much as possible. The above embodiment lists an example of subtraction.
[0368] 3. Perform the aforementioned template matching optimization.
[0369] 4. For some CPUs, DSPs, ASICs and other processors that do not support FMA instructions, the optimized FMA instructions are further converted into a multiplication instruction and an addition instruction, and the overall number of instructions is still reduced.
[0370] 5. Continue with other backend optimizations.
[0371] It should be understood that the specific embodiments described herein are only for explaining the present invention and are not limiting to the present invention. For example, those of ordinary skill in the art can easily further expand the optimized content in the instruction template according to the method of the present invention, perform more similar optimizations to expand the applicable scope of instruction optimization, and further improve the execution efficiency of various types of actual applications.
[0372] It should be noted that the compilation optimization process of the FMA instruction optimization of the present invention is a process of optimizing an intermediate representation and then outputting it again during the compilation process. It can be used as a general module and added to any compatible compiler, and is not limited to a specific processor, having very wide applicability. Intuitively and from actual tests, the method in the present invention can significantly improve the optimization effect of the compiler.
[0373] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A GPU compiler optimization method, characterized in that, Including: Traverse the basic blocks of each function in the program, traverse the instructions in the basic blocks, and find the FMA instructions; If the number of instructions in the basic block where the FMA instruction is located is less than or equal to the preset value, then use all the instructions in the basic block as the template matching window; otherwise, use the FMA instruction, the instruction before the FMA instruction, and the two instructions after the FMA instruction together as the template matching window; According to the dependency relationships of the instructions in the template matching window, check whether there is at least one temporary instruction in the template matching window; wherein, the temporary instruction is an instruction not depended on by any instruction outside the template matching window; Use the template matching window with at least one temporary instruction as the first window, and use each optimization template to match the first window. If the first window matches the corresponding optimization template successfully, then optimize the instructions in the first window according to the optimization rules corresponding to the optimization template; Wherein, each optimization template includes at least a temporary instruction and a first instruction, and the first instruction depends on the temporary instruction; the optimization rules corresponding to each optimization template replace the temporary variable in the first instruction with other parameters and delete the temporary instruction.
2. The GPU compiler optimization method according to claim 1, wherein The step of using each optimization template to match the first window, if the first window matches the corresponding optimization template successfully, then optimizing the instructions in the first window according to the optimization rules corresponding to the optimization template includes: If there is a temporary instruction and a second FMA instruction before the first FMA instruction in the first window, and the temporary instruction is an ADD instruction, then perform the following process: If one of the multiplication variables in the second FMA instruction participates in the addition operation in the temporary instruction as the first variable, the other multiplication variable and the temporary variable participate in the multiplication operation in the first FMA instruction together, and the addition variable in the first FMA instruction is the same as the addition variable in the second FMA instruction; Then replace the temporary variable in the first FMA instruction with the second variable, replace the addition variable in the first FMA instruction with the result variable of the second FMA instruction, and delete the temporary instruction; wherein, the temporary variable is the result variable of the temporary instruction, and the second variable is the other addition variable in the temporary instruction.
3. The GPU compiler optimization method according to claim 1, characterized in that The step of using each optimization template to match the first window, if the first window matches the corresponding optimization template successfully, then optimizing the instructions in the first window according to the optimization rules corresponding to the optimization template includes: If there is a temporary instruction and a second FMA instruction before the first FMA instruction in the first window, and the temporary instruction is an FMA instruction, then perform the following process: If the two multiplication variables in the second FMA instruction are the same as the two multiplication variables in the first FMA instruction, the addition variable in the first FMA instruction is a temporary variable, and the addition variable in the temporary variable is the same as the addition variable in the second FMA instruction; Then replace the two multiplication variables in the first FMA instruction with the two multiplication variables in the temporary instruction, replace the addition variable in the first FMA instruction with the result variable of the second FMA instruction, and delete the temporary instruction; wherein, the temporary variable is the result variable of the temporary instruction.
4. The GPU compiler optimization method according to claim 1, wherein When using each optimization template to match the first window, if the first window successfully matches the corresponding optimization template, then optimizing each instruction in the first window according to the optimization rule corresponding to the optimization template includes: If there are a temporary instruction and a second MUL instruction before the first MUL instruction in the first window, and the temporary instruction is an FMA instruction, then the following process is executed: If one multiplication variable of the second MUL instruction participates in the multiplication operation in the temporary instruction together with the first variable and the second variable, and the other multiplication variable participates in the multiplication operation in the first MUL instruction together with the third variable and the temporary variable, and the addition variable in the temporary instruction is 1 or the first variable; Then change the first MUL instruction to a first FMA instruction, and use the result variable of the second MUL instruction and the second variable as the two multiplication variables of the first FMA instruction; When the addition parameter in the temporary instruction is 1, use the third variable as the addition variable of the first FMA instruction; when the addition parameter in the temporary instruction is the first variable, then use the result variable of the second MUL instruction as the addition variable of the first FMA instruction, and delete the temporary instruction.
5. A GPU compiler optimization method, characterized in that, Perform compiler optimization using the GPU compiler optimization method according to any one of claims 1 to 4, and optimize the instructions in the first window that contain constant parameters.
6. The GPU compiler optimization method according to claim 5, wherein The optimization of the instructions in the first window that contain constant parameters includes: If the temporary instruction in the first window is an FMA instruction, there is a first constant participating in the multiplication operation in the FMA instruction, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first MUL instruction that uses the temporary variable and a third constant to participate in the multiplication operation, Change the first MUL instruction to a first FMA instruction, participate in the multiplication operation in the first FMA instruction with the first value and the multiplication variable in the temporary instruction, participate in the addition operation in the first FMA instruction with the second value, and delete the temporary instruction; wherein, the first value is the product of the first constant and the third constant, and the second value is the product of the second constant and the third constant.
7. The GPU compiler optimization method according to claim 5, wherein The optimization of the instructions in the first window that contain constant parameters includes: If the temporary instruction in the first window is an FMA instruction, there is a first constant and a first variable participating in the multiplication operation in the FMA instruction, and there is a second constant participating in the addition operation, and after the temporary instruction, there is a first ADD instruction that uses the temporary variable to participate in the addition operation, then the following process is executed: If in the first ADD instruction, the temporary variable participates in an addition operation together with a third constant, then change the first ADD instruction to a first FMA instruction. The first FMA instruction uses the first constant in the temporary instruction to participate in a multiplication operation together with a first variable, uses a first value to participate in an addition operation, and deletes the temporary instruction; wherein, the first value is the sum of a second constant and a third constant. If in the first ADD instruction, the temporary variable participates in an addition operation together with a first variable, then change the first ADD instruction to a second FMA instruction. The second FMA instruction uses a second value and the first variable in the temporary instruction to participate in a multiplication operation, uses the second variable in the temporary instruction to participate in an addition operation, and deletes the temporary instruction; wherein, the second value is the value obtained by adding 1 to the first constant.
8. The GPU compiler optimization method according to claim 5, wherein The optimization of the instructions containing constant parameters in the first window includes: If the temporary instruction in the first window is an FMA instruction, in the FMA instruction, a first constant and a first variable participate in a multiplication operation together, a second constant participates in an addition operation, and after the temporary instruction, there is a first FMA instruction using the temporary variable to participate in an addition operation, If in the first FMA instruction, a third constant and a temporary variable participate in a multiplication operation together, and a fourth constant participates in an addition operation, then replace the two multiplication parameters in the first FMA instruction with a first value and the first variable respectively, replace the addition parameter in the first FMA instruction with a second value, and delete the temporary instruction; wherein, the first value is the product of the first constant and the third constant, and the second value is the sum obtained by adding the product of the second constant and the third constant to the fourth constant. If after the temporary instruction, there is a second FMA instruction using the temporary variable to participate in a multiplication operation, and in the second FMA instruction, a third constant and a first variable participate in a multiplication operation together, and the temporary variable participates in an addition operation, then replace the second constant in the second FMA instruction with a third value, replace the addition parameter in the second FMA instruction with the second constant, and delete the temporary instruction; wherein, the third value is the sum of the first constant and the second constant.
9. The GPU compiler optimization method according to claim 5, wherein The optimization of the instructions containing constant parameters in the first window includes: If in the first FMA instruction in the first window, a temporary variable and a first constant are used to participate in a multiplication operation, a second constant is used to participate in an addition operation, and the temporary command corresponding to the temporary variable is a MUL command or an ADD command, If the temporary command is a MUL command, and in the temporary command, a third constant and a first variable are used to participate in a multiplication operation, then replace the two multiplication parameters in the first FMA instruction with a first value and the first variable respectively, and delete the temporary instruction; wherein, the first value is the product of the third constant and the first constant. If the said temporary instruction is an ADD instruction, and in the said temporary instruction, a third constant and a first variable are used for addition operation, then replace the temporary variable in the said first FMA instruction with the first variable, replace the second constant in the said first FMA instruction with a second value, and delete the temporary instruction; wherein, the said second value is the sum of the product of the first constant and the third constant plus the second constant.
10. The GPU compiler optimization method according to claim 5, wherein The said method further includes: Before performing FMA instruction template matching optimization, check whether the instructions in the basic block contain division operations. If so, make the second operand become its reciprocal, turn the division into a multiplication operation, and then match with the defined multiply-add instruction template; Expand the above instruction template, replace the plus sign in the instruction template with a minus sign correspondingly, increase the number of templates, and expand the scope of template matching and optimization; After performing template matching optimization, for CPU, DSP, and ASIC processor backends that do not support FMA instructions, convert the optimized FMA instruction into a multiplication instruction and an addition instruction.
Citation Information
Patent Citations
Low-power-consumption register allocation compiling optimization method
CN112445481A
Compiling method for masked vector instruction, electronic equipment and medium
CN115328493A