Compiler optimization method, and device and storage medium
Patent Information
- Application Number
- PCT/CN2025/132074
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2025-11-03
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025132074_01102026_PF_FP_ABST
Abstract
Description
Compiler optimization methods, devices, and storage media Technical Field
[0001] This application relates to the field of computer technology, and in particular to a compiler optimization method, device and storage medium.
[0002] This application claims priority to Chinese Patent Application No. 202510353075.6, filed on March 24, 2025, entitled “Compiler Optimization Method, Apparatus and Storage Medium”, the entire contents of which are incorporated herein by reference. Background Technology
[0003] Alias analysis is a technique commonly used in compiler optimization. It determines whether two or more pointers in a program might point to the same memory location (or memory address). Accurate alias analysis can significantly improve the compiler's optimization capabilities, for example, by eliminating duplicate load or store memory operations.
[0004] The current alias analysis method will give an analysis conclusion of "yes" or "no". "Yes" means that there is a memory dependency between the two pointers and the related operations cannot be eliminated, while "no" means that there is no memory dependency between the two pointers and the related operations can be eliminated.
[0005] However, current alias analysis methods, in order to reduce false positives, typically overestimate the range of memory addresses and give a "yes" analysis conclusion, thus avoiding optimization of the compiled language. While this conservative analysis process can reduce errors, it also leads to lower accuracy in compiler optimization results. Technical issues
[0006] This application provides a compiler optimization method, device, and storage medium, which improves the accuracy of compiler optimization results. The technical solution is as follows:
[0007] Firstly, a compiler optimization method is provided, the method comprising: obtaining the vector length corresponding to an access operation in a compiled language; the vector length being related to the number of elements included in the memory address to be accessed; if the vector length corresponding to the access operation is not a power of a preset value and is greater than a preset threshold, splitting the access operation according to the vector length to obtain multiple sub-access operations; and optimizing repetitive operations in the compiled language according to the multiple sub-access operations to obtain a compilation optimization result.
[0008] Secondly, a compiler optimization apparatus is provided, comprising: an acquisition module for acquiring the vector length corresponding to an access operation in a compiled language; the vector length being related to the number of elements included in the memory address to be accessed; a splitting module for splitting the access operation into multiple sub-access operations based on the vector length when the vector length corresponding to the access operation is not a power of a preset value and is greater than a preset threshold; and an optimization module for optimizing repetitive operations in the compiled language based on the multiple sub-access operations to obtain a compilation optimization result.
[0009] Thirdly, a computer device is provided, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program implementing the method described in the first aspect when executed by the processor.
[0010] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0011] Fifthly, a computer program product containing instructions is provided that, when run on a computer, causes the computer to perform the method described in the first aspect.
[0012] This application provides a compiler optimization method, device, and storage medium. According to the scheme provided, the vector length corresponding to an access operation in the compiled language is obtained. The vector length is related to the number of elements included in the memory address to be accessed. When the vector length is not a power of a preset value and is greater than a preset threshold, the access operation is split according to the vector length to obtain multiple sub-access operations. Based on these multiple sub-access operations, repetitive operations in the compiled language are optimized to obtain the compilation optimization result. Standardizing the access operations in the compiler's intermediate representation (IR) and splitting them into multiple sub-access operations not only eliminates repetitive sub-access operations in the IR, but also, during instruction scheduling after register allocation, eliminates hardware-specific bank conflicts based on the instructions corresponding to each sub-access operation, thus improving the accuracy of the compiler optimization result. Attached Figure Description
[0013] Figure 1 is a flowchart of a compiler optimization method provided in an embodiment of this application;
[0014] Figure 2 is a flowchart of another compiler optimization method provided in an embodiment of this application;
[0015] Figure 3 is a flowchart of another compiler optimization method provided in an embodiment of this application;
[0016] Figure 4 is a flowchart of another compiler optimization method provided in an embodiment of this application;
[0017] Figure 5 is a flowchart of another compiler optimization method provided in an embodiment of this application;
[0018] Figure 6 is a flowchart of another compiler optimization method provided in an embodiment of this application;
[0019] Figure 7 is a schematic diagram of a compiler optimization device provided in an embodiment of this application;
[0020] Figure 8 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Embodiments of the present invention
[0021] The compiler optimization method provided in this application is applicable to at least the following scenarios during compiler alias analysis: code optimization, memory management and security, concurrent programming and multithreading, and specific domains (e.g., graphics processing and rendering, database query optimization). The compiler can be an open-source compiler framework, such as a low-level virtual machine (LLVM), the GNU Compiler Collection (GCC), or Clang. The compiler translates the intermediate language (IR) representation into executable machine instructions.
[0022] Alias analysis can be used to optimize the compiler, such as eliminating duplicate load or store memory operations, eliminating bank conflicts, and parallelizing instructions. Bank conflicts refer to the phenomenon of performance degradation caused by conflicts when multiple processing units try to access the same memory module.
[0023] Alias analysis is used to determine whether two or more pointers in a program might point to the same memory location. Current alias analysis methods can improve compiler optimization capabilities, but they still suffer from insufficient precision. For example, current alias analysis methods treat different pointers indiscriminately, treating pointers with constant base addresses as ordinary pointers (i.e., pointers to variables). For instance, suppose the first step is a load operation, the second a write operation, and the third a load operation. In one scenario, the second step performs a write operation on the address in the first step, and the third step performs a load operation on the address after the write operation. The second step depends on the first step, and the third step depends on the second step; therefore, the third step depends on the first step. Since the third step and the first step have the same memory address, the value read from that memory address in the third step is different from the value read from the memory address in the first step, and the third step cannot be deleted. Current alias analysis methods do not know whether the third step depends on the first step, but for the sake of conservatism, they give a "yes" conclusion, indicating that the third step depends on the first step, and do not optimize the compiler accordingly. Thus, the optimization result will definitely be correct. In another scenario, although the second step depends on the first, and the third step depends on the second, if the address of the third step's load operation is different from the address of the second step's write operation, then the value read from the memory address in the third step is the same as the value read from the memory address in the first step, and the third step can be deleted. However, the current alias analysis method still gives a "yes" conclusion and does not optimize the compiled code. Therefore, it is evident that the accuracy of compiler optimization results using the current alias analysis method is not high.
[0024] Current alias analysis methods still suffer from insufficient precision when handling complex pointer operations, subtraction, and negative immediate values. For example, the calculation of memory addresses for load or store operations within loops typically involves loop variables. Current alias analysis methods do not consider the number of iterations in the loop variable; they simply subtract two loop variables and calculate whether the memory address range of the difference (starting address plus vector length) overlaps with the memory addresses of the two loop variables. However, subtraction is signed, while computers use binary arithmetic, which is unsigned. Current alias analysis methods treat the difference as unsigned, resulting in very large numbers and thus larger estimated memory address ranges. This leads to overlapping analysis conclusions, i.e., interdependence, without optimizing the compiled code. Furthermore, current alias analysis methods do not consider the number of iterations in the loop variable, only providing a "yes" or "no" final conclusion, without providing the calculation results of the loop variable in each iteration of the LLVM IR expression. Consequently, during instruction scheduling after register allocation, hardware-specific bank conflicts cannot be completely eliminated. This shows that the accuracy of current alias analysis methods is not high.
[0025] This application provides a compiler optimization method, as shown in Figure 1. Figure 1 is a flowchart of a compiler optimization method provided in this application, which includes:
[0026] S101. Obtain the vector length corresponding to the access operation in the compiled language; the vector length is related to the number of elements included in the memory address to be accessed.
[0027] In this embodiment, the access operation can be a load operation or a store operation. A load operation refers to reading data from a memory address into a register. A store operation refers to writing data from a register into a memory address.
[0028] In this embodiment, the vector length is positively correlated with the number of elements included in the memory address to be read or written. Vector length can refer to the number of elements included in the memory address to be read or written. An element is the basic unit of a data structure, such as a value in an array. An element can be an integer, a character, or other data type. The size of an element depends on the data type; an integer (int) occupies 4 bytes, and a double-precision floating-point number (double) occupies 8 bytes. Vector length can also refer to the number of bytes occupied by the memory address to be read or written. A byte is a unit of measurement for storage space, used to measure the size of data. For example, for a memory address read in a load operation, assuming the memory address includes 100 elements, the vector length is 100 elements. If the data type is int, the vector length is 400 bytes; if the data type is double, the vector length is 800 bytes.
[0029] S102. If the length of the vector corresponding to the access operation is not a power of the preset value but is greater than the preset threshold, the access operation is split according to the vector length to obtain multiple sub-access operations.
[0030] In this embodiment, access operations are standardized. If the vector length corresponding to an access operation is a power of a preset value, the memory address of the access operation does not need to be split. If the vector length is less than or equal to a preset threshold, it indicates that the memory address length is not very long, and the access operation is not further split. During splitting, after splitting the access operation into multiple sub-access operations, if the remaining vector length is less than or equal to the preset threshold, further splitting is not performed.
[0031] It should be noted that the preset value can be set to a power of 2, such as 2, 4, 8, 16, etc. The preset threshold can be set to a power of 2, such as 64, 128, 256, etc., and this embodiment of the application does not impose any restrictions on this.
[0032] For example, taking a preset value of 2 and a preset threshold of 64 as an example, the splitting rule is that if the vector length is not a power of 2, or is less than or equal to 64, no splitting is performed. For example, if the vector length is 2000, 2000 = 1024 (2^2000 = 1024). 10 )+512(2 9 )+256(2 8 )+128 (2 7 )+64(2 6 ) +16; For example, if the vector length is 1030, 1030 = 1024 (2 10 )+6.
[0033] In some embodiments, the plurality of sub-access operations include at least a first sub-access operation and a second sub-access operation, wherein the vector length corresponding to the first sub-access operation is greater than half of the vector length corresponding to the access operation, and the vector length corresponding to the second sub-access operation is less than or equal to a preset threshold.
[0034] In this embodiment of the application, when splitting access operations, they can be split sequentially in descending order of powers of a preset value. Thus, after the first split, the vector length of the first sub-access operation is greater than half the vector length of the access operation itself. After the last split, the vector length of the second sub-access operation is less than or equal to a preset threshold. For example, 2000 = 1024 (2... 10 )+512(2 9 )+256(2 8 )+128 (2 7 )+64(2 6 )+16, the vector length of the first sub-access operation is 1024, and the vector length of the second sub-access operation is 16.
[0035] When splitting access operations, if the length of the vector corresponding to the access operation is not an integer multiple of a preset power of a preset value, but is greater than a preset threshold, the preset threshold is set to a preset power of the preset value. The access operation is then split according to the vector length, resulting in multiple sub-access operations. That is, it is always split according to a preset power of the preset value. For example, if the preset value is 2 and the preset power is 8, the vector length is 2000, and 2000 = 256 (2... 8 ) × 7 + 208; For example, if the vector length is 1030, 1030 = 256 (2 8 )×4+6.
[0036] It should be noted that the above-described scheme, which splits data according to a preset power of a preset value, is a small-granular split. This scheme eliminates duplicate access operations in the compiled language, improving the optimization effect at the LLVM IR level, but it does not affect the optimization effect at the machine instruction level. However, this scheme leads to an increasing number of sub-access operations, thus increasing compilation time.
[0037] In this embodiment, optimizing the compiled language includes not only eliminating duplicate access operations, but also instruction selection (also known as instruction optimization) during the conversion of the compiled language into machine instructions, instruction scheduling before register allocation, elimination of duplicate access operations in machine instructions, and instruction scheduling after register allocation to eliminate bank conflicts. The splitting of access operations occurs before eliminating duplicate access operations in the compiled language. This splitting is a large-granularity process; the granularity of subsequent instruction selection, instruction scheduling, elimination of duplicate access operations in machine instructions, and elimination of bank conflicts decreases sequentially, thus completing the entire optimization process. Therefore, the splitting of access operations here can follow a large-granularity method, for example, splitting them sequentially in descending order of powers of a preset value to eliminate most duplicate access operations in the compiled language and improve the accuracy of the compiler optimization results.
[0038] S103. Based on multiple sub-access operations, optimize repetitive operations in the compiled language to obtain the compilation optimization result.
[0039] In this embodiment, by splitting a long-vector access operation into multiple shorter-vector sub-access operations, and then eliminating duplicate operations within these sub-access operations, sub-access operations with duplicate memory addresses can be eliminated, improving the accuracy of the compiler's optimization results. Furthermore, these multiple sub-access operations can also be applied to subsequent instruction selection, instruction scheduling, elimination of duplicate access operations in machine instructions, and elimination of bank conflicts, further improving the accuracy of the optimization results in each subsequent optimization process.
[0040] In this embodiment, after eliminating redundant operations in multiple sub-access operations to obtain the compilation optimization result (which includes the optimized multiple sub-access operations), instruction conversion is performed on the optimized multiple sub-access operations to obtain machine instructions; then, physical register allocation (i.e., hardware register allocation) is performed on the machine instructions to complete the compiler optimization process. Before and after the physical register allocation process, instruction scheduling rules are combined to perform instruction scheduling before and after register allocation for the machine instructions. These will be described separately below.
[0041] The optimization process of a compiled language includes the process of eliminating duplicate operations in multiple sub-access operations (see the descriptions in S101-S103 above), the instruction selection process, the instruction scheduling process before register allocation, the process of eliminating duplicate access operations in machine instructions, the register allocation process, and the instruction scheduling process after register allocation.
[0042] The instruction selection process refers to the instruction conversion of multiple optimized sub-access operations to generate machine instructions related to the computer device, so as to ensure that the machine instructions generated after instruction conversion can be supported by the computer device.
[0043] The instruction scheduling process before register allocation refers to rearranging the execution order of machine instructions to obtain the instruction scheduling result, so as to ensure that the execution order of machine instructions does not depend on each other, reduce waiting, and improve resource utilization.
[0044] Eliminating duplicate access operations in machine instructions refers to optimizing the memory addresses that are repeatedly accessed in the instruction scheduling results, resulting in multiple optimized sub-access operations, each corresponding to a different instruction.
[0045] The register allocation process refers to allocating physical registers to the corresponding instructions of multiple optimized sub-access operations based on the number of physical registers in the hardware platform, so as to efficiently utilize limited register resources and improve program execution efficiency.
[0046] The instruction scheduling process after register allocation refers to scheduling the instructions corresponding to multiple sub-access operations after the allocated physical registers. In other words, the execution order of machine instructions in the allocated physical registers is rearranged to obtain the instruction scheduling result.
[0047] The compiler optimization process described above is based on the previously decomposed multiple sub-access operations, thereby improving the accuracy of the compiler optimization results.
[0048] According to the scheme provided in this application, the vector length corresponding to the access operation in the compiled language is obtained; the vector length is related to the number of elements included in the memory address to be accessed; when the vector length is not a power of a preset value and the vector length is greater than a preset threshold, the access operation is split according to the vector length to obtain multiple sub-access operations; based on the multiple sub-access operations, the repetitive operations in the compiled language are optimized to obtain the compilation optimization result. Standardizing the access operations in the compiler's intermediate language (IR) and splitting them into multiple sub-access operations not only eliminates repetitive sub-access operations in the intermediate language, but also, in the subsequent instruction scheduling process after register allocation, based on the instructions corresponding to each of the multiple sub-access operations, eliminates hardware-specific bank conflicts, improving the accuracy of the compiler optimization result.
[0049] In some embodiments, before S101, the compiler optimization method may further include the following steps: converting the base address of the pointer in the access operation to obtain a conversion result; if the conversion result indicates that the base address is a constant, replacing the base address with a constant and then executing subsequent S101-S103.
[0050] In this embodiment, an access operation can be understood as an expression of operands. Operands refer to data that participates in calculations within a program; these data can be constants, variables, or the result of an expression. In an access operation, a pointer represents an abstract memory address, and the operand is the data stored at the memory address pointed to by the pointer. By associating operands with pointers, operations and access to the data pointed to by the pointers are achieved.
[0051] When optimizing a compiled language, the compiler attempts to convert the base address of a pointer in an access operation to an integer constant. If the conversion is successful, the constant base address is used for subsequent optimization, reducing unnecessary memory dependencies in alias analysis (including eliminating duplicate access operations in the compiled language, instruction selection, instruction scheduling, eliminating duplicate access operations in machine instructions, and eliminating bank conflicts). If the conversion fails, the pointer is treated as a regular pointer. In other words, if the pointer's base address is constant, subsequent alias analysis directly uses that constant, improving the accuracy of the compiler's optimization results.
[0052] Current alias analysis methods indiscriminately process different pointers, leading to low accuracy in compiler optimization results. However, the compiler optimization method provided in this application utilizes information that can be determined at compile time. If the base address is constant, it is replaced with a constant during the compiler optimization process, eliminating unnecessary memory dependencies between the original load and store operations. For example, constants can be applied to processes such as eliminating duplicate load and store operations, instruction selection, instruction scheduling before register allocation, eliminating duplicate load and store operations in machine instructions, and instruction scheduling after register allocation (i.e., eliminating bank conflicts) in compiled languages (e.g., LLVM IR languages), thereby improving the accuracy of compiler optimization results.
[0053] In some embodiments, based on Figure 1 above, as shown in Figure 2, Figure 2 is a flowchart of another compiler optimization method provided by an embodiment of this application.
[0054] S201. Obtain the vector length corresponding to the access operation in the compiled language.
[0055] S201 is the same as S101 in Figure 1 above. Its implementation method and the technical effect achieved can be found in the description of S101 above, and will not be repeated here.
[0056] S202. Determine the type of access operation.
[0057] The types of access operations include load operations and store operations. Correspondingly, multiple sub-access operations are multiple sub-load operations and / or multiple sub-store operations.
[0058] When the access operation is a load operation, execute S203-S204; when the access operation is a storage operation, execute S205-S206.
[0059] S203. If the length of the first vector corresponding to the loading operation is not a power of a preset value but is greater than a preset threshold, the loading operation is split according to the length of the first vector to obtain multiple sub-loading operations; the length of the first vector is related to the number of elements included in the memory address to be read.
[0060] When splitting a loading operation, the splits are performed sequentially in descending order of powers of a preset value. Thus, after the first split, the vector length of the first sub-loading operation is greater than half the vector length of the loading operation. After the last split, the vector length of the second sub-loading operation is less than or equal to a preset threshold. After splitting a loading operation into multiple sub-loading operations, if the remaining vector length is less than or equal to the preset threshold, further splitting is not performed. Each sub-loading operation must include at least the first and second sub-loading operations.
[0061] S204. Merge multiple sub-load operations to obtain a merged operation, and then delete the load operation.
[0062] For loading operations, after splitting, multiple sub-loading operations need to be merged to obtain a merged operation, and then the loading operation itself needs to be deleted. When the computer executes the code, it merges the loading results corresponding to multiple sub-loading operations and replaces the original loading result corresponding to the loading operation. In other words, it merges the loading results corresponding to multiple sub-loading operations and uses them as new input, replacing the original input. However, the code of the original loading operation still exists, so it also needs to be deleted. This is called deleting dead code, which indicates code segments that do not affect the program's result.
[0063] S205. If the length of the second vector corresponding to the storage operation is not a power of a preset value but is greater than a preset threshold, the storage operation is split according to the length of the second vector to obtain multiple sub-storage operations; the length of the second vector is related to the number of elements included in the memory address to be written.
[0064] When splitting a storage operation, the splits are performed sequentially in descending order of powers of a preset value. Thus, after the first split, the vector length of the first sub-storage operation is greater than half the vector length of the storage operation. After the last split, the vector length of the second sub-storage operation is less than or equal to a preset threshold. After splitting a storage operation into multiple sub-storage operations, if the remaining vector length is less than or equal to the preset threshold, further splitting is not performed. Each sub-storage operation must include at least the first and second sub-storage operations.
[0065] S206, Delete storage operation.
[0066] Since a load operation reads data from a memory address into a register and then performs other operations, multiple sub-load operations need to be merged. A store operation, on the other hand, writes data to a memory address, so merging is unnecessary. For store operations, after splitting, the store operations also need to be deleted. When the computer executes code, it replaces the original store operation's result with the result of the multiple sub-store operations. However, the original store operation's code still exists, so it also needs to be deleted. This is called deleting dead code, which refers to code segments that do not affect the program's result.
[0067] S207 is executed after S204 or S206.
[0068] S207. Based on the merge operation and / or multiple sub-store operations, optimize the repetitive operations in the compiled language to obtain the compilation optimization result.
[0069] In this embodiment, by splitting long-vector loading and / or storage operations into multiple short-vector sub-loading and / or sub-storage operations, and thereby eliminating duplicate operations in the multiple sub-loading and / or sub-storage operations, sub-loading and / or sub-storage operations with duplicate memory addresses can be eliminated, thus improving the accuracy of compiler optimization results.
[0070] It should be noted that when the access operation type is a load operation, S203-S204 are executed, and then the repetitive operations in the compiled language are optimized based on the merge operation to obtain the compilation optimization result. When the access operation type is a store operation, S205-S206 are executed, and then the repetitive operations in the compiled language are optimized based on multiple sub-store operations to obtain the compilation optimization result. When the access operation type includes both load and store operations, S203-S206 are executed, and then the repetitive operations in the compiled language are optimized based on the merge operation and multiple sub-store operations to obtain the compilation optimization result. In this embodiment, the execution order of S203-S204 and S205-S206 is not limited; they can be executed sequentially or simultaneously.
[0071] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0072] Based on Figures 1 and 2 above, and exemplarily as shown in Figure 3, Figure 3 is a flowchart of another compiler optimization method provided in an embodiment of this application. Taking the LLVM IR language as an example, with a preset value of 2 and a preset threshold of 30, the splitting of load and / or store operations in the LLVM IR language is explained.
[0073] S301. Normalize the LLVM IR language to obtain the vector length of the LLVM IR language.
[0074] Before breaking down the load and / or store operations in the LLVM IR language, the instructions involving memory address operations in the LLVM IR language are first normalized to obtain the vector length. The vector length refers to the length of the number of elements (or bytes) stored in the memory address to be loaded or stored.
[0075] S302. Determine the instruction type corresponding to the LLVM IR language.
[0076] The LLVM IR language includes load and store operations as instruction types. For load operations, execute instructions S303-S305; for store operations, execute instructions S306-S307.
[0077] S303. Determine if the length of a vector is a power of 2.
[0078] S304. The loading operation is split into multiple sub-loading operations of vector length that are powers of 2.
[0079] For loading operations where the vector length is not a power of 2, the loading operation is split into multiple smaller sub-loading operations where the vector length is a power of 2. The remaining vectors with a length of less than 30 are not split.
[0080] S305. Merge multiple sub-loading operations, and use the output of the merged operation as a new instruction to replace the output of the original loading operation.
[0081] These smaller sub-load operations are merged into a merge operation of the original vector length, and the load result corresponding to the merge operation (i.e., the new instruction) replaces the load result of the original load operation.
[0082] S306. Determine if the length of a vector is a power of 2.
[0083] S307. Break down the storage operation into multiple sub-storage operations of vector length that are powers of 2.
[0084] For storage operations where the vector length is not a power of 2, the storage operation is split into multiple smaller sub-storage operations where the vector length is a power of 2. The remaining vectors with a length of less than 30 are not split.
[0085] S308 is executed after S305 or S307.
[0086] S308. Delete the code for the original load operation and / or the code for the original store operation.
[0087] In S305 above, replacement refers to using the loading result corresponding to the merge operation as input, but the old code (i.e., the code of the loading operation) still exists, so it is also necessary to delete the dead code generated during the splitting process (i.e., the code of the original loading operation). Similarly, for the storage operation, it is also necessary to delete the dead code generated during the splitting process (i.e., the code of the original storage operation).
[0088] It should be noted that for the splitting of load operations, S303-S305 are executed, and then repetitive operations in the LLVM IR language are optimized based on the multiple sub-load operations after S308. For the splitting of storage operations, S306-S307 are executed, and then repetitive operations in the LLVM IR language are optimized based on the multiple sub-storage operations after S308. For the splitting of load and storage operations, S303-S308 are executed, and then repetitive operations in the LLVM IR language are optimized based on the merge operation and multiple sub-storage operations. This example does not restrict the execution order of S303-S305 and S306-S307; they can be executed sequentially or simultaneously.
[0089] In this embodiment, load and store operations in the LLVM IR language are split into lengths that are powers of 2, and any remaining vectors with lengths less than a preset threshold are not split. This allows compiler optimization to eliminate sub-load and / or sub-store operations with duplicate memory addresses, improving the accuracy of compiler optimization results.
[0090] In some embodiments, the compilation optimization result includes multiple optimized sub-access operations. After S103, the compiler optimization method further includes the following steps: Instruction translation is performed on the multiple optimized sub-access operations to generate machine instructions; instruction scheduling is performed on the execution order of the machine instructions to obtain an instruction scheduling result; the instruction scheduling result includes the memory address corresponding to the machine instruction; the repeatedly accessed memory addresses in the instruction scheduling result are optimized to obtain the instructions corresponding to each of the multiple optimized sub-access operations.
[0091] In this embodiment of the application, when optimizing the compiled language, the backend processing flow includes not only the process of eliminating duplicate access operations in the compiled language as shown in Figures 1-3 above, but also the instruction selection process, the instruction scheduling process before register allocation, the process of eliminating duplicate access operations in machine instructions, and the instruction scheduling process after register allocation (this process includes, but is not limited to, the process of eliminating bank conflicts), which will be described below.
[0092] Because different computer devices support different machine instruction sets, it is necessary to perform instruction translation on the compiled language (i.e., optimized multiple sub-access operations). Instruction translation includes operand type instruction translation and operation instruction translation to ensure that the generated machine instructions can be supported by the computer device. The goal of the instruction selection process is to generate machine instructions relevant to the computer device. The process of mapping the compiled language to machine instructions is called the instruction selection process. During the instruction selection process, optimizations such as merging and deleting machine instructions are performed to improve the efficiency of subsequent machine instruction scheduling. The instruction selection process can also be called the instruction optimization process.
[0093] Instruction scheduling refers to the process of rearranging generated machine instructions during compilation to determine their execution order, aiming to optimize execution timing. For example, machine instructions may have data dependencies—the output of one instruction might be the input of another. Without sufficient scheduling, subsequent instructions might have to wait for the previous one to complete, leading to performance degradation. Therefore, instruction scheduling can reduce waiting caused by data dependencies by changing the execution order. Proper scheduling of machine instructions allows for the parallel execution of independent instructions, improving instruction-level parallelism and fully utilizing the multiple execution units of the central processing unit (CPU), thereby increasing CPU resource utilization and program throughput.
[0094] The instruction scheduling process is divided into pre-register allocation (pre-RA) instruction scheduling and post-register allocation (post-RA) instruction scheduling. Here, the instruction scheduling process refers to pre-register allocation instruction scheduling, which sorts machine instructions according to the topology, considers instruction-level parallelism, and generates an instruction sequence (i.e., the instruction scheduling result).
[0095] Instruction scheduling is based on the parallelism between machine instructions, without considering the memory addresses corresponding to the machine instructions. The instruction scheduling result includes the memory addresses of the machine instructions. Next, it is necessary to eliminate duplicate load and / or store operations in the machine instructions to optimize them, resulting in optimized instructions for each of the multiple sub-access operations.
[0096] After instruction optimization (i.e., each of the optimized sub-access operations corresponds to an instruction), an unlimited number of virtual registers can be used. However, the number of physical registers on the hardware platform is finite. If the number of physical registers is insufficient to accommodate all virtual registers, the virtual registers will spill into memory. Therefore, instruction scheduling after register allocation is also required. The purpose of this scheduling is to sort the machine instructions that have been allocated physical registers, reduce bank conflicts, and improve the accuracy of the compiler's optimization results.
[0097] In summary, the compiler backend processing flow converts the compiled language into assembly code or binary code for a specific target architecture, thereby generating the final executable file.
[0098] In this embodiment, the compiler optimization process during backend processing (including instruction selection, instruction scheduling before register allocation, elimination of duplicate access operations in machine instructions, and instruction scheduling after register allocation) is based on the aforementioned split into multiple sub-access operations, thereby improving the accuracy of the compiler optimization results.
[0099] In some embodiments, the compilation optimization result includes multiple optimized sub-access operations. After S103, the compiler optimization method further includes S401-S404. In this embodiment, the optimized multiple sub-access operations are subjected to instruction conversion, instruction scheduling before register allocation, and elimination of duplicate load and / or store operations in the machine instructions to obtain the instructions corresponding to each of the optimized multiple sub-access operations. The instruction scheduling after register allocation based on the instructions corresponding to each of the optimized multiple sub-access operations will be described next with reference to FIG4. As shown in FIG4, FIG4 is a flowchart of another compiler optimization method provided by an embodiment of this application.
[0100] S401. Obtain the iteration count of each instruction corresponding to the optimized multiple sub-access operations.
[0101] In this embodiment, analysis tools or functions can be used to predict the number of iterations of an instruction. For example, the compiler's application programming interface (API) provides a wealth of analysis methods and tools. The compiler's API can be used to analyze the loop variable of an instruction with a loop structure and predict the number of iterations.
[0102] In some embodiments, S401 described above can be implemented in the following way: Perform loop variable analysis on the instructions corresponding to each of the optimized sub-access operations to obtain the iteration count for each of the optimized sub-access operations; if the loop execution count is predicted, use the loop execution count as the iteration count; if the loop execution count is not predicted, set the iteration count to a preset number.
[0103] The above process involves converting operations (store or load operations) into machine instructions. Machine instructions define the types of operations to be executed, while operations are the means to implement these instructions. Based on this, analysis tools or functions are used to perform loop transformation analysis on each machine instruction to predict the number of iterations for each corresponding operation. The number of iterations for some operations can be predicted, while the number for others cannot. Therefore, if the number of loop executions for an operation can be predicted, this number is used as the number of iterations for that operation. If the number of loop executions for an operation cannot be predicted, a preset number of iterations needs to be set based on experience, and this preset number is used as the number of iterations for that operation.
[0104] It should be noted that the preset number of times can be appropriately set by those skilled in the art based on a large number of experimental thresholds, such as 10, 20, 100, 200, etc., and this application embodiment does not limit this.
[0105] In this embodiment, analysis tools or functions are used to perform loop variable analysis on each instruction with a loop structure to predict the number of iterations for each instruction's corresponding operation. The predicted or set number of iterations is used in the instruction scheduling process after register allocation. By considering the number of loop executions in the loop variable, bank conflicts can be reduced.
[0106] S402. When assigning the target instruction to the sequence, the dependency between the target instruction and each first instruction that has been assigned in the sequence is analyzed to obtain the calculation result of each iteration when running the target instruction; the target instruction is any instruction that has not been assigned, and the sequence includes multiple first instructions that do not depend on each other and are executed in parallel.
[0107] An instruction that has been assigned represents a machine instruction whose order has been determined within its sequence. During instruction scheduling, each machine instruction is assigned to a sequence, which can be understood as an instruction packet or issue packet containing multiple instructions to be executed in parallel. Taking a target instruction as an example, to assign a target instruction to a sequence, assuming there are no already assigned instructions in that sequence, the target instruction can be assigned to it. Assume that the sequence includes one or more first instructions that have been assigned. When there is only one first instruction, the dependency between the target instruction and that first instruction needs to be analyzed to obtain the result of each iteration. When there are multiple first instructions, these multiple first instructions are independent of each other and can be executed in parallel. In this case, the dependency between the target instruction and each of the multiple first instructions in the sequence needs to be analyzed individually, that is, the dependency between the target instruction and each of the first instructions needs to be analyzed.
[0108] S403. If the calculation result indicates that the target instruction depends on any first instruction, continue to analyze the dependency between the target instruction and each second instruction that has been allocated in the next sequence, until the calculation results of multiple iterations all indicate that the target instruction does not depend on the instructions that have been allocated in the target sequence, and then allocate the target instruction in the target sequence.
[0109] For any first instruction, when executing the target instruction, if the result of any iteration indicates that the target instruction depends on that first instruction, then the target instruction cannot be assigned to that sequence. Continue analyzing the dependencies between the target instruction and the next sequence. If there are no already assigned instructions in the next sequence, then the target instruction can be assigned to that sequence. Assume the next sequence includes one or more already assigned second instructions. When there is only one second instruction, the dependency between the target instruction and that second instruction needs to be analyzed to obtain the result of each iteration. When there are multiple second instructions, the dependency between the target instruction and each of the multiple second instructions in the sequence needs to be analyzed one by one, that is, the dependency between the target instruction and each first instruction needs to be analyzed. This process continues until a target sequence is found where there are no already assigned instructions, or where the results of multiple iterations all indicate that the target instruction does not depend on any already assigned instructions in the target sequence. In this case, the target instruction can be assigned to the target sequence.
[0110] S404. If the results of multiple iterations all indicate that the target instruction does not depend on each first instruction, the target instruction is allocated in the sequence, and the allocation of the next unallocated instruction continues until the optimized sub-access operations have their respective corresponding instructions allocated, thus obtaining the instruction scheduling result.
[0111] During the process of allocating a target instruction to a sequence, assuming that there are no instructions yet to be allocated in the sequence, or that the results of multiple iterations indicate that the target instruction does not depend on the first instruction already allocated in the sequence, the target instruction can be allocated to that sequence. Following the allocation process of the target instruction, the allocation of the next unallocated instruction continues until all sub-access operations have their corresponding instructions allocated, thus completing the instruction scheduling process and obtaining the instruction scheduling result. The instruction scheduling result is a series of machine instructions, which are then processed by an assembler to generate assembly code or binary code.
[0112] This example illustrates the scenario where the target instruction undergoes 50 iterations, and the desired sequence includes three already allocated first instructions. Example 1: For each first instruction, the result of each iteration of the target instruction's operation needs to be obtained. For the first first instruction, assuming that the result at the 23rd iteration indicates a dependency between the target instruction and the first first instruction, then the target instruction cannot be allocated to the sequence. Assuming that the results of all 50 iterations indicate no dependency between the target instruction and the first first instruction, we continue with the second first instruction. If the result of the third iteration indicates a dependency between the target instruction and the second first instruction, then the target instruction cannot be allocated to the sequence. If the results of all 50 iterations of the operation of each of the three first instructions indicate no dependency between them, then the target instruction can be allocated to the sequence. This process continues until all machine instructions have been allocated.
[0113] Example 2: Obtain the calculation result of the target instruction at each iteration. For the three first instructions, assume that the calculation results of 22 iterations all indicate that the target instruction is independent of the three first instructions. However, in the 23rd iteration, the calculation result indicates that the target instruction is dependent on the first first instruction, so the target instruction cannot be assigned to this sequence. Assuming that the calculation results of 23 iterations all indicate that the target instruction is independent of these three first instructions, continue to check the calculation result of the 24th iteration for dependency on each of the three first instructions, until the calculation results of each of the three first instructions and the target instruction in 50 iterations all indicate that the target instruction is independent of that first instruction, and the target instruction can be assigned to this sequence. Continue in this manner until all machine instructions have been assigned.
[0114] In Example 1 above, for each first instruction, it checks whether the target instruction is dependent on any of the first instructions in each iteration. If the results of multiple iterations indicate that the target instruction is not dependent on the first first instruction, then it checks whether the target instruction is dependent on the second first instruction, and so on. In Example 2 above, for multiple first instructions, it checks whether the target instruction is dependent on all of the first instructions in each iteration, and so on.
[0115] The instruction scheduling process after register allocation is as follows: For machine instruction A, we want to allocate it to sequence A. Assuming there are no already allocated machine instructions in sequence A, we allocate machine instruction A to sequence A. For machine instruction B, we want to allocate it to sequence A. Since sequence A includes already allocated machine instructions of A, we need to analyze the dependency between machine instruction B and machine instruction A. Assuming machine instruction B iterates 10 times, we need to obtain the result of each iteration of machine instruction B. If, in the 6th iteration, the result indicates that machine instruction B and machine instruction A are dependent, then machine instruction B cannot be allocated to sequence A. If, in all 10 iterations, the result indicates that machine instruction B and machine instruction A are not dependent, then machine instruction B can be allocated to sequence A. This process continues until all machine instructions are allocated.
[0116] Current alias analysis methods do not consider the number of iterations in the loop variable, resulting in the inability to completely eliminate hardware-specific bank conflicts. However, the compiler optimization method provided in this application analyzes the iteration count of instructions and the dependencies between instructions and the already allocated instructions in the sequence. This ensures that instructions are allocated to appropriate sequences where they are independent and can be executed in parallel. This minimizes bank conflicts, reduces hardware overhead caused by bank conflicts, and improves the accuracy of compiler optimization results.
[0117] In some embodiments, when assigning the target instruction to the sequence in S402, the dependency between the target instruction and the first instructions is analyzed for each first instruction in the sequence, and the computation result of running the target instruction in each iteration is obtained. Assuming the number of iterations for the target instruction is K, and the sequence includes M first instructions, then for each first instruction, the target instruction needs to be run K times. Of course, the number of runs may be less than K before the target instruction becomes dependent on the first instructions, indicating that the target instruction cannot be assigned to this sequence, and there is no need to analyze the dependency between the target instruction and the first instructions. Thus, if the computation results of K×M iterations all indicate that the target instruction does not depend on any of the first instructions, then the target instruction can be assigned to this sequence. The process of determining whether the computation result of each iteration indicates that the target instruction depends on the first instructions will be explained next.
[0118] For the Nth iteration, where N is a positive integer, after the Nth iteration, if the addressing mode of the target instruction is base addressing, the memory address corresponding to the target instruction is determined. The memory address includes the starting address and the vector length. The starting address is determined based on the base address and the offset. If the memory address corresponding to the target instruction does not overlap with the memory address corresponding to the first instruction, the target instruction is used as the result of the Nth iteration when running the target instruction, independent of the first instruction. If the memory address corresponding to the target instruction overlaps with the memory address corresponding to the first instruction, the target instruction is used as the result of the Nth iteration when running the target instruction, dependent on the first instruction.
[0119] The base address acts like a pointer, pointing to the starting position of the memory region to be accessed. The starting address is the base address plus an offset, representing a specific location in memory. The vector length represents the number of bytes occupied by the accessed memory. Based on the starting address and the vector length, the range of memory addresses (which can be simply referred to as memory addresses) can be calculated. In the Nth iteration, the target instruction is executed, and its memory address is calculated. When the target instruction's addressing mode is base addressing, the memory address corresponding to the target instruction is determined, and it is checked whether this memory address overlaps with the memory address corresponding to the first instruction. If the memory address of the target instruction does not overlap with the memory address of the first instruction, the target instruction is independent of the first instruction and is used as the result of the Nth iteration. If the memory address of the target instruction overlaps with the memory address of the first instruction, the target instruction depends on the first instruction and is used as the result of the Nth iteration.
[0120] In this embodiment, the memory address of the target instruction is calculated to determine whether the memory addresses of the target instruction and the first instruction overlap, thereby determining whether the target instruction and the first instruction are interdependent, and ultimately determining whether the target instruction can be allocated into the sequence. This approach not only considers multiple iterations but also the possibility of parallel processing between the target instruction and each of the allocated first instructions in the sequence. Therefore, during instruction scheduling after register allocation, it helps eliminate bank conflicts and improves the accuracy of compiler optimization results.
[0121] In some embodiments, after each iteration, it is determined whether the addressing mode of the target instruction is base addressing. If the addressing mode of the target instruction is not base addressing, recursive analysis is performed on each operand in the target instruction to obtain the expression characteristics of each operand; if the expression characteristics of any operand are not base addresses, recursive analysis is continued on each sub-operand in the target operand to obtain the expression characteristics of each sub-operand, until the expression characteristics of each sub-operand are all base addresses; the target operand indicates the operand whose expression characteristics are not base addresses; if the expression characteristics of each operand are all base addresses, it is determined that the addressing mode of the target instruction is base addressing.
[0122] When the addressing mode of the target instruction is not base addressing, recursive analysis is performed on all operands in the expression of the target instruction to obtain the expression characteristics of each operand. For example, taking the expression of the target instruction as a+b, if a and b are not base addresses, then recursive analysis is performed on a and b. Assuming a=a1×a2+a3 and b=b1+5, the analysis continues to determine if a1, a2, a3, and b1 are base addresses. Assuming a1, a2, and b1 are base addresses, and a3=a4+a5, the analysis continues to determine if a4 and a5 are base addresses. Assuming a4 and a5 are base addresses, then the addressing mode of the target instruction is base addressing, and the memory address corresponding to the target instruction can be obtained accordingly.
[0123] In this embodiment, by analyzing the addressing mode of the target instruction, when the target instruction is not the base address, all operands in the expression of the target instruction are recursively analyzed until the target instruction becomes the base address, thus obtaining the memory address of the target instruction. In this way, the target instruction whose starting address is the base address plus an offset can be obtained. Combined with the vector length, the memory address of the target instruction can be obtained. This allows for comparison with the memory address of the first instruction to determine whether their memory addresses overlap. Compared to current alias analysis techniques that only provide a "yes" or "no" final conclusion, by outputting the memory address and then comparing the memory addresses to determine whether the target instruction and the first instruction are interdependent, the accuracy of the judgment result is improved.
[0124] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0125] In this embodiment, when scheduling instructions after register allocation, the result of the iteration is queried for each instruction. Assuming the API analysis tool gives an evaluation result indicating that the instruction requires 20 iterations, then the iteration result should be 20. If the result of the 15th iteration indicates a dependency between the instruction and already allocated instructions in the sequence, then the position of this instruction needs to be readjusted, attempting to allocate the instruction to the next sequence to isolate it from the instructions in the sequence, i.e., to prevent parallel processing, thereby eliminating bank conflicts during scheduling. Figure 5 is a flowchart of another compiler optimization method provided in this embodiment. The instruction scheduling process after register allocation is explained below using Figure 5.
[0126] S501, Predict the number of iterations for the target instruction.
[0127] The above process involves instruction translation from operations (store or load operations) to obtain machine instructions. Machine instructions define the types of operations to be executed, while operations are the means to implement these instructions. The target instruction is any instruction to be allocated, and the process of addressing the memory address of the instruction's expression is essentially the process of addressing the operation. The calculation of the memory address of load and / or store operations within a loop typically involves loop variables. Therefore, it is necessary to evaluate whether the number of times the loop body is executed (i.e., the number of iterations of the predicted instruction) is constant. The compiler's API analysis tools are used to evaluate the expressions in the IR language and predict the number of iterations of the target instruction.
[0128] If the number of times the loop body is executed can be determined at compile time, then all values of the loop variable are successively substituted into the calculation expression of the memory addresses of the load and / or store operations to calculate the result of the expression at each iteration when the program actually runs these operations. If the number of times the loop body is executed cannot be determined at compile time, then a preset number needs to be specified by a person skilled in the art as the iteration number; that is, the result of the expression at the specified iteration number can only be calculated at compile time.
[0129] In other words, for some expressions, the number of loop iterations can be determined at compile time, and API analysis tools can estimate the number of iterations, such as 10 or 20. For other expressions, the number of loop iterations cannot be determined at compile time, and API analysis tools will give an unknown conclusion. Therefore, to eliminate bank conflicts during instruction scheduling, those skilled in the art need to specify a preset number, such as 100 or 8. That is, the number of iterations is set to a preset number.
[0130] S502. Set the iteration variable s, s=1.
[0131] The iteration variable s is a loop variable used to determine whether the iteration has ended.
[0132] S503. Determine if s exceeds the number of iterations.
[0133] If not, then execute S504.
[0134] In each iteration, the process of determining whether the memory address corresponding to the expression of the target instruction overlaps with the memory address of any instruction in the sequence is performed. If so, the process ends, and the process of determining whether the memory address of the target instruction overlaps with the memory address of the next instruction in the sequence is executed again. This can be understood as executing S502-S507 again, and the object compared with the memory address of the target instruction is the next instruction.
[0135] In each iteration, the process of determining whether the memory address corresponding to the expression of the target instruction overlaps with the memory addresses of all instructions in the sequence is performed. If so, the process ends and the process of determining whether the memory addresses of the target instruction overlap with the memory addresses of all instructions in the next sequence is performed. This can be understood as executing S502-S507 again, and the objects compared with the memory address of the target instruction are all instructions in the next sequence.
[0136] S504. Determine whether the expression of the target instruction is a base address.
[0137] If yes, then execute S505; otherwise, execute S508. The purpose of each determination of whether the expression of the target instruction is a base address is to calculate the memory address of the expression. The memory address includes the starting address and the vector length. The starting address is obtained by adding an offset to the base address.
[0138] S505. Calculate the memory address corresponding to the expression of the target instruction.
[0139] S506. Determine if memory addresses overlap.
[0140] Whether memory addresses overlap is determined by evaluating whether the expressions of the target instruction overlap with the expressions of the allocated instructions in the sequence in each dimension.
[0141] If not, then execute S507.
[0142] In each iteration, the process checks whether the memory address corresponding to the expression of the target instruction overlaps with the memory address of any instruction in the sequence. If so, the process ends, and the process of checking whether the memory address of the target instruction overlaps with the memory address of the next instruction in the sequence is executed. This can be understood as executing S502-S507 again, and the object compared with the memory address of the target instruction is the next instruction.
[0143] In each iteration, the process of determining whether the memory address corresponding to the expression of the target instruction overlaps with the memory addresses of all instructions in the sequence is performed. If so, the process ends and the process of determining whether the memory addresses of the target instruction overlap with the memory addresses of all instructions in the next sequence is performed. This can be understood as executing S502-S507 again, and the objects compared with the memory address of the target instruction are all instructions in the next sequence.
[0144] S507, s = s + 1.
[0145] S508. Calculate the result of all operands in the expression of the target instruction.
[0146] Memory addresses are typically composed of a series of expressions. Therefore, the process of calculating memory addresses requires recursively analyzing the values of the operands of these expressions until the base address is found.
[0147] The process calculates the result of all operands in an expression. For example, in the expression d = a + b, the operands are a and b. It then determines whether operands a and b are base addresses. If they are, the memory address of the expression can be calculated. If not, the process continues calculating the results of all sub-operands included in a and b until the base address is found, at which point the memory address of the expression is calculated.
[0148] After S507 or S508, continue executing S503.
[0149] This application proposes a compiler optimization method to improve compiler alias analysis methods. By predicting the iteration count of the loop body at compile time and accurately calculating the memory address of expressions in the compiled language (e.g., LLVM IR) at each iteration, complex pointer operations, subtraction between different base pointers, and operations containing negative immediate values can be handled. This can improve the performance of current compiler optimization techniques and can also be used to eliminate hardware-specific compiler optimizations such as bank conflicts.
[0150] Based on the access operation splitting process shown in Figures 1-3 above, and the bank conflict elimination process shown in Figures 4 and 5, the compiler's multiple optimization processes are explained here based on the alias analysis method. Figure 6 is a flowchart of another compiler optimization method provided in this application embodiment. The explanation uses LLVM IR language as an example.
[0151] S601. Decompose the load and / or store operations in the LLVM IR language to obtain multiple sub-load operations and / or multiple sub-store operations.
[0152] The scheme for splitting load operations and / or storage operations as described in Figures 1-3 above can be used in S601.
[0153] S602. Eliminate duplicate subload and / or substore operations in the LLVM IR language.
[0154] Based on the standardization of the LLVM IR language described in any of the embodiments in Figures 1-3, repetitive load and / or store operations in the LLVM IR language are eliminated.
[0155] S603: Convert LLVM IR language into machine instructions through instruction selection.
[0156] Instruction selection involves translating IR language into computer-specific assembly code (i.e., pseudo-instructions or machine instructions). During instruction selection, LLVM's alias analysis process is performed. Based on the alias analysis results, unnecessary dependencies between machine instructions are eliminated, allowing these machine instructions to be executed in parallel during subsequent instruction scheduling.
[0157] Here are a few scenarios to illustrate the instruction selection process. In one scenario, if both addition and subtraction are performed, and the result is 0, then no operation is needed, and the instruction is deleted. In another scenario, if both addition and multiplication are performed, these two instructions can be combined into one. In yet another scenario, when accessing a memory address, assuming a previous load instruction has already retrieved the content from that address and the content hasn't changed, subsequent accesses to that address will, through alias analysis, reveal that the content has already been retrieved and can be used directly; in this case, the load instruction is deleted.
[0158] S604: Order machine instructions through instruction scheduling before register allocation.
[0159] During instruction scheduling before register allocation, LLVM's alias analysis process is executed. Based on the alias analysis results, the order of machine instructions is adjusted. Some instructions can be executed in parallel, while others have a specific execution order.
[0160] S605, Eliminate duplicate sub-load operations and / or sub-store operations in machine instructions.
[0161] After S603 converts the IR language into machine instructions, it can further determine the starting address and vector length (or vector size) of each machine instruction. When eliminating duplicate operations in machine instructions, the LLVM aliasing process is executed, and duplicate loaded and / or stored machine instructions are deleted based on the aliasing analysis results.
[0162] After deleting duplicate loaded and / or stored machine instructions, physical registers are allocated to the machine instructions. This involves mapping virtual registers in the machine instructions to a limited number of physical registers to efficiently utilize limited register resources, reduce memory accesses, and thus improve program execution efficiency. Register allocation strategies include, but are not limited to, graph coloring algorithms, linear scan algorithms, and greedy algorithms.
[0163] S606. Schedule machine instructions that have been allocated physical registers through instruction scheduling after register allocation.
[0164] The instruction scheduling scheme described in any of the embodiments in Figures 4 and 5 above can be used in S606 to implement instruction scheduling after register allocation.
[0165] Instruction scheduling after register allocation involves sorting the machine instructions in the allocated physical registers. Instruction scheduling is performed based on the results of each iteration (i.e., whether the memory addresses of two instructions overlap), as shown in Figure 5. The scheduling principle is that two parallel instructions cannot have bank conflicts. After allocating each instruction, the iteration result is checked. Assuming there are 10 allocated instructions in the instruction packet or issue packet, and this instruction requires 20 iterations, then the 20 × 10 iterations must not result in a bank conflict; then the next instruction is allocated. Assuming there can be 4 parallel instructions in the instruction packet or issue packet, when allocating the second instruction, it is parallelized with the first instruction. The iteration result is checked again. If there is no bank conflict, the third instruction is allocated, parallelized with the first and second instructions. The iteration result is checked again. If there is no bank conflict, the fourth instruction is allocated to the next instruction packet or issue packet, and so on.
[0166] It should be noted that the granularity of instruction decomposition, instruction selection, instruction scheduling before register allocation, and instruction scheduling after register allocation in IR language becomes increasingly finer. Therefore, although the scheme in Figures 1-3, which decomposes instructions sequentially from powers of a preset value in descending order, eliminates duplicate operations through large-granularity decomposition, subsequent smaller-granularity decompositions can further eliminate duplicate operations. In this way, duplicate load and / or store operations can be eliminated sequentially without affecting the elimination effect.
[0167] In this embodiment, by splitting the load and / or store operations in the LLVM IR language into multiple sub-load and / or sub-store operations, it is beneficial to eliminate redundant operations in the LLVM IR language and improve the effect of redundant operation elimination. Then, the LLVM IR language is converted into instructions, and instruction optimization is performed in conjunction with alias analysis results during the conversion to machine instructions. After conversion to machine instructions, the order of machine instructions is adjusted in conjunction with alias analysis results to complete instruction scheduling before register allocation. After adjusting the order of machine instructions, redundant operations in machine instructions are eliminated in conjunction with alias analysis results. Then, the memory address of each iteration of the LLVM IR expression in the loop is calculated and compared with the memory address of the allocated instruction to obtain the calculation result. The calculation result can indicate whether the memory addresses of the two overlap, and thus determine whether the instruction can be allocated to the sequence. This process is repeated to complete instruction scheduling after register allocation, eliminate bank conflicts, and improve the accuracy of compiler optimization results.
[0168] Based on the compiler optimization method provided in the above embodiments, Figure 7 is a schematic diagram of a compiler optimization device provided in an embodiment of this application. This device can be implemented as part or all of a computer device by software, hardware, or a combination of both. Referring to Figure 7, the compiler optimization device 70 includes: an acquisition module 701, used to acquire the vector length corresponding to an access operation in the compiled language; the vector length is related to the number of elements included in the memory address to be accessed; a splitting module 702, used to split the access operation according to the vector length when the vector length corresponding to the access operation is not a power of a preset value and is greater than a preset threshold, to obtain multiple sub-access operations; and an optimization module 703, used to optimize repetitive operations in the compiled language based on the multiple sub-access operations to obtain a compilation optimization result.
[0169] Optionally, the multiple sub-access operations include at least a first sub-access operation and a second sub-access operation, wherein the vector length corresponding to the first sub-access operation is greater than half of the vector length corresponding to the access operation, and the vector length corresponding to the second sub-access operation is less than or equal to a preset threshold.
[0170] Optionally, the access operation includes a load operation and / or a store operation; the splitting module 702 is further configured to split the load operation according to the length of the first vector when the length of the first vector corresponding to the load operation is not a power of a preset value and is greater than a preset threshold, to obtain multiple sub-load operations; the length of the first vector is related to the number of elements included in the memory address to be read; and / or, when the length of the second vector corresponding to the store operation is not a power of a preset value and is greater than a preset threshold, to split the store operation according to the length of the second vector, to obtain multiple sub-store operations; the length of the second vector is related to the number of elements included in the memory address to be written; the multiple sub-access operations are multiple sub-load operations and / or multiple sub-store operations.
[0171] Optionally, the optimization module 703 is further configured to merge multiple sub-load operations to obtain a merged operation; delete load operations and optimize repetitive operations in the compiled language based on the merged operation to obtain a compilation optimization result; or delete store operations and optimize repetitive operations in the compiled language based on multiple sub-store operations to obtain a compilation optimization result; or merge multiple sub-load operations to obtain a merged operation; delete load operations and store operations and optimize repetitive operations in the compiled language based on the merged operation and multiple sub-store operations to obtain a compilation optimization result.
[0172] Optionally, the compilation optimization result includes multiple optimized sub-access operations; the compiler optimization device 70 also includes an allocation module 704.
[0173] The acquisition module 701 is used to acquire the iteration count of the instructions corresponding to the optimized sub-access operations. The allocation module 704 is used to analyze the dependency between the target instruction and each first instruction that has been allocated in the sequence when allocating the target instruction to the sequence, and to obtain the calculation result of each iteration when running the target instruction. The target instruction is any instruction that has not been allocated, and the sequence includes multiple first instructions that do not depend on each other and are executed in parallel. If the calculation result indicates that the target instruction depends on any first instruction, the dependency between the target instruction and each second instruction that has been allocated in the next sequence is analyzed until the calculation results of multiple iterations all indicate that the target instruction does not depend on the instructions that have been allocated in the target sequence, and the target instruction is allocated to the target sequence. If the calculation results of multiple iterations all indicate that the target instruction does not depend on any first instruction, the target instruction is allocated to the sequence, and the allocation of the next instruction that has not been allocated continues until the instructions corresponding to the optimized sub-access operations are allocated, and the instruction scheduling result is obtained.
[0174] Optionally, the allocation module 704 is further configured to, after the Nth iteration, determine the memory address corresponding to the target instruction when the addressing mode of the target instruction is base addressing, wherein the memory address includes the starting address and the vector length, the starting address being determined based on the base address and the offset, and N being a positive integer; if the memory address corresponding to the target instruction does not overlap with the memory address corresponding to the first instruction, use the target instruction independently of the first instruction as the result of the Nth iteration when running the target instruction; if the memory address corresponding to the target instruction overlaps with the memory address corresponding to the first instruction, use the target instruction dependent on the first instruction as the result of the Nth iteration when running the target instruction.
[0175] Optionally, the allocation module 704 is further configured to: recursively analyze each operand in the target instruction to obtain the expression characteristics of each operand when the addressing mode of the target instruction is not base addressing; continue to recursively analyze each sub-operand in the target operand to obtain the expression characteristics of each sub-operand when the expression characteristics of any operand are not base addresses, until the expression characteristics of each sub-operand are all base addresses; indicate operands whose expression characteristics are not base addresses in the target operand; and determine that the addressing mode of the target instruction is base addressing when the expression characteristics of each operand are all base addresses.
[0176] Optionally, the compiler optimization device 70 further includes a prediction module 705; the prediction module 705 is used to perform loop variable analysis on the instructions corresponding to the optimized multiple sub-access operations to obtain the iteration number of each of the optimized multiple sub-access operations; if the loop execution number is predicted, the loop execution number is used as the iteration number; if the loop execution number is not predicted, the iteration number is set to a preset number.
[0177] Optionally, the compiler optimization device 70 further includes a conversion module 706; the conversion module 706 is used to convert the base address of the pointer in the access operation to obtain the conversion result; if the conversion result indicates that the base address is a constant, the base address is replaced with a constant and then the compiler optimization process is executed.
[0178] Optionally, the compilation optimization result includes multiple optimized sub-access operations; the optimization module 703 is also used to perform instruction conversion on the multiple optimized sub-access operations to generate machine instructions; to perform instruction scheduling on the execution order of the machine instructions to obtain instruction scheduling results; the instruction scheduling results include the memory addresses corresponding to the machine instructions; and to optimize the memory addresses that are accessed repeatedly in the instruction scheduling results to obtain the instructions corresponding to each of the multiple optimized sub-access operations.
[0179] It should be noted that the compiler optimization device provided in the above embodiments is only illustrated by the division of the above functional modules when optimizing the compiler. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0180] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.
[0181] The compiler optimization apparatus and compiler optimization method embodiments provided in the above embodiments belong to the same concept. The specific working process and technical effects of the units and modules in the above embodiments can be found in the method embodiment section, and will not be repeated here.
[0182] Based on the compiler optimization method provided in the above embodiments, FIG8 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. As shown in FIG8, the computer device 80 includes: a processor 801, a memory 802, and a computer program 803 stored in the memory 802 and executable on the processor 801. When the processor 801 executes the computer program 803, it implements the steps in the compiler optimization method in the above embodiments.
[0183] Computer device 80 can be a general-purpose computer device or a special-purpose computer device. In specific implementations, computer device 80 can be a desktop computer, portable computer, network server, handheld computer, mobile phone, tablet computer, wireless terminal device, communication device, or embedded device. The embodiments of this application do not limit the type of computer device 80. Those skilled in the art will understand that FIG8 is merely an example of computer device 80 and does not constitute a limitation on computer device 80. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0184] Processor 801 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0185] In some embodiments, memory 802 may be an internal storage unit of computer device 80, such as a hard disk or RAM of computer device 80. In other embodiments, memory 802 may be an external storage device of computer device 80, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on computer device 80. Furthermore, memory 802 may include both internal and external storage units of computer device 80. Memory 802 is used to store operating systems, applications, boot loaders, data, and other programs. Memory 802 may also be used to temporarily store data that has been output or will be output.
[0186] This application also provides a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.
[0187] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the various method embodiments described above.
[0188] This application provides a computer program product that, when run on a computer, causes the computer to perform the steps described in the various method embodiments above.
[0189] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above method embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices. The computer-readable storage medium mentioned in this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.
[0190] It should be understood that all or part of the steps of the above embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-described computer-readable storage medium.
Claims
1. A compiler optimization method, characterized in that, The method includes: Obtain the vector length corresponding to the access operation in the compiled language; the vector length is related to the number of elements included in the memory address to be accessed; If the length of the vector corresponding to the access operation is not a power of a preset value but is greater than a preset threshold, the access operation is split according to the vector length to obtain multiple sub-access operations. Based on the multiple sub-access operations, the repetitive operations in the compiled language are optimized to obtain the compilation optimization result.
2. The method as described in claim 1, characterized in that, The plurality of sub-access operations include at least a first sub-access operation and a second sub-access operation, wherein the vector length corresponding to the first sub-access operation is greater than half of the vector length corresponding to the access operation, and the vector length corresponding to the second sub-access operation is less than or equal to the preset threshold.
3. The method as described in claim 1, characterized in that, The access operations include load operations and / or store operations; When the length of the vector corresponding to the access operation is not a power of a preset value but is greater than a preset threshold, the access operation is split according to the vector length to obtain multiple sub-access operations, including: If the length of the first vector corresponding to the loading operation is not a power of the preset value but is greater than the preset threshold, the loading operation is split according to the length of the first vector to obtain multiple sub-loading operations; the length of the first vector is related to the number of elements included in the memory address to be read. And / or, if the length of the second vector corresponding to the storage operation is not a power of the preset value but is greater than the preset threshold, the storage operation is split according to the length of the second vector to obtain multiple sub-storage operations; the length of the second vector is related to the number of elements included in the memory address to be written; The plurality of sub-access operations are the plurality of sub-load operations and / or the plurality of sub-store operations.
4. The method as described in claim 3, characterized in that, The step of optimizing repetitive operations in the compiled language based on the multiple sub-access operations to obtain a compilation optimization result includes: The multiple sub-loading operations are merged to obtain a merged operation; the loading operations are deleted, and the repetitive operations in the compiled language are optimized based on the merged operation to obtain the compilation optimization result; Alternatively, the storage operation can be deleted, and the repetitive operations in the compiled language can be optimized based on the multiple sub-storage operations to obtain the compilation optimization result; Alternatively, the multiple sub-load operations can be merged to obtain a merged operation; the load operation and the store operation can be deleted, and the repetitive operations in the compiled language can be optimized based on the merged operation and the multiple sub-store operations to obtain the compilation optimization result.
5. The method according to any one of claims 1-4, characterized in that, The compilation optimization result includes multiple optimized sub-access operations; After optimizing repetitive operations in the compiled language based on the multiple sub-access operations to obtain the compilation optimization result, the method further includes: Obtain the iteration count of the instruction corresponding to each of the optimized sub-access operations; When assigning a target instruction to a sequence, the dependency between the target instruction and each first instruction that has been assigned in the sequence is analyzed to obtain the computation result of each iteration when running the target instruction; the target instruction is any instruction that has not been assigned, and the sequence includes multiple first instructions that are independent of each other and are executed in parallel; If the calculation result indicates that the target instruction depends on any first instruction, the dependency analysis between the target instruction and each second instruction that has been allocated in the next sequence continues until the calculation results of multiple iterations all indicate that the target instruction does not depend on the instructions that have been allocated in the target sequence, and the target instruction is allocated in the target sequence. If the results of multiple iterations all indicate that the target instruction does not depend on the first instructions, the target instruction is allocated in the sequence, and the allocation of the next unallocated instruction continues until the instructions corresponding to the optimized sub-access operations are allocated, thus obtaining the instruction scheduling result.
6. The method as described in claim 5, characterized in that, The method further includes: After the Nth iteration, if the addressing mode of the target instruction is base addressing, the memory address corresponding to the target instruction is determined. The memory address includes the starting address and the vector length. The starting address is determined based on the base address and the offset, where N is a positive integer. If the memory address corresponding to the target instruction does not overlap with the memory address corresponding to the first instruction, the target instruction is used as the result of the Nth iteration when the target instruction is executed, independent of the first instruction. If the memory address corresponding to the target instruction overlaps with the memory address corresponding to the first instruction, the target instruction is used as the result of the Nth iteration when the target instruction is executed, depending on the first instruction.
7. The method as described in claim 6, characterized in that, The method further includes: If the addressing mode of the target instruction is not the base address addressing mode, recursively analyze each operand in the target instruction to obtain the expression characteristics of each operand. If any operand's representation feature is not the base address, the recursive analysis continues for each sub-operand in the target operand to obtain the representation features of each sub-operand until the representation features of each sub-operand are all base addresses; the target operand indicates the operand whose representation feature is not the base address. When all operands are represented by base addresses, the addressing mode of the target instruction is determined to be base addressing.
8. The method according to any one of claims 1-4, characterized in that, The compilation optimization result includes multiple optimized sub-access operations; After optimizing repetitive operations in the compiled language based on the multiple sub-access operations to obtain the compilation optimization result, the method further includes: The optimized sub-access operations are converted into machine instructions. The execution order of the machine instructions is scheduled to obtain an instruction scheduling result; the instruction scheduling result includes the memory address corresponding to the machine instruction. The repeatedly accessed memory addresses in the instruction scheduling result are optimized to obtain the instructions corresponding to each of the optimized sub-access operations.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.