An Automatic Optimization Method for Mixed-Precision Operators with Controllable Errors
Through the basic framework of LLVM compiler, the operator is automatically inserted into the deep learning compiler for matrix quantization and accuracy recovery, which solves the problem of low computational efficiency of matrix multiplier in the existing technology, and realizes the acceleration of matrix multiplier quantization of user-controllable errors, which significantly improves the computing performance.
Patent Information
- Application Number
- CN202111551663.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing deep learning compilers and hardware for hybrid precision computing are difficult to improve the computational efficiency of matrix multiplier in non-neural network programs, and lacks the automation and user-controllable error-based matrix multiplier quantization acceleration method.
It provides a hybrid precision operator automation optimization method with controllable errors. Through the basic framework of the LLVM compiler, operators are automatically inserted before and after matrix operation function call for quantization and accuracy recovery, realizing the automated optimization of hybrid precision GEMM operators.
When the user provides error thresholds and operator function names, the matrix multiplication operation is automatically optimized, which improves the calculation efficiency and significantly improves the performance of matrix multiplication operation in the program.
Smart Images

Figure CN114217817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of compilers, and in particular to an automated optimization method for mixed-precision operators with controllable errors. Background Art
[0002] The effectiveness of deep learning in image processing and natural language processing tasks has led to the development of more and more applications. Existing deep learning compilers and hardware for mixed-precision computing are difficult to be used by non-neural network programs to improve the computing efficiency of matrix multiplication operators. Moreover, there is currently no automated, user-controllable error matrix multiplication operator quantization acceleration method under general programs. Summary of the Invention
[0003] In order to solve the above technical problems, the purpose of the present invention is to provide an automated optimization method for mixed-precision operators with controllable errors, and to construct an automated compiler framework. As long as the user provides the matrix operation error threshold and the function name of the general matrix multiplication (GEMM) operator in the matrix library, LLVM can automatically insert operators before and after the function call to quantize the input matrix and restore the accuracy of the output matrix, so that the mixed-precision GEMM operator can help improve performance in more complex and extensive high-performance computing programs.
[0004] LLVM represents a set of compiler infrastructure frameworks, including a series of modular compiler components and toolchains. It compiles source code into an executable program, and at the same time can analyze all information in the source code and provide developers with functional interfaces to modify the code.
[0005] The first technical solution adopted by the present invention is: an automated optimization method for mixed-precision operators with controllable errors, including the following steps:
[0006] Obtain the source code;
[0007] The first optimization: Based on the source code, under the given input, optimize the general matrix multiplication function through the ID conversion method, and record the running order of the general matrix multiplication function and its corresponding dimensions, to obtain an optimized general matrix multiplication function ID list and a general matrix multiplication function order list;
[0008] The second optimization: According to the optimized general matrix multiplication function ID list and the general matrix multiplication function order list, transfer the input matrix of each unoptimized general matrix multiplication call function to the fast matrix multiplication function with different parameter combinations preset for performance timing and error calculation, to obtain the fast matrix multiplication function under the best parameter combination that replaces the original general matrix multiplication function;
[0009] The third optimization is to analyze the obtained performance data and select the best parameter combination for function replacement. Replace the single-precision matrix multiplication algorithm with a mixed-precision matrix multiplication operator and output the optimized code.
[0010] Furthermore, the specific steps of the first optimization include:
[0011] Generate the corresponding LLVM bitcode file according to the source code;
[0012] Scan the instructions of the LLVM IR and convert the general matrix multiplication call function into a wrapper function with a unique ID;
[0013] IR represents a set of intermediate languages provided by LLVM that are suitable for all language compiler systems. All languages can be transformed into the same intermediate language, and the compiler can implement general optimizations based on this intermediate language.
[0014] Output the ID of the wrapper function to the function call ID file before the next instruction of the wrapper function;
[0015] Output the matrix scale of the general matrix multiplication call function to the matrix dimension record file before the next instruction of the wrapper function;
[0016] Generate an executable file and run it with the given input to obtain a list of optimized general matrix multiplication GEMM function IDs.
[0017] Furthermore, the step of generating the corresponding LLVM bitcode file according to the source code specifically includes:
[0018] Process all compilation units in the source code based on clang - emit - llvm to generate corresponding multiple LLVM bitcode files;
[0019] Connect all LLVM bitcode files to generate an overall bitcode file.
[0020] Furthermore, the specific steps of the second optimization include:
[0021] Read and organize the function call ID file to obtain the organized function call ID file;
[0022] Record the ID of the target function to be replaced and its test effect of the fast matrix multiplication function under different parameter combinations to obtain a list of IDs of the target functions to be replaced;
[0023] Traverse each ID in the organized function call ID file in order;
[0024] Perform IR traversal on the overall bitcode file;
[0025] When the target general matrix multiplication function is found and its ID is in the list of target function IDs to be replaced, it is replaced with the fast matrix multiplication function under the parameter combination with the best performance under the user-preset error threshold;
[0026] When the target general matrix multiplication function is found and its ID appears in the function call ID list but does not exist in the list of target function IDs to be replaced, a timing function is inserted before and after it, and the timing result and ID are output to the optimization data record file;
[0027] Replace the general matrix multiplication function with a fast matrix multiplication function with multi-layer loop nesting;
[0028] Insert a timer and call the error calculation function to calculate the average error and minimum error between the output matrix of its fast matrix multiplication function and the result matrix of the original target function;
[0029] Output the current target function ID, current parameter combination, running time of the current fast matrix multiplication function, and error result to the optimization data record file;
[0030] Generate an executable file according to the overall bitcode file and run it under the given input to obtain the fast matrix multiplication function under the best parameter combination for each replacement of the original general matrix multiplication function.
[0031] Furthermore, the specific steps of the third optimization include:
[0032] Read the matrix dimension record file, calculate the maximum required additional memory space, and generate a memory pool;
[0033] Read the optimization data record file, read the data of each function ID in order, calculate the parameter combination with the best performance under the user-given error threshold, and replace the original target function corresponding to the current function ID with the fast matrix multiplication function under the specific parameter combination to obtain the optimized code;
[0034] Generate an executable file according to the optimized code.
[0035] The beneficial effects of the method and system of the present invention are: The present invention can enable any program that calls a specific GEMM operator to automatically complete matrix quantization and precision recovery when the user provides information such as specified precision and operator function name, improve development efficiency, give full play to the mixed-precision calculation ability of GPGPU, and significantly improve the efficiency of matrix multiplication operations in the program while controlling the degree of precision loss of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the flowchart of the steps of an automatic optimization method for a mixed-precision operator with controllable error of the present invention;
[0037] Figure 2 It is a schematic flow chart of the first optimization of a specific embodiment of the present invention;
[0038] Figure 3 It is a schematic flow chart of the second optimization of a specific embodiment of the present invention;
[0039] Figure 4 It is a schematic flow chart of the third optimization of a specific embodiment of the present invention. Specific Embodiments
[0040] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0041] Refer to Figure 1 , the present invention provides an automated optimization method for a mixed-precision operator with controllable error, and the method includes the following steps:
[0042] S0. Obtain the source code;
[0043] S1. First optimization: Based on the source code, under a given input, optimize the general matrix multiplication function through the ID conversion method, and record the running order of the general matrix multiplication function and its corresponding dimensions, to obtain an optimized general matrix multiplication function ID list and a general matrix multiplication function order list;
[0044] S2. Second optimization: According to the optimized general matrix multiplication function ID list and the general matrix multiplication function order list, transfer the input matrix of each unoptimized general matrix multiplication call function to a fast matrix multiplication function with different preset parameter combinations for performance timing and error calculation, to obtain a fast matrix multiplication function under the best parameter combination for each replacement of the original general matrix multiplication function;
[0045] S3. Third optimization: Analyze the obtained performance data and select the best parameter combination for function replacement, replace the single-precision matrix multiplication algorithm with a mixed-precision matrix multiplication operator, and output the optimized code.
[0046] Optimization pass one, refer to Figure 2 :
[0047] 1. Use clang-emit-llvm to generate corresponding multiple LLVM bitcode files for all compilation units;
[0048] 2. Use llvm-link to link all the LLVM bitcode files to generate a single LLVM bitcode file for the overall program;
[0049] 3. Use the LLVM optimization tool opt and the first optimization pass shared library to optimize the bitcode file obtained in the second step;
[0050] 4. Scan each instruction of the entire module of the LLVM IR;
[0051] 5. When encountering each general matrix multiplication call function instruction, replace it with a wrapper function with a unique ID (e.g., starting from ID 0, increment the ID by 1 for each encountered general matrix multiplication call function);
[0052] 6. Insert code before the next instruction of the wrapper function to output the ID of the wrapper function to the function call ID file;
[0053] 7. Insert code before the next instruction of the wrapper function to output the matrix sizes M, N, K of the general matrix multiplication call function into the matrix dimension record table;
[0054] 8. Use llc to generate the final executable file and run it with the given input.
[0055] Optimization pass two, refer to Figure 3 :
[0056] 1. Read the function call ID file generated by optimization pass 1, organize it into an ID list that follows the original order and contains no duplicates, and overwrite the original file;
[0057] 2. Create a file to record the list of target function IDs to be replaced by the fast matrix multiplication function under different parameter combinations that have been tried and their test effects (including performance time and error analysis results);
[0058] 3. Use a script to traverse each ID in the function call ID file from front to back;
[0059] 4. Use the LLVM optimization tool opt to perform an IR traversal on the overall bitcode file;
[0060] 5. Whenever encountering a target general matrix multiplication function and its ID appears in the list of target function IDs to be replaced in the second step, replace it with the fast matrix multiplication function under the parameter combination with the best performance under the user-specified error threshold;
[0061] 6. When encountering a target general matrix multiplication function and its ID appears in the function call ID list in the first step but does not exist in the list of target function IDs to be replaced, insert timing functions before and after it, and insert code after the last timing function to output the timing results and the ID to the optimization data record file;
[0062] 7. Add the hipMalloc function after the second timing function inserted in the sixth step to allocate a memory space with a size equal to twice the size of its input matrix C, and copy two copies of its input matrix C into this memory. Then add the hipFree function after the hipMalloc function to release this memory space;
[0063] 8. Add a triple loop for adjusting parameters before the hipFree function inserted in the seventh step, with all maximum values being 1. The innermost loop is the fast matrix multiplication function that accepts three parameters, and its input matrix C is the second copy of matrix C in the memory allocated in the seventh step;
[0064] 9. Insert code before the innermost fast matrix multiplication function in the eighth step to copy each element of the first matrix in the backup memory to the corresponding element in the second matrix one by one;
[0065] 10. Insert timing functions before and after the fast matrix multiplication function in the eighth step, and insert code after the latter timing function to call the error calculation function to calculate the average error and maximum error between its result and the result matrix of the original target function;
[0066] 11. Insert code after the error calculation function in the ninth step to output the current target function ID, the current parameter combination, the running time of the current fast matrix multiplication function, and the error result to the optimization data record file;
[0067] 12. Use llc to turn the entire LLVM bitcode file into an executable target file and run it under the specified input;
[0068] 13. Return to the third step until all called function IDs have been traversed;
[0069] Optimization pass 3, refer to Figure 4 :
[0070] 1. Use the LLVM optimization tool opt to process the overall bitcode file;
[0071] 2. Read the matrix dimension record file generated by optimization pass 1, calculate the maximum additional memory space required, and insert the hipMalloc function at the beginning of the main() function to generate a memory pool;
[0072] 3. Modify the signature of the user-defined function to add a parameter for transmitting a memory address pointer, and modify the parameter passing at all its call sites to add a memory pool pointer parameter;
[0073] 4. Traverse all call function instructions to establish a mapping between the function IR address and the function ID for all general matrix multiplication functions;
[0074] 5. Read the optimized data record file, sequentially read the data for each function ID, calculate the parameter combination with the best performance under the error threshold given by the user, and replace the original objective function corresponding to the current function ID with a fast matrix multiplication function under the specific parameter combination;
[0075] 6. Use llc to generate an executable object file and run it with specific inputs.
[0076] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. An automated optimization method for mixed-precision operators with controllable errors, characterized in that, it includes the following steps: Obtain the source code; The first optimization: Based on the source code, under a given input, optimize the general matrix multiplication function through the ID conversion method, and record the running order of the general matrix multiplication function and its corresponding dimensions, to obtain an optimized general matrix multiplication function ID list and a general matrix multiplication function order list; The second optimization: According to the optimized general matrix multiplication function ID list and the general matrix multiplication function order list, transfer the input matrix of each unoptimized general matrix multiplication call function to the fast matrix multiplication function with different preset parameter combinations for performance timing and error calculation, to obtain the fast matrix multiplication function under the best parameter combination for each replacement of the original general matrix multiplication function; The third optimization: Analyze the obtained performance data and select the best parameter combination for function replacement, replace the single-precision matrix multiplication algorithm with a mixed-precision matrix multiplication operator, and output the optimized code; The specific steps of the first optimization include: Generate the corresponding LLVM bitcode file according to the source code; Scan the instructions of the LLVM IR and convert the general matrix multiplication call function into a wrapper function with a unique ID; Output the ID of the wrapper function to the function call ID file before the next instruction of the wrapper function; Output the matrix scale of the general matrix multiplication call function to the matrix dimension record file before the next instruction of the wrapper function; Generate an executable file and run it under a given input to obtain an optimized general matrix multiplication GEMM function ID list; The specific steps of the second optimization include: Read and organize the function call ID file to obtain an organized function call ID file; Record the ID of the target function to be replaced and its test effect of the fast matrix multiplication function under different parameter combinations to obtain a list of IDs of the target functions to be replaced; Traverse each ID in the organized function call ID file in sequence; Perform an IR traversal on the overall bitcode file; When the target general matrix multiplication function is found and its ID is in the list of IDs of the target functions to be replaced, replace it with the fast matrix multiplication function under the best parameter combination with a user-preset error threshold; When the target general matrix multiplication function is found and its ID appears in the function call ID list but does not exist in the list of IDs of the target functions to be replaced, insert a timing function before and after it and output the timing result and ID to the optimization data record file; Replace the general matrix multiplication function with a fast matrix multiplication function with multi-layer loop nesting; Insert a timer and call the error calculation function to calculate the average error and minimum error between the output matrix of the fast matrix multiplication function and the result matrix of the original target function; Output the current target function ID, the current parameter combination, the running time of the current fast matrix multiplication function, and the error result to the optimization data record file; Generate an executable file according to the overall bitcode file and run it under a given input to obtain the fast matrix multiplication function under the best parameter combination for each replacement of the original general matrix multiplication function; The specific steps of the third optimization include: Read the matrix dimension record file, calculate the maximum required additional memory space, and generate a memory pool; Read the optimized data record file, read the data of each function ID in order, calculate the parameter combination with the best performance under the error threshold given by the user, and replace the original objective function corresponding to the current function ID with a fast matrix multiplication function under the specific parameter combination to obtain the optimized code; Generate an executable file according to the optimized code.
2. According to the method for automatically optimizing a hybrid precision operator with controllable error described in claim 1, It is characterized in that The step of generating the corresponding LLVM bitcode file according to the source code specifically includes: Process all compilation units in the source code based on clang-emit-llvm to generate corresponding multiple LLVM bitcode files; Connect all LLVM bitcode files to generate an overall bitcode file.
Citation Information
Patent Citations
Method for automatically generating dense matrix multiplication assembly code based on x86 architecture
CN102750150A
Code protection method, code protection device, storage medium and electronic equipment
CN110059456A