Performance optimization method and device of graphics processor, electronic equipment and storage medium

By collecting and analyzing the basic block execution data of the graphics processor, combining its microarchitecture characteristics, optimizing control flow and resource allocation, the inefficiency and resource imbalance at the microarchitecture level of the graphics processor are solved, and performance improvement and energy consumption reduction are achieved.

CN120339035APending Publication Date: 2025-07-18JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510397778.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Graphics processors have problems with inefficiency in control flow and resource imbalance at the microarchitecture level, which limits further improvements in their performance.

Method used

The basic block execution data is collected through instrumentation logic compiled based on static library, analyzed performance characteristics, and established quantitative mapping relationships to optimize control flow and resource allocation.

Benefits of technology

It improves the computing efficiency of the graphics processor, optimizes resource utilization, reduces energy consumption, and improves the overall performance of the program.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339035A_ABST
    Figure CN120339035A_ABST
Patent Text Reader

Abstract

The invention discloses a performance optimization method and device of a graphics processor, electronic equipment and a storage medium, and relates to the technical field of computers, the performance optimization method comprises the steps that basic block execution data of a target program can be dynamically collected based on instrumentation logic compiled by a static library, and performance bottlenecks are accurately positioned by analyzing the data; meanwhile, the control flow and resource allocation are optimized in combination with the characteristics of the graphic processor micro-system structure, so that the technical problems of low efficiency of the control flow, resource imbalance, low memory access efficiency and the like of the graphic processor micro-system structure level can be solved; the technical effects of improving the calculation efficiency of the graphics processor, optimizing the resource utilization rate, reducing the energy consumption and improving the overall performance of the program are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for optimizing the performance of a graphics processing unit, an electronic device, and a storage medium. Background Art

[0002] With its excellent computing power and efficient parallel processing mechanism, the Graphics Processing Unit (GPU) is a key technical support in the field of compute-intensive applications. The optimization of GPU performance not only depends on its macroscopic architecture design, but is also closely related to its microarchitecture. At present, GPU technology has defects such as inefficient control flow and resource imbalance at the microarchitecture level, which limit the further improvement of GPU performance. Summary of the Invention

[0003] The present application provides a method, apparatus, electronic device, and storage medium for optimizing the performance of a graphics processing unit, so as to at least solve the problem that defects such as inefficient control flow and resource imbalance at the microarchitecture level in the related art limit the performance of the GPU.

[0004] The present application provides a method for optimizing the performance of a graphics processing unit, including:

[0005] Running a target program, and collecting basic block execution data based on a instrumentation logic; wherein, the instrumentation logic is compiled from a static library;

[0006] Analyzing the performance characteristics of the target program through the basic block execution data, and establishing a quantitative mapping relationship between the performance characteristics and the microarchitecture of the graphics processing unit;

[0007] Performing performance optimization on the graphics processing unit based on the quantitative mapping relationship.

[0008] The present application also provides an apparatus for optimizing the performance of a graphics processing unit, including:

[0009] A collection unit, configured to run a target program and collect basic block execution data based on a instrumentation logic; wherein, the instrumentation logic is compiled from a static library;

[0010] A establishment unit, configured to analyze the performance characteristics of the target program through the basic block execution data, and establish a quantitative mapping relationship between the performance characteristics and the microarchitecture of the graphics processing unit;

[0011] An optimization unit, configured to perform performance optimization on the graphics processing unit based on the quantitative mapping relationship.

[0012] The present application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned performance optimization methods of the graphics processor when executing the computer program.

[0013] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above-mentioned performance optimization methods of the graphics processor when executed by a processor.

[0014] The present application also provides a computer program product including a computer program, and the computer program implements the steps of any of the above-mentioned performance optimization methods of the graphics processor when executed by a processor.

[0015] Through the present application, since the instrumentation logic compiled based on the static library can dynamically collect the basic block execution data of the target program, accurately locate the performance bottleneck by analyzing this data, and at the same time, combine the characteristics of the graphics processor microarchitecture to optimize the control flow and resource allocation, it is possible to solve technical problems such as inefficient control flow, resource imbalance, and low memory access efficiency at the graphics processor microarchitecture level, and achieve the technical effects of improving the computing efficiency of the graphics processor, optimizing resource utilization, reducing energy consumption, and improving the overall performance of the program. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of a method for optimizing the performance of a graphics processor provided by an embodiment of the present disclosure;

[0018] Figure 2 It is a schematic structural diagram of a device for optimizing the performance of a graphics processor provided by an embodiment of the present disclosure;

[0019] Figure 3 It is a schematic structural diagram of a device for optimizing the performance of a graphics processor provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0021] It should be noted that in the description of this application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail in conjunction with the accompanying drawings and specific implementation manners.

[0023] The following describes a method, apparatus, electronic device, and storage medium for optimizing the performance of a graphics processor according to embodiments of the present disclosure with reference to the accompanying drawings.

[0024] Figure 1 It is a schematic flowchart of a method for optimizing the performance of a graphics processor provided by an embodiment of the present disclosure.

[0025] As Figure 1 shown, the method includes the following steps:

[0026] Step 101, run the target program, and collect basic block execution data based on the instrumentation logic; wherein, the instrumentation logic is compiled from a static library.

[0027] In some embodiments, the instrumentation logic is compiled according to a static library in accordance with specific instrumentation strategies and requirements, and includes code logic for collecting basic block execution data.

[0028] In some embodiments, the operating environment also needs to meet requirements, including the dependent libraries and tools required for the normal operation of the target program and the instrumentation logic. When the target program depends on third-party libraries, ensure that these third-party libraries are correctly installed and the environment variables are configured correctly.

[0029] During the running of the target program, the instrumentation logic will automatically collect basic block execution data. The collected data usually includes information such as the identifier, execution count, and execution time of the basic block. These data will be stored in a specified file or database. The storage format can be a text file, a comma-separated values file (CSV file), a binary file, etc., which specifically depends on the configuration of the instrumentation logic. This application embodiment does not limit this.

[0030] Step 102: Perform data analysis on the performance characteristics of the target program through the basic blocks, and establish a quantitative mapping relationship between the performance characteristics and the microarchitecture of the graphics processing unit.

[0031] In some embodiments, before performing step 102, data preprocessing may be performed on the data of the basic blocks, and cleaning work may be carried out on the data of the basic blocks to remove noise, errors, and duplicate data. For example, when the execution time appears negative or an abnormally large value, it needs to be corrected or excluded; then the data is converted into a format convenient for analysis, such as unifying the data type and organizing the data structure. When performing performance characteristic extraction, statistical analysis methods may be used, such as calculating statistics such as the mean, variance, and median of the basic block execution time. Machine learning algorithms such as decision trees and neural networks may also be used to model the data to predict program performance; in practical applications, the analysis method may be determined according to actual needs, and the embodiments of the present application do not limit this.

[0032] Identify the performance characteristics of the target program, analyze the distribution of the basic block execution time, find the basic blocks with longer execution times, and determine the possible performance bottlenecks; according to the execution order and frequency of the basic blocks, analyze the control flow and data flow characteristics of the program, and judge the parallelism and locality of the program; finally, establish a quantitative mapping relationship. In some embodiments, select appropriate quantitative indicators, such as the proportional relationship between the basic block execution time and the GPU cache hit rate; construct a mathematical model, such as a linear regression model, a neural network model, etc., to describe the relationship between the performance characteristics and the microarchitecture parameters; use experimental data to train and verify the model, and adjust the model parameters to improve the accuracy of the mapping relationship.

[0033] Step 103: Optimize the performance of the graphics processing unit based on the quantitative mapping relationship.

[0034] In some embodiments, through the quantitative mapping relationship, identify the key microarchitecture factors affecting program performance, such as problems like too high cache miss rate, memory bandwidth bottleneck, or insufficient utilization of parallel computing units. After clarifying these key factors, formulate corresponding optimization strategies for different performance bottlenecks. For example, if the quantitative mapping relationship shows that the memory access latency is the main factor affecting performance, then redesign the data layout to make the data in memory more compliant with the GPU cache mechanism and reduce the cache miss situation.

[0035] In some embodiments, the code of the target program may also be optimized specifically according to the quantitative mapping relationship. For example, reduce unnecessary calculations and memory operations to improve the execution efficiency of the code; adopt more efficient algorithms and data structures to reduce the computational complexity, etc. It should be noted that this narrative method is only an exemplary illustration and is not a specific limitation on the specific performance optimization method.

[0036] In some embodiments, before running the target program and collecting basic block execution data based on the instrumentation logic, the following steps are further included:

[0037] Obtain the target compilation tool, and compile the source code based on the target compilation tool and the definition rules to obtain the static library;

[0038] To build the static library libnvbit.a required for a microarchitecture-independent basic block analysis tool, first, the source code directory of the tool needs to be located. This directory contains all the source code files, header files, and compilation scripts required to implement the basic block analysis function. After locating the source code directory, determine the build configuration file responsible for compiling the static library. The configuration file is usually Makefile or CMakeLists.txt, which defines the tasks to be executed during compilation, the dependencies, and the options for the compiler and linker. In some embodiments, if the Makefile is included in the source code directory, the make command can be used to compile the static library. The make command will automatically handle the compilation, linking, etc. of the source code according to the rules defined in the Makefile, and finally generate the target file libnvbit.a. When executing the make command, various information during compilation is usually displayed, including the files being compiled, the compiler options used, and any possible compilation errors.

[0039] In another implementable manner of the embodiments of the present application, if the build configuration file in the source code directory is CMakeLists.txt, the compilation process is to use the cmake command to configure the project and generate a Makefile or other build system files that match the current system environment. In the cmake command, the source code directory and the target build directory can be specified, as well as any necessary CMake options. After configuration, the make command can be used to compile the static library libnvbit.a just like dealing with a normal Makefile. Regardless of which build system is adopted, compiling the static library libnvbit.a is one of the important steps to ensure the correct operation of the microarchitecture-independent basic block analysis tool. This library file contains all the code and data required to implement the basic block analysis function and is the basis for subsequent development of custom plugins or extension of tool functions.

[0040] Compile the source code based on the application programming interface of the static library to obtain a dynamic link library.

[0041] The source code will utilize the application programming interface (API) and functions provided by libnvbit.a to implement specific analysis logic or operations. After completing the writing or modification of the plug-in source code, use appropriate compilation tools to compile the source code into an executable file or a library file. For NVBit-based plug-in development, use the compilation tools provided by NVBit, or you can also use standard compilation commands (such as gcc or g++) to compile the plug-in code.

[0042] The goal of the compilation process is to generate a dynamic link library (.so file) so that the plug-in can be dynamically loaded into the main program at runtime, thereby expanding the functionality of the main program. During compilation, it is necessary to ensure that the libnvbit.a library is linked so that the plug-in can call the APIs and functions it provides. In addition, it is also necessary to pay attention to setting the correct compiler options and linker options to ensure that the generated dynamic link library is compatible with the main program and can interact correctly with the main program.

[0043] In some embodiments, before running the target program and collecting basic block execution data based on the preset instrumentation logic, the following steps are further included:

[0044] Insert the dynamic link library at the function entry of the target program and at the start position of the basic block;

[0045] In some embodiments, for the insertion at the function entry, it is necessary to locate the start position of each function and then add code to call the initialization function of the dynamic link library. This initialization function will be called when the function starts to execute. When inserting the dynamic link library at the start position of the basic block, it is necessary to analyze the control flow of the program to determine the start point of each basic block. A basic block is a series of instructions executed sequentially in the program without branches and jumps. Insert code to call the data collection function in the dynamic link library at the start position of the basic block. When the program executes to this basic block, the data collection operation will be triggered.

[0046] In the performance analysis and optimization process of CUDA programs (Compute Unified Device Architecture Programs), instrumentation technology is a commonly used method. It allows developers to insert additional code (i.e., instrumentation logic) into the execution flow of the program to collect information about program behavior. To achieve basic block counting and analysis independent of microarchitecture, developers need to write instrumentation logic at appropriate positions in CUDA programs (parallel computing platforms and programming models). The core task of this instrumentation logic is to call functions provided by the microarchitecture-independent basic block plug-in. These functions encapsulate the specific implementation of basic block counting and analysis. Through them, developers can obtain the execution status and statistical information of each basic block in the CUDA program. The writing of instrumentation logic requires precise control of the insertion position to ensure that the boundaries and execution paths of basic blocks can be accurately captured.

[0047] Generally, instrumentation logic is inserted at the function entry, the start of a basic block, or specific instructions in the CUDA program. Inserting instrumentation logic at the function entry can capture function call situations, including the number of calls and call times, etc. Inserting instrumentation logic at the start of a basic block can accurately count the execution times and execution times of each basic block, which is very helpful for analyzing the control flow and hot code of the program. In addition, according to the analysis requirements, developers can also insert instrumentation logic at specific instructions to collect more detailed execution information. When writing instrumentation logic, developers need to ensure that the instrumentation code has as little impact on the original program as possible, while ensuring that enough information can be collected for subsequent analysis. Therefore, choosing the appropriate instrumentation position and writing efficient instrumentation code are the keys to achieving basic block counting and analysis independent of microarchitecture.

[0048] Set environment variables and establish the link relationship between the environment variables and the dynamic link library; so that when the target program runs, the internal functions of the dynamic link library are called to perform data collection on the basic blocks.

[0049] Environment variables are variables used to store configuration information in the operating system. By setting environment variables, the system can know the location and loading method of the dynamic link library. In different operating systems, the methods of setting environment variables are different. LD_PRELOAD is an environment variable widely used in Linux systems. It allows users to specify one or more shared libraries (.so files), and these libraries will be preferentially loaded by the dynamic linker (ld.so) when the program starts. By setting LD_PRELOAD, users can ensure that specific libraries are loaded before the program uses any other shared libraries, thus having the opportunity to overwrite or extend the original function implementation in the program.

[0050] In the context of the CUDA program, set the LD_PRELOAD environment variable and point it to the compiled instrumentation plugin dynamic link library. The instrumentation plugin usually contains function implementations for basic block counting and analysis. Through the LD_PRELOAD mechanism, these functions can be dynamically injected and executed in the main function or other key locations of the CUDA program. Specifically, the user needs to set the LD_PRELOAD variable in the shell environment before running the CUDA program to point to the .so file containing the instrumentation logic. For example, in the bash shell, you can use the following command:

[0051] export LD_PRELOAD= / path / to / your / plugin.so

[0052] When a CUDA program is executed, the dynamic linker will first load and initialize the specified instrumentation plugin library. Before the main function of the CUDA program or any function called through dynamic linking is executed, the initialization code and function implementation in the instrumentation plugin are ready. Therefore, when the CUDA program executes to the instrumentation point (that is, the code location where the instrumentation logic is inserted), it can successfully call the function provided by the instrumentation plugin to achieve basic block counting and analysis.

[0053] When the target program is running, the system will find the dynamic link library according to the set environment variables and load it into the memory. When the program executes to the function entry or the beginning of the basic block where the dynamic link library call code is inserted, the internal function of the dynamic link library will be called. These internal functions will perform data collection operations on the basic blocks. The collected data may include the number of executions of the basic blocks, the execution time, the memory address accessed, and other information. The collected data can be stored in a file or database for subsequent analysis and processing.

[0054] In some embodiments, when installing the microarchitecture-independent basic block feature mining tool and its dependencies to support the analysis and optimization of CUDA programs, the first step is to ensure the availability of the CUDA development environment. First, install the NVIDIA CUDA Toolkit, which not only provides the nvcc compiler required to compile CUDA programs, but also contains the library files required to run CUDA applications. The installation of the CUDA Toolkit is crucial for subsequent steps because it is the basis for GPU accelerated computing and application development.

[0055] To ensure that key tools such as the CUDA Toolkit can be correctly recognized and used by the system during compilation and operation, their installation paths need to be added to the system's Path Environment Variable (PATH). The PATH environment variable is a list of directories used by the operating system to search for executable files. By modifying this variable, users can include the paths of custom tools or third-party software, enabling these tools to be globally accessible from the command line.

[0056] In addition, for basic block analysis tools that are independent of the microarchitecture, a specific environment variable is set. This environment variable should point to the root directory of the tool, which contains all the necessary library files, configuration files, scripts, etc. By doing so, it is not only convenient to reference the header files and library files provided by the tool during compilation but also easy to locate the executable file or other resources of the tool during runtime.

[0057] In some embodiments, the performance optimization of the graphics processing unit based on the quantization mapping relationship includes the following steps:

[0058] According to the basic block count result, identify the hot regions in the target program, where the hot regions are basic blocks determined according to the execution count and / or execution time;

[0059] Analyze the control flow complexity of the hot regions, reorganize the code structure, and reduce the levels of nested branches and loops;

[0060] Evaluate the memory access patterns of the hot regions, improve the data storage layout, and reduce the number of memory accesses.

[0061] In some embodiments, the basic block execution data includes at least one of the execution count, execution time, and execution path of the basic block.

[0062] In some embodiments, the hot regions are basic blocks determined according to the execution count and / or execution time. Usually, basic blocks with frequent execution counts or long execution times may constitute hot regions. By conducting a detailed statistical analysis of the basic block count results and setting appropriate thresholds, such as marking basic blocks with an execution count exceeding a certain value or an execution time exceeding a certain range as hot regions. These hot regions are often the locations of program performance bottlenecks, and optimizing them can significantly improve the performance of the entire program.

[0063] After identifying the hotspots, analyze the control flow complexity of the hotspots. A high control flow complexity usually manifests as a large number of nested branches and loop levels, which can increase the execution time and understanding difficulty of the program. First, perform a detailed control flow analysis on the code of the hotspots, draw a control flow graph, and clearly show the branch and loop structures in the code. Then, reorganize the code structure according to the analysis results. Methods such as splitting complex nested branches into multiple simple conditional judgments or merging or unfolding deep loops can be adopted. For example, for multi-layer nested `if-else` statements, consider using `switch-case` statements for replacement to improve the readability and execution efficiency of the code. At the same time, try to avoid unnecessary loop nesting, and reduce the number of loop levels by optimizing algorithms or data structures. This can not only make the code more concise and understandable, but also reduce the execution paths of the program, thereby improving the running speed of the program.

[0064] Evaluate the memory access pattern of the hotspots. Memory access is often a performance bottleneck during program execution. Frequent memory access will slow down the program running speed. Therefore, perform a memory access analysis on the code of the hotspots to understand the data reading and writing patterns, such as whether there are a large number of random memory accesses and whether there are poor data locality situations. Based on the analysis results, improve the data storage layout. Technologies such as data alignment and data prefetching can be adopted to improve the data access efficiency. For example, store the data that is often accessed together in adjacent memory locations to make full use of the cache mechanism and reduce the number of cache misses. In addition, by optimizing algorithms and data structures, reduce unnecessary memory access operations. Use appropriate data structures to avoid repeated memory reads and writes, or adopt cache technology to reduce access to external storage devices. Through these measures, the overhead of memory access can be effectively reduced, and the overall performance of the program can be improved. During the entire optimization process, compare and test the program performance before and after optimization to ensure that the optimization measures can indeed achieve the expected effect. If the optimization effect is not obvious, it is necessary to re-examine the steps of hotspot identification, control flow analysis, and memory access evaluation, and further adjust the optimization strategy.

[0065] In some embodiments, after performing performance optimization on the graphics processor based on the quantization mapping relationship, the method further includes:

[0066] Rerun the target program, and re-collect the execution data of the optimized basic blocks based on the instrumentation logic and

[0067] Compare the execution data of the basic blocks before optimization and the execution data of the optimized basic blocks to evaluate the performance improvement effect.

[0068] When developing CUDA programs, in order to integrate the basic block analysis function independent of the microarchitecture, it is necessary to use NVIDIA's nvcc compiler to compile the CUDA source code. nvcc is a compiler designed specifically for CUDA programs. It can handle CUDA-specific syntax and extensions and generate machine code that can be executed efficiently on NVIDIA GPUs. During the compilation of CUDA programs, it is crucial to ensure that the runtime libraries based on microarchitecture-independent basic blocks and the previously compiled instrumentation plugins are linked. These runtime libraries and plugins provide the core functions required to implement basic block counting and analysis and are the basis for CUDA programs to utilize these advanced analysis features.

[0069] To achieve this, developers need to explicitly specify the options for linking these libraries and plugins in the nvcc compilation command. This usually involves using the -L option to specify the search path for library files and the -l option to specify the library names to be linked. For instrumentation plugins, if they are provided in the form of dynamic link libraries (.so files), they also need to be included in the compilation command through appropriate linker options. In addition, developers need to ensure that the instrumentation logic in the CUDA program can correctly call the functions provided by the instrumentation plugin. This usually means including the header file of the instrumentation plugin in the CUDA program and calling the corresponding functions where basic block analysis needs to be performed.

[0070] When the CUDA program is ready and all necessary environments and dependencies are set up, execute the CUDA program. Whenever the program reaches the location where the instrumentation code is inserted, these instrumentation logics will be automatically triggered. These logics internally call the functions provided by the microarchitecture-independent basic block plugin and are responsible for performing basic block counting and analysis tasks. The execution of the instrumentation logic is transparent and does not interfere with the main logic flow of the CUDA program, but it will collect data on the execution of basic blocks in the background. These data include, but are not limited to, the execution count, execution time of each basic block, and other possible statistical metrics, which together constitute an in-depth analysis of the execution behavior of the CUDA program.

[0071] After the data collection is completed, the instrumentation logic will output the analysis results to the specified location according to the preset requirements. This can be a log file in the file system or the console window, depending on the developer's configuration. Either way, the purpose of the output is to facilitate subsequent analysis work. Developers can use these analysis results to evaluate the performance bottlenecks of CUDA programs, optimize hot code regions, or conduct more in-depth research on program behavior.

[0072] In some embodiments, before running the target program and collecting basic block execution data based on the instrumentation logic, the following steps are further included:

[0073] Generate an executable file based on a preset benchmark set and a preset renderer; wherein, test cases are included in the executable file;

[0074] Running the target program and collecting basic block execution data based on the instrumentation logic includes:

[0075] Run the test cases in the executable file based on the target program, and collect basic block execution data based on the instrumentation logic.

[0076] In some embodiments, the preset benchmark set is the Rodinia benchmark set, the RayBench benchmark set. To compile the benchmark set and the Octane renderer to form the final executable file, it is first necessary to ensure that the necessary compilation tools are installed on the system, such as the GNU Compiler Collection (GCC) or the Clang compiler (Clang), as well as the relevant dependency libraries. For each benchmark set, its source code needs to be downloaded separately and configured according to their respective build instructions. Generally, this includes running a configuration script (such as. / configure) in the source code directory, and then using the make command to compile the source code. For software like the Octane renderer that contains multiple projects and complex dependencies, it is also necessary to use CMake or other build systems to manage the compilation process. Finally, the binary file generated by compilation is the required executable file, which can be directly run for performance testing or rendering tasks.

[0077] After the CUDA program runs and the basic block count results output by the instrumentation logic are collected, the next key step is to deeply analyze these results. The basic block count results provide the detailed execution conditions of each function or basic block in the CUDA program, including key information such as their execution times and execution paths. By carefully analyzing these data, developers can deeply understand the running characteristics and complexity of the CUDA program. Specifically, developers can observe which functions or basic blocks are frequently executed, which indicates the hot spots of the program, that is, where the performance bottlenecks are. At the same time, by comparing the execution times of different functions or basic blocks, developers can also evaluate the workload distribution of each part of the program and identify possible uneven resource allocation problems.

[0078] In addition, it is also very important to compare the basic block count results with the expectations. This helps to verify whether the program is running as expected, or reveals possible logical errors or performance issues in the program. If there are significant differences between the actual execution and the expectations, developers need to further investigate the reasons, which may be improper algorithm implementation, unreasonable data access patterns, or inefficient parallelization strategies, etc. Based on these analysis results, developers can formulate targeted optimization strategies. For example, for functions or basic blocks with excessive execution times or long execution durations, it is possible to consider reducing their execution burden through algorithm optimization, code refactoring, or data parallelization, etc. At the same time, more detailed optimizations can also be carried out for performance bottleneck areas, such as improving the memory access pattern, reducing unnecessary synchronization overhead, or optimizing the control flow structure, etc.

[0079] The number of basic blocks reflects the control flow complexity of the program. If a program has a large number of basic blocks, it indicates that the control flow of the program is relatively complex, containing multiple branches and jump statements. This leads to an increase in branch prediction errors, thus affecting the performance of the program. Basic Block Count can help understand the structure and logic of the program. By analyzing the number and order of basic blocks, the control flow graph of the program can be obtained, and the dependency relationships and execution order between different parts of the program can be understood. This is very helpful for program understanding, debugging, and optimization.

[0080] Please refer to Table 1, which is a schematic table of the number of basic blocks of Rodinia provided by the embodiments of this application. As shown in Table 1:

[0081] Table 1

[0082] Program Name Number of Basic Blocks Program Name Number of Basic Blocks backprop 3 particlefilter 43 hotspot 50 bfs 6 kmeans 14 dwt2d 62 hotspot3D 23 gaussian 27 lavaMD 34 huffman 15 myocyte 702 leukocyte 175 b+tree 49 srad_v1 40 lud 10 nw 64 heartwall 512 streamcluster 16 cfd 130 pathfinder 45

[0083] Based on the statistical results of the number of basic blocks of the Rodinia programs listed in Table 1, the following analysis and conclusions can be drawn. First, the number of basic blocks is one of the important indicators for measuring program complexity and structure. It can be observed from the table that there are significant differences in the number of basic blocks among different Rodinia programs. For example, the backprop program has only 3 basic blocks, while the myocyte program has 702 basic blocks, showing a huge difference between the two. This indicates that there are obvious differences in the structures of different programs, involving different algorithms or implementation methods. Second, there is a certain correlation between the number of basic blocks and the computational complexity and parallelism of the program. A smaller number of basic blocks means the program is relatively simple and the computational tasks are relatively single, while a larger number of basic blocks indicates that the program is more complex, involving more computational tasks and data dependencies. This has important guiding significance for program performance optimization and parallel design. Program designers can adjust the optimization strategy according to the differences in the number of basic blocks and select appropriate optimization methods and parallel patterns for different program structures. In addition, the number of basic blocks can also reflect the control flow and loop structure in the program. For example, programs with a larger number of basic blocks often contain more complex conditional judgments and loop structures, requiring more refined control flow analysis and optimization. This is very important for program performance improvement and maintainability considerations.

[0084] Please refer to Table 2 and Table 3. Table 2 is a schematic table of the number of basic blocks provided by an embodiment of the present application, and Table 3 is another schematic table of the number of basic blocks provided by an embodiment of the present application. As shown in Table 2 and Table 3:

[0085] Table 2

[0086]

[0087] Table 3

[0088]

[0089] In GPU rendering, a basic block is a consecutive sequence of instructions in the compiled code. It has only one entry and one exit, without internal jumps or branches. The number of basic blocks can reflect the complexity and structural characteristics of the program. Table 2 and Table 3 show the number of basic blocks of 100 GPU rendering programs in RayBench. For GPU rendering programs, the number of basic blocks is affected by various factors, including the complexity of the rendering algorithm, the geometric details of the scene, the lighting model, texture mapping, etc. GPU rendering utilizes the parallel processing ability of the graphics processor to accelerate the rendering process of three-dimensional scenes, which involves multiple stages such as geometric transformation, lighting calculation, texture mapping, and depth testing. Each stage corresponds to one or more basic blocks, especially when these stages involve complex calculations or conditional branches.

[0090] Analyzing the basic block sizes of representative GPU rendering programs through examples can better understand the relationship between the number of basic blocks and the complexity of rendering tasks. First, consider a ray tracing renderer based on Perlin noise. Perlin noise is often used to generate natural textures such as clouds and terrain. Ray tracing is a rendering technique that simulates the propagation of light in a scene and calculates its interaction with objects. The number of basic blocks in this rendering core is 103. Since ray tracing involves a large number of random accesses and conditional branches, such as determining whether a ray intersects an object, this program contains more basic blocks to handle these complex logics. Another example is the CUDA-OpenGL-ray-trace program, which has 537 basic blocks. This is a ray tracing rendering program that combines CUDA and OpenGL, and utilizes the parallel computing power of NVIDIA GPUs to implement an efficient rendering algorithm. During the rendering process, complex operations such as parallel task partitioning, data transfer, and memory management need to be handled, which result in the use of a large number of basic blocks. The complexity of the ray tracing algorithm itself, including ray-object intersection detection and material property calculation, also increases the number of basic blocks. Overall, the relatively large number of basic blocks in this program reflects its complexity in parallel computing and graphics rendering.

[0091] Another example is the Mandelbrot-rendering program, which has 98 basic blocks. This program explores the rendering method of the Mandelbrot set and its application in computer graphics. The Mandelbrot set is a typical fractal pattern, and its rendering process involves a large number of iterative calculations and conditional judgments. To accurately determine whether a point belongs to the Mandelbrot set, multiple iterative operations need to be performed and corresponding processing is carried out according to the results. These iterative and conditional branch operations lead to the use of basic blocks. However, compared with other more complex rendering tasks, the calculation of the Mandelbrot set is relatively independent and simple, so the number of basic blocks in this program is relatively small. Finally, consider the Luminous-sphere-render program, which has 258 basic blocks. This program studies a method for rendering a luminous sphere, involving the calculation of lighting models and the simulation of effects such as reflection and refraction of surface materials. These calculations need to be comprehensively processed according to factors such as the geometric shape of the sphere, the position and intensity of the light source, and the surface material properties. To achieve high-quality rendering effects, this program adopts complex lighting models and surface material processing algorithms, which result in the use of multiple basic blocks. If advanced effects such as shadow calculation and ambient occlusion are further considered, the number of basic blocks will increase further. In summary, the number of basic blocks in this program reflects its complexity in lighting calculation and surface material processing.

[0092] Please refer to Table 4. Table 4 is a schematic table of the number of basic blocks of an Octane renderer provided by an embodiment of the present application, as shown in Table 4:

[0093] Table 4

[0094]

[0095] As shown in Table 4, the number of basic blocks in the Octane renderer varies greatly among different cores. Its average value is 187.22, the median is 7.00, the standard deviation is 616.06, and the coefficient of variation is 0.304. This wide difference reflects the complexity and functional diversity of each core. First, this paper observes the cores with a relatively large number of basic blocks. For example, kernel36_s has 2981 basic blocks, far exceeding other cores. This is because kernel36_s is responsible for processing complex rendering tasks, involving a large amount of calculation and operations. For example, it includes complex ray tracing algorithms or texture mapping processes, and these algorithms require a large number of basic blocks to implement. On the contrary, a core with a relatively small number of basic blocks, such as kernel23, only has 3 basic blocks. This means that this core performs simple tasks such as initialization or post-processing operations. In addition, this paper can also observe that some cores have relatively average numbers of basic blocks, such as kernel27 and kernel49_simple, which have 1324 and 401 basic blocks respectively. This means that these cores perform tasks of medium complexity, involving some calculations and operations, but not as complex as kernel36_s. For example, kernel27 is responsible for detecting the intersection of light and objects, which requires certain calculations and logical operations, but is not as complex as some advanced rendering algorithms.

[0096] After obtaining the basic block count results independent of the microarchitecture, developers can optimize the CUDA program targeted according to these data. The goals of the optimization work usually include reducing the number of basic blocks, optimizing the control flow, improving the memory access pattern, etc., in order to improve the execution efficiency and performance of the program. Reducing the number of basic blocks is an effective means of optimization. Basic blocks are the smallest execution units in a program. Reducing the number of basic blocks means reducing the control flow complexity of the program, which helps to simplify the execution path of the program and reduce the instruction cache miss rate, thereby improving performance. Developers can reduce the number of basic blocks by merging similar code blocks, eliminating unnecessary branch judgments, etc.

[0097] Optimizing control flow is also the key to improving the performance of CUDA programs. Reducing the complexity of control flow can reduce instruction prediction errors and improve the execution efficiency of the pipeline. Developers can simplify the control flow by reorganizing the code structure and adopting more efficient loop unrolling or conditional optimization techniques. Improving memory access patterns is particularly important for CUDA programs. Due to the particularity of GPU architecture, the impact of memory access patterns on performance is particularly significant. Developers should pay attention to data locality, reduce memory access conflicts, optimize memory bandwidth utilization, etc. to ensure that data can be efficiently transferred between the CPU and GPU, reduce waiting time, and improve the overall performance of the program. After completing the initial optimization, developers need to repeatedly run the basic block analysis tool based on microarchitecture independence to analyze the optimized CUDA program again. The purpose of this step is to verify the optimization effect, confirm whether the number of basic blocks has been reduced, whether the control flow has been optimized, and whether the memory access pattern has been improved. By comparing the results before and after the analysis, developers can evaluate the effectiveness of the optimization measures and make further adjustments and optimizations as needed.

[0098] Corresponding to the above-mentioned method for optimizing the performance of a graphics processor, the present invention further provides a device for optimizing the performance of a graphics processor. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be described in detail in the present invention.

[0099] Figure 2 A schematic diagram of a performance optimization device for a graphics processor provided by an embodiment of the present disclosure is shown in FIG. Figure 2 As shown, including:

[0100] A collection unit 21 is used to run the target program and collect basic block execution data based on the instrumentation logic; wherein the instrumentation logic is compiled according to the static library;

[0101] An establishing unit 22, configured to analyze the performance characteristics of the target program through the basic block execution data, and establish a quantitative mapping relationship between the performance characteristics and the micro-architecture of the graphics processor;

[0102] The optimization unit 23 is used to optimize the performance of the graphics processor based on the quantization mapping relationship.

[0103] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 3 As shown, the device also includes:

[0104] The acquisition unit 24 is used to acquire the target compilation tool before the acquisition unit 21 runs the target program and acquires the basic block execution data based on the instrumentation logic, and compile the source code based on the target compilation tool and the definition rules to obtain the static library.

[0105] A compilation unit 25, configured to compile the source code based on the application programming interface of the static library to obtain a dynamic link library.

[0106] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 3 shown, the device further includes:

[0107] An insertion unit 26, configured to insert the dynamic link library at the function entry of the target program and at the start position of the basic block before the acquisition unit 21 runs the target program and acquires basic block execution data based on a preset instrumentation logic;

[0108] A setting unit 27, configured to set an environment variable and establish a link relationship between the environment variable and the dynamic link library; so as to call an internal function of the dynamic link library to perform data acquisition on the basic block when the target program runs.

[0109] Further, in a possible implementation manner of the embodiments of the present disclosure, the optimization unit 23 is further configured to:

[0110] Identify a hot spot area in the target program according to the basic block count result, where the hot spot area is a basic block determined according to the execution times and / or execution time;

[0111] Analyze the control flow complexity of the hot spot area, reorganize the code structure, and reduce the nested branch and loop levels;

[0112] Evaluate the memory access pattern of the hot spot area, improve the data storage layout, and reduce the number of memory accesses.

[0113] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 3 shown, the device further includes:

[0114] An acquisition unit 28, configured to re-run the target program after the optimization unit 23 optimizes the performance of the graphics processor based on the quantization mapping relationship, and re-acquire the execution data of the optimized basic block based on the instrumentation logic and

[0115] An evaluation unit 29, configured to compare the execution data of the basic block before optimization and the execution data of the optimized basic block, and evaluate the performance improvement effect.

[0116] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 3 shown, the device further includes:

[0117] A generation unit 210, configured to generate an executable file based on a preset benchmark test set and a preset renderer before the acquisition unit 28 runs a target program and collects basic block execution data based on a stubbing logic; wherein, test cases are included in the executable file.

[0118] The acquisition unit 28 is further configured to:

[0119] Run the test cases in the executable file based on the target program, and collect basic block execution data based on the stubbing logic.

[0120] Further, in a possible implementation manner of the embodiments of the present disclosure, the basic block execution data includes at least one of the execution times, execution times, and execution paths of basic blocks.

[0121] For the descriptions of the features in the corresponding embodiments of the performance optimization device of the graphics processing unit, reference may be made to the relevant descriptions in the corresponding embodiments of the performance optimization method of the graphics processing unit, which will not be elaborated herein one by one.

[0122] The embodiments of the present application further provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the performance optimization method of the graphics processing unit.

[0123] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned embodiments of the performance optimization method of the graphics processing unit when running.

[0124] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, and other media that can store computer programs.

[0125] The embodiments of the present application further provide a computer program product. The above-mentioned computer program product includes a computer program, and the steps in any of the above-mentioned embodiments of the performance optimization method of the graphics processing unit are implemented when the computer program is executed by a processor.

[0126] The embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps in any of the above-mentioned embodiments of the performance optimization method of the graphics processing unit are implemented when the computer program is executed by a processor.

[0127] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0128] The above has introduced in detail a method and device for optimizing the performance of a graphics processor, an electronic device, and a storage medium provided by this application. Specific examples are used herein to illustrate the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for optimizing the performance of a graphics processor, characterized in that, including: Run the target program and collect basic block execution data based on the instrumentation logic; wherein, the instrumentation logic is compiled from a static library; Analyze the performance characteristics of the target program through the basic block execution data, and establish a quantitative mapping relationship between the performance characteristics and the microarchitecture of the graphics processor; Perform performance optimization on the graphics processor based on the quantitative mapping relationship.

2. The performance optimization method of the graphics processor according to claim 1, characterized in that Before running the target program and collecting basic block execution data based on the instrumentation logic, the method further includes: Obtain a target compilation tool, and compile the source code based on the target compilation tool and definition rules to obtain the static library Compile the source code based on the application programming interface of the static library to obtain a dynamic link library.

3. The performance optimization method of the graphics processor according to claim 2, wherein Before running the target program and collecting basic block execution data based on a preset instrumentation logic, the method further includes: Insert the dynamic link library at the function entry of the target program and the start position of the basic block; Set an environment variable and establish a link relationship between the environment variable and the dynamic link library; so that when the target program runs, the internal function of the dynamic link library is called to perform data collection on the basic block.

4. The performance optimization method of the graphics processor according to claim 1, wherein The performing performance optimization on the graphics processor based on the quantitative mapping relationship includes: According to the basic block count result, identify the hot spots in the target program, where the hot spots are basic blocks determined according to the execution times and / or execution time; Analyze the control flow complexity of the hot spots, reorganize the code structure, and reduce the nested branch and loop levels; Evaluate the memory access pattern of the hot spots, improve the data storage layout, and reduce the number of memory accesses.

5. The performance optimization method of the graphics processor according to claim 3, wherein After performing performance optimization on the graphics processor based on the quantitative mapping relationship, the method further includes: Rerun the target program, and re-collect the execution data of the optimized basic blocks based on the instrumentation logic and Compare the execution data of the basic blocks before optimization and the execution data of the optimized basic blocks, and evaluate the performance improvement effect.

6. The performance optimization method of the graphics processor according to claim 3, wherein Before running the target program and collecting basic block execution data based on the instrumentation logic, the method further includes: Generate an executable file based on a preset benchmark test set and a preset renderer; wherein, the executable file contains test cases; The running the target program and collecting basic block execution data based on the instrumentation logic includes: Run the test cases in the executable file based on the target program, and collect basic block execution data based on the instrumentation logic.

7. The performance optimization method of the graphics processor according to any one of claims 1-6, characterized in that The basic block execution data includes at least one of the execution times, execution time, and execution path of the basic block.

8. A performance optimization device for a graphics processor, characterized in that, including: A collection unit for running the target program and collecting basic block execution data based on the instrumentation logic; wherein, the instrumentation logic is compiled from a static library; A establishment unit for analyzing the performance characteristics of the target program through the basic block execution data and establishing a quantitative mapping relationship between the performance characteristics and the microarchitecture of the graphics processor; An optimization unit for performing performance optimization on the graphics processor based on the quantitative mapping relationship.

9. An electronic device, characterized in that, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are for causing the computer to execute the method according to any one of claims 1-7.

Citation Information

Cited By

  • Compiling method, electronic equipment and storage medium

    CN120704693A