MLIR code performance analysis method, system, device and equipment and storage medium
By inserting timing events into the MLIR code file and obtaining performance data, the shortcomings of existing tools in MLIR code performance analysis are solved, and efficient and accurate performance analysis and bottleneck positioning are achieved.
Patent Information
- Application Number
- CN202510798627.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing code performance analysis tools lack support for multi-level IRs such as MLIR (Multi-Level Intermediate Representation), making it difficult to perform performance analysis effectively.
Insert timing events in the MLIR code file, insert timing events for the target operation by declaring the timing functions and traversing the code file, obtaining performance data, and using the MLIR-OPT tool for code conversion and compilation, dynamically call the dynamic link library, detecting operation pointers, and cumulative performance data.
It realizes accurate performance analysis of MLIR code, can efficiently locate performance bottlenecks, reduce unnecessary overhead, simplify the ease of use and development efficiency of tools, and accurately obtain performance data for each target operation.
Smart Images

Figure CN120336194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of development technologies, and particularly to a method, system, device, equipment, and storage medium for analyzing the performance of MLIR code. Background Art
[0002] A profiler is a tool used to measure the performance characteristics of a program during execution. It helps developers understand the running behavior of the program by collecting data on function call frequencies, execution times, memory usage, and resource utilization. This information is crucial for identifying performance bottlenecks, optimizing code, and improving the overall performance of an application.
[0003] Existing code performance analysis methods or tools usually target a single-level code representation and lack support for multi-level IRs such as MLIR (Multi-Level Intermediate Representation). Therefore, there is an urgent need for a solution to analyze the performance of MLIR code. Summary of the Invention
[0004] The present invention provides a method, system, device, equipment, and storage medium for analyzing the performance of MLIR code to solve the defects in the prior art.
[0005] The present invention provides a method for analyzing the performance of MLIR code, including: Obtaining an MLIR code file, and inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; Executing the target MLIR code file to obtain performance data corresponding to the target operations; Performing performance analysis based on the performance data.
[0006] According to the method for analyzing the performance of MLIR code provided by the present invention, the step of inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file includes: Declaring a timing function in the MLIR code file, and defining a first function representing the start of timing and a second function representing the end of timing; Traversing the MLIR code file, determining the operation information of the target operations, and inserting the corresponding timing events for the target operations based on the first function and the second function to obtain the target MLIR code file.
[0007] A method for analyzing the performance of MLIR code provided by the present invention, inserting corresponding timing events for the target operations based on the first function and the second function, includes: Create corresponding timing event objects for each of the target operations, and associate the timing event objects with the corresponding target operations; Insert code calling the first function at the start position of each of the target operations, and create a first constant operation. Pass the address of the timing event object as a parameter to the corresponding first function through the first constant operation; Insert code calling the second function at the end position of each of the target operations, and create a second constant operation. Pass the address of the timing event object as a parameter to the corresponding second function through the second constant operation.
[0008] A method for analyzing the performance of MLIR code provided by the present invention, executing the target MLIR code file, includes: Perform code conversion and compilation on the target MLIR code file through the MLIR-OPT tool to obtain a code compilation file; Dynamically call the method in the dynamic link library in the code compilation file to complete the execution of the target MLIR code file.
[0009] A method for analyzing the performance of MLIR code provided by the present invention, executing the target MLIR code file to obtain the performance data corresponding to the target operations, includes: During the execution of the target MLIR code file, detect whether the current target operation is a null pointer; If it is not a null pointer, obtain the operation pointer associated with the current target operation; Detect whether the operation pointer is a null pointer; If it is not a null pointer, obtain the operation name and operation duration of the current target operation, and accumulate the operation duration to obtain the total operation duration. Use the total operation duration as the performance data.
[0010] A method for analyzing the performance of MLIR code provided by the present invention, performing performance analysis based on the performance data, includes: Perform performance analysis based on the performance data to obtain a performance analysis result; Display the performance data and the performance analysis result on a preset page.
[0011] The present invention also provides an MLIR code performance analysis system, including: The instrumentation module is used to insert corresponding timing events for target operations in the MLIR code file to obtain the target MLIR code file; The compilation module is used to execute the target MLIR code file to obtain the performance data corresponding to the target operation; The time management module is used to manage the timing events and perform performance analysis based on the performance data.
[0012] The present invention also provides an MLIR code performance analysis device, including: The acquisition module is configured to acquire an MLIR code file and insert corresponding timing events for target operations in the MLIR code file to obtain the target MLIR code file; The execution module is configured to execute the target MLIR code file to obtain the performance data corresponding to the target operation; The performance analysis module is configured to perform performance analysis based on the performance data.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the MLIR code performance analysis method described in any one of the above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the MLIR code performance analysis method described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the MLIR code performance analysis method described in any one of the above is implemented.
[0016] The MLIR code performance analysis method, system, device, equipment, and storage medium provided by the present invention insert corresponding timing events for target operations in the MLIR code file, so that each target operation has a corresponding timing function, obtain the target MLIR code file, execute the target MLIR code file, obtain the performance data corresponding to the target operation, and perform performance analysis based on the obtained performance data. By directly inserting timing events in MLIR, the performance data of each target operation can be accurately obtained, and thus the performance analysis of the MLIR code can be realized. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of the MLIR code performance analysis method provided by the present invention.
[0019] Figure 2 It is a schematic diagram of the MLIR code performance analysis system provided by the present invention.
[0020] Figure 3 It is a schematic diagram of the time management module provided by the present invention.
[0021] Figure 4 It is a schematic structural diagram of the MLIR code performance analysis device provided by the present invention.
[0022] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0024] Figure 1 It is a flowchart of a MLIR code performance analysis method shown according to an exemplary embodiment. As Figure 1 shown, in an exemplary embodiment, the MLIR code performance analysis method includes steps 110 to 130, which are introduced in detail as follows.
[0025] Step 110, obtain an MLIR code file, and insert corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; In the embodiments of the present invention, MLIR (Multi-Level Intermediate Representation) is an intermediate representation (IR) system between a language (such as C) or a library (such as TensorFlow) and a compiler backend (such as LLVM). It allows code reuse between different compiler stacks of different languages and has other performance and usability advantages.
[0026] The target operation can be all operations in the MLIR code file or certain specified operations in the MLIR code file. In the MLIR code file, corresponding timing events are inserted for the target operations so that each target operation has a corresponding timing function.
[0027] Step 120: Execute the target MLIR code file to obtain the performance data corresponding to the target operation.
[0028] In the embodiment of the present invention, the target MLIR code file is executed to obtain the performance data corresponding to the target operation.
[0029] Step 130: Perform performance analysis based on the performance data.
[0030] In the embodiment of the present invention, performance analysis is performed based on the obtained performance data. By directly inserting timing events in MLIR, the performance data of each target operation is accurately obtained, and thus the performance analysis of the MLIR code is realized.
[0031] In an exemplary embodiment of the present invention, inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file includes: Declare a timing function in the MLIR code file, and define a first function representing the start of timing and a second function representing the end of timing; Traverse the MLIR code file to determine the operation information of the target operation, and insert the corresponding timing events for the target operation based on the first function and the second function to obtain the target MLIR code file.
[0032] In the embodiment of the present invention, a builder is created according to the interface provided by the MLIR official, and operations are inserted through the OpBuilder constructor. The OpBuilder constructor can insert new operations in the MLIR module for subsequent insertion of timing function declarations and function calls.
[0033] Declare a timing function in the MLIR code file. Specifically, set the insertion point to the beginning of the MLIR code file, and then define the first function timingStart and the second function timingEnd. The first function timingStart is used to record the start time of the operation, and the second function timingEnd is used to record the end time of the operation. Declare these two functions in MLIR for subsequent timestamp insertion.
[0034] The first function and the second function accept a 64-bit integer type parameter and have no return value. The role of the parameter is to pass the unique identifier of each timing event, that is, the address of the TimeEvent pointer.
[0035] Traverse the MLIR code file, determine the operation information of the target operation, such as the operation name, and insert corresponding timing events TimeEvent for each target operation based on the first function and the second function to obtain the target MLIR code file. In this embodiment, a timing function is inserted for the target operation, so that the execution time of the operation can be recorded.
[0036] In an exemplary embodiment of the present invention, the inserting the corresponding timing event for the target operation based on the first function and the second function includes: Create a corresponding timing event object for each target operation, and associate the timing event object with the corresponding target operation; Insert code calling the first function at the start position of each target operation, and create a first constant operation. Pass the address of the timing event object as a parameter to the corresponding first function through the first constant operation; Insert code calling the second function at the end position of each target operation, and create a second constant operation. Pass the address of the timing event object as a parameter to the corresponding second function through the second constant operation.
[0037] In the embodiment of the present invention, traverse the MLIR code module, match the Dialect or operation (Op), check each operation (Op) in the function body one by one, and obtain its dialect name. For each operation, determine whether its dialect name belongs to the specified target dialect (targetDialect). If the dialect name of the operation matches the target dialect, it is determined as the target operation.
[0038] For each qualified target operation, create a new timing event object, associate it with the target operation, and add the timing event object to the TimeManager for management.
[0039] Insert code calling the first function timingStart at the start position of the target operation to record the start time of the target operation. Create a constant operation ConstantOp, and pass the address of the timing event object (converted to a 64-bit integer) as a parameter to the first function timingStart.
[0040] Insert code calling the second function timingEnd at the end position of the target operation to record the end time of the target operation. The passed parameter is the same as that of the first function timingStart, which is used to match the start time and end time of the same target operation.
[0041] After all eligible target operations in the MLIR code file have completed the timing event insertion, return a string of the execution result (e.g., "Instrumentation completed successfully") indicating that the insertion process has been completed.
[0042] Through the above steps, the timing function is inserted for the target operations of the specified dialect, recording the start time and end time of each operation. Finally, the time data of all timing events is recorded in the TimeManager and can be used for subsequent analysis and performance measurement.
[0043] In the technical solution provided by the embodiments of the present invention, the operations to be analyzed can be flexibly specified, improving the pertinence of performance analysis, that is, the specified operation names can be recorded in the MLIR file, and timestamps can be selectively inserted before and after these operations. This solution allows developers to focus on specific operations. Compared with the existing method of uniformly analyzing all operations by the Profiler, this solution can more efficiently locate performance bottlenecks and reduce unnecessary overhead.
[0044] In an exemplary embodiment of the present invention, the execution of the target MLIR code file includes: Perform code conversion and compilation on the target MLIR code file through the MLIR-OPT tool to obtain a code compilation file; Dynamically call the method in the dynamic link library in the code compilation file to complete the execution of the target MLIR code file.
[0045] In the embodiments of the present invention, after generating the target MLIR code file with timestamps, it is desired to call the code in the profiler to run the code for performance testing. In the MLIRProfiler, a multi-threaded method is used to support dynamic compilation, ensuring that the pointers of the previously recorded timing event objects are correct and can be used.
[0046] Specifically, a new thread is created using std::thread. In the new thread, the std::system system call is used to call the makefile tool to compile the target MLIR code file. In the makefile, the mlir-opt tool is used for operations such as code conversion and compilation to obtain a code compilation file.
[0047] The compiled library provides a C interface, and the dlopen method is used to dynamically call the method in the dynamic link library (such as.so files) in the compiled code compilation file.
[0048] The technical solution provided by the embodiments of the present invention can be closely integrated with the compilation and execution processes of MLIR, facilitating the compilation and running of MLIR files with timestamps. At the same time, compared with the existing Profilers that require additional configuration or complex integration, this solution simplifies the usage process and greatly improves the usability and development efficiency of the tool.
[0049] In an exemplary embodiment of the present invention, executing the target MLIR code file to obtain the performance data corresponding to the target operation includes: During the execution of the target MLIR code file, detecting whether the current target operation is a null pointer; If it is not a null pointer, obtaining the operation pointer associated with the current target operation; Detecting whether the operation pointer is a null pointer; If it is not a null pointer, obtaining the operation name and operation duration of the current target operation, accumulating the operation duration to obtain the total operation duration, and using the total operation duration as the performance data.
[0050] In the embodiments of the present invention, after the code execution ends, the performance data of the target operation is obtained, and these performance data are stored in the timing event list. The function processTimingData is used to process the event data in the TimeManager into a JSON file containing the target operation and its corresponding time information.
[0051] Specifically, the function checks whether the timing event list is empty. If it is empty, it directly outputs the error message "Events list is empty.", and then returns without performing subsequent operations, ensuring that unnecessary calculations are not performed when there is no performance data.
[0052] A times variable of type std::map<std::string, double> is defined to store the operation name and its corresponding total operation duration. At the same time, a duration variable of type double is defined to temporarily save the operation duration of each operation.
[0053] Obtaining the size of the timing event list and outputting its size, and then looping through each element in the timing event list. The loop variable i ranges from 0 to size - 1, and each loop processes a timing event object.
[0054] Inside the loop, first check whether the current target operation events[i] corresponding to the current index i is a null pointer nullptr. If it is a null pointer, output the error message "Null pointer found in events at index [i]" and skip the current loop to ensure the processing of valid timing event pointers.
[0055] If the current target operation events[i] is not a null pointer, call the events[i]->getOpPtr() method to obtain the operation pointer opPtr associated with the current timing event. Next, check whether opPtr is a null pointer. If it is a null pointer, output "Null pointer found in op Ptratindex[i]" and skip the current loop to ensure the processing of valid operation pointers.
[0056] Obtain the operation name opName of the current target operation through the events[i]->getOpName() method and output the name. At the same time, obtain the operation duration duration of the current target operation through the events[i]->getDuration() method. At this time, opName and duration store the operation name and operation duration of the current target operation respectively.
[0057] Check whether there is an entry with the operation name opName in the times container. If it already exists, add the current operation duration duration to times[opName]; otherwise, add the operation name opName and the operation duration duration to times to ensure that the operation durations with the same name can be accumulated together to obtain the total operation duration.
[0058] After the traversal is completed, times stores all operation names and their corresponding total operation durations. Then use the nlohmann::json library to convert times into a JSON object jsonObject. Try to open the specified file path outputFilepath for writing. If the file is successfully opened, convert jsonObject into a formatted JSON string and write it into the file, and finally close the file, outputting the success message "Results is saved in[outputFilepath]"; if the file cannot be opened, output the error message "Can’t open[outputFilepath]".
[0059] The technical solution provided by the embodiments of the present invention establishes a direct association between performance data and operations at the MLIR level, directly records the corresponding MLIR Op information in the performance data, realizes the precise mapping between performance data and high-level code, solves the problem that existing Profilers are difficult to map performance data back to high-level abstract operations, and facilitates developers to perform performance analysis and optimization.
[0060] The present invention allows for the direct collection of performance data at the intermediate representation stage of the compiler by tracking the operation duration of each target and binding it to the corresponding target operation, and performs fine-grained performance measurement in MLIR, achieving precise performance analysis of high-level abstract operations.
[0061] In an exemplary embodiment of the present invention, performing performance analysis based on the performance data includes: Performing performance analysis based on the performance data to obtain a performance analysis result; Displaying the performance data and the performance analysis result on a preset page.
[0062] In the embodiments of the present invention, performance analysis is performed based on performance data to obtain a performance analysis result, and the performance analysis result is displayed on a preset page to intuitively display the performance analysis result.
[0063] Figure 2 It is a schematic diagram of an MLIR code performance analysis system shown according to an exemplary embodiment. As Figure 2 shown, in an exemplary embodiment, the MLIR code performance analysis system includes: An instrumentation module 210, configured to insert corresponding timing events for target operations in an MLIR code file to obtain a target MLIR code file; A compilation module 220, configured to execute the target MLIR code file to obtain performance data corresponding to the target operations; A time management module 230, configured to manage the timing events and perform performance analysis based on the performance data.
[0064] In the embodiments of the present invention, after obtaining the MLIR code file, the time management module (TimeManager) 230 and the timing events are initialized, the MLIR code file is read and parsed, the instrumentation module 210 inserts timing events for target operations according to user requirements to obtain a target MLIR code file, the target MLIR code file is dynamically compiled by the compilation module 220, and the target MLIR code file is called and executed. During the process of calling and executing the MLIR code file, the time management module 230 is as Figure 3As shown, the TimeManager instance is obtained by calling the getTimeManager() function. The TimeManager is used to uniformly manage all timing events and record the start time and end time inside the function.
[0065] When the target MLIR code file executes to the timestamp, the time manager will call the runtime function to record the timestamp and the corresponding target operation. This integration method enables the runtime to control and monitor the execution process of MLIR and collect performance data in real time. A close connection is established between the runtime and the MLIR execution, achieving efficient performance data collection.
[0066] In another embodiment of the present invention, the MLIR code performance analysis system further includes: A data visualization module for displaying performance data and performance analysis results.
[0067] In the embodiment of the present invention, the data visualization module provides visualization of performance data and intuitively displays the performance analysis results.
[0068] The MLIR code performance analysis device provided by the present invention will be described below. The MLIR code performance analysis device described below can be correspondingly referred to the MLIR code performance analysis method described above. It should be noted that the device provided in the following embodiments and the method provided in the above embodiments belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiments and will not be repeated here.
[0069] In an exemplary embodiment of the present invention, please refer to Figure 4 , Figure 4 is a MLIR code performance analysis device shown according to an exemplary embodiment, including the following modules.
[0070] An acquisition module 410, configured to acquire an MLIR code file and insert corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; An execution module 420, configured to execute the target MLIR code file to obtain performance data corresponding to the target operation; A performance analysis module 430, configured to perform performance analysis based on the performance data.
[0071] In an exemplary embodiment of the present invention, the acquisition module 410 includes: A definition sub-module, configured to declare a timing function in the MLIR code file and define a first function representing the start of timing and a second function representing the end of timing; A determination sub-module, configured to traverse the MLIR code file, determine operation information of the target operation, and insert corresponding timing events for the target operation based on the first function and the second function, to obtain the target MLIR code file.
[0072] In an exemplary embodiment of the present invention, the determination sub-module includes: A creation unit, configured to create corresponding timing event objects for each of the target operations, and associate the timing event objects with the corresponding target operations; A first insertion unit, configured to insert codes for calling the first function at the start positions of each of the target operations respectively, and create a first constant operation, and pass the address of the timing event object as a parameter to the corresponding first function through the first constant operation; A second insertion unit, configured to insert codes for calling the second function at the end positions of each of the target operations respectively, and create a second constant operation, and pass the address of the timing event object as a parameter to the corresponding second function through the second constant operation.
[0073] In an exemplary embodiment of the present invention, the execution module 420 includes: A compilation sub-module, configured to perform code conversion and compilation on the target MLIR code file through the MLIR-OPT tool, to obtain a code compilation file; A call sub-module, configured to dynamically call methods in the dynamic link library in the code compilation file to complete the execution of the target MLIR code file.
[0074] In an exemplary embodiment of the present invention, the execution module 420 includes: A first detection sub-module, configured to detect whether the current target operation is a null pointer during the execution of the target MLIR code file; An acquisition sub-module, configured to, if it is not a null pointer, acquire the operation pointer associated with the current target operation; A second detection sub-module, configured to detect whether the operation pointer is a null pointer; An accumulation sub-module, configured to, if it is not a null pointer, acquire the operation name and operation duration of the current target operation, and accumulate the operation duration to obtain the total operation duration, and use the total operation duration as the performance data.
[0075] In an exemplary embodiment of the present invention, the performance analysis module 430 includes: A performance analysis sub-module, configured to perform performance analysis based on the performance data to obtain a performance analysis result; A display sub-module, configured to display the performance data and the performance analysis result on a preset page.
[0076] Figure 5 Fig. illustrates a schematic physical structure diagram of an electronic device, as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute the MLIR code performance analysis method, which includes: obtaining an MLIR code file, and inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; Executing the target MLIR code file to obtain performance data corresponding to the target operation; Performing performance analysis based on the performance data.
[0077] In addition, when the logic instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0078] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the MLIR code performance analysis method provided by the above-mentioned various methods. The method includes: obtaining an MLIR code file, and inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; Executing the target MLIR code file to obtain performance data corresponding to the target operation; Performing performance analysis based on the performance data.
[0079] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the MLIR code performance analysis method provided by the above-mentioned various methods. The method includes: obtaining an MLIR code file, and inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; Executing the target MLIR code file to obtain performance data corresponding to the target operations; Performing performance analysis based on the performance data.
[0080] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0081] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for analyzing the performance of MLIR code, characterized in that, Including: Obtain the MLIR code file, and insert corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; Execute the target MLIR code file to obtain performance data corresponding to the target operations; Perform performance analysis based on the performance data.
2. The MLIR code performance analysis method according to claim 1, wherein The step of inserting corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file includes: Declare a timing function in the MLIR code file, and define a first function representing the start of timing and a second function representing the end of timing; Traverse the MLIR code file to determine the operation information of the target operations, and insert the corresponding timing events for the target operations based on the first function and the second function to obtain the target MLIR code file.
3. The MLIR code performance analysis method according to claim 2, wherein The step of inserting the corresponding timing events for the target operations based on the first function and the second function includes: Create corresponding timing event objects for each of the target operations, and associate the timing event objects with the corresponding target operations; Insert code calling the first function at the start position of each of the target operations, and create a first constant operation to pass the address of the timing event object as a parameter to the corresponding first function through the first constant operation; Insert code calling the second function at the end position of each of the target operations, and create a second constant operation to pass the address of the timing event object as a parameter to the corresponding second function through the second constant operation.
4. The MLIR code performance analysis method according to claim 1, wherein The step of executing the target MLIR code file includes: Perform code conversion and compilation on the target MLIR code file through the MLIR-OPT tool to obtain a code compilation file; Dynamically call the methods in the dynamic link library in the code compilation file to complete the execution of the target MLIR code file.
5. The MLIR code performance analysis method according to claim 1, wherein The step of executing the target MLIR code file to obtain performance data corresponding to the target operations includes: During the execution of the target MLIR code file, detect whether the current target operation is a null pointer; If it is not a null pointer, obtain the operation pointer associated with the current target operation; Detect whether the operation pointer is a null pointer; If it is not a null pointer, obtain the operation name and operation duration of the current target operation, and accumulate the operation duration to obtain the total operation duration, and use the total operation duration as the performance data.
6. The MLIR code performance analysis method according to any one of claims 1 to 5, characterized in that, The step of performing performance analysis based on the performance data includes: Perform performance analysis based on the performance data to obtain a performance analysis result; Display the performance data and the performance analysis result on a preset page.
7. An MLIR code performance analysis system, characterized in that Including: An instrumentation module for inserting corresponding timing events for target operations in an MLIR code file to obtain a target MLIR code file; A compilation module for executing the target MLIR code file to obtain performance data corresponding to the target operations; A time management module for managing the timing events and performing performance analysis based on the performance data.
8. An MLIR code performance analysis device, characterized in that Including: An acquisition module, configured to acquire an MLIR code file, and insert corresponding timing events for target operations in the MLIR code file to obtain a target MLIR code file; An execution module, configured to execute the target MLIR code file to obtain performance data corresponding to the target operations; A performance analysis module, configured to perform performance analysis based on the performance data.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the MLIR code performance analysis method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a processor, it implements the MLIR code performance analysis method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Function running time measuring method applied to parallel scientific calculation program
CN112882912A
Low-efficiency code detection method and device
CN114168428A
Code-oriented application program performance analysis method and system
CN114968779A
Software program running time delay monitoring method and device, storage medium and equipment
CN115033460A
Performance test method and device, computer equipment and storage medium
CN117271371A