Method for automatically extracting hot spot codes of parallel application programs

By statically analyzing the source code of parallel applications, dividing them into multiple code snippets, and building a program running diagram to identify and extract code snippets whose resource consumption characteristics are greater than the threshold, it solves the inaccuracy and inefficiency of existing tools when identifying hotspot codes of parallel applications, and achieves more efficient and accurate hotspot code extraction.

CN119938309AActive Publication Date: 2025-05-06ZHEJIANG UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411795933.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-06
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Existing dynamic and static performance analysis tools have additional overhead, lack of runtime context and data granularity issues in identifying hotspot code for parallel applications, resulting in inaccurate and inefficient analysis results.

Method used

By performing static analysis of parallel applications, dividing the source code into multiple code snippets, and building a program run diagram based on the runtime resource consumption characteristics, and feature identification and extraction of code snippets whose resource consumption characteristics are greater than the set threshold.

Benefits of technology

It improves the accuracy and efficiency of hotspot code extraction, helps developers identify real performance bottleneck code snippets, and facilitates subsequent selection of specific hotspot codes in the simulator for analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938309A_ABST
    Figure CN119938309A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically extracting hot codes of parallel application programs, which is characterized by comprising the following steps of: statically analyzing the parallel application programs to be tested to obtain source codes of the parallel application programs to be tested, and dividing the source codes into a plurality of code snippets according to resource characteristics consumed when the parallel application programs to be tested run; marking according to the occurrence time of the running states when the parallel application program to be tested runs, marking the running states as vertexes, marking the running process of transferring from one running state to another running state as edges between the vertexes, and constructing a program running diagram; all the code snippets are attached to all vertexes and edges in the program running diagram, feature identification is conducted on the code snippets with consumed resource features larger than a set threshold value, and the code snippets subjected to feature identification serve as hotspot codes and are extracted; the method has the advantages that the hot spot codes can be conveniently and quickly extracted, and the hot spot code extraction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of computer technology, and in particular to a method for automatically extracting hotspot codes of parallel application programs. Background Art

[0002] A parallel application is an application that can be executed simultaneously on multiple processors or multiple computing nodes. In a parallel application, a task is broken down into multiple subtasks, which can be executed independently and simultaneously on different processors or computing nodes. Certain communication and synchronization mechanisms are required between the subtasks to coordinate their execution and share data.

[0003] Hotspot code identification and extraction is a performance analysis and optimization technology that is mainly used to identify and extract code snippets that are executed frequently, for a long time, and have a greater impact on the overall performance, as well as source code snippets that can represent the running behavior of the application. During the performance analysis and optimization process, developers identify hotspot codes, find code snippets that become system performance bottlenecks, and optimize them to reduce execution time and resource consumption, thereby improving system performance.

[0004] To identify and extract hot code in common workloads, developers usually use dynamic performance analysis tools or static performance analysis tools, but they still have the following shortcomings:

[0005] 1. Existing dynamic performance analysis tools introduce large overhead when collecting runtime data, which may affect the actual execution speed and behavior of the program. This means that the collected data may be different from the situation when it is not monitored, which may distort the analysis results, especially in time-sensitive applications and fine-grained parallel programs;

[0006] 2. Static performance analysis tools analyze code without executing it, resulting in a lack of runtime context and inability to consider dynamic behaviors, such as actual input data, runtime conditions, or dynamic interactions between threads, which are critical for identifying the real hot code in parallel programs;

[0007] 3. Many existing dynamic or static performance analysis tools have the problem of data granularity: too coarse granularity may miss specific bottlenecks, while too fine granularity may lead to huge data volumes and difficult analysis.

[0008] Identifying and extracting hot code for parallel applications is more complicated than identifying and extracting hot code for ordinary workloads because factors such as inter-process communication, computing resources, memory resources, and load balancing must be considered. Existing tools and methods are not efficient and accurate enough when dealing with large-scale parallel applications. Summary of the invention

[0009] The technical problem to be solved by the present invention is to provide a method for automatically extracting hotspot codes of parallel application programs, which can not only extract hotspot codes conveniently and quickly, but also improve the accuracy of extracting hotspot codes.

[0010] The technical solution adopted by the present invention to solve the above technical problem is: a method for automatically extracting hotspot codes of parallel application programs, comprising the following steps:

[0011] Step ①, by performing static analysis on the parallel application to be tested, the source code of the parallel application to be tested is obtained, and the source code is divided into multiple code fragments according to the resource characteristics consumed by the parallel application to be tested during runtime;

[0012] Step ②, marking the running states in order according to the time when the parallel application to be tested appears when running, marking the running states as vertices, marking the running process of transferring from one running state to another as edges between vertices, and constructing a program running graph;

[0013] Step ③, attach all code snippets to the corresponding vertices and edges in the program running graph, identify the code snippets whose resource consumption characteristics are greater than the set threshold, and extract the code snippets after feature identification as hot code.

[0014] Compared with the prior art, the advantages of the present invention are that the source code is divided into multiple code fragments according to the resource characteristics consumed by the parallel application to be tested during runtime, which helps to accurately evaluate the amount of resources consumed by each code fragment, thereby helping to identify and extract hot code; constructing a program operation diagram helps developers clearly understand the entire parallel application operation process and the synchronization or asynchronous process between threads; the code fragments are marked with feature identifiers according to the consumed resource characteristics, and the code fragments with feature identifiers are used as hot code and extracted, which improves the accuracy of hot code extraction and makes it convenient for subsequent developers to select corresponding hot code to run in the simulator as needed. For example, if the developer only wants to explore computing power, he only needs to select the code fragment containing the C identifier from the hot code set to run, which saves more time for the developer.

[0015] Furthermore, the resource characteristics include inter-thread communication resources, computing resources and input and output resources.

[0016] Furthermore, in step ③, the code snippets whose resource consumption characteristics are greater than the set threshold are feature-identified, and the feature-identified code snippets are used as hot code and extracted. The specific operation process is as follows:

[0017] Step ③-1, construct a feature vector V for each code snippet, V = (M, C, IO), where M represents the inter-thread communication resources consumed by the code snippet when the parallel application is running, C represents the computing resources consumed by the code snippet when the parallel application is running, and IO represents the input and output resources consumed by the code snippet when the parallel application is running;

[0018] Step ③-2, normalize each dimension of the feature vector of each code snippet to obtain a normalized feature vector, wherein the normalized feature vector corresponding to the i-th code snippet is denoted as V i , n represents the total number of code snippets, M i represents the inter-thread communication resources consumed by the i-th code snippet, C i Indicates the computing resources consumed by the i-th code snippet, IO i Represents the input and output resources consumed by the i-th code snippet;

[0019] Step ③-3, determine whether the total number of code snippets on each vertex or each edge in the program running graph is single or multiple. If it is single, execute step ③-4; if it is multiple, execute step ③-5;

[0020] Step ③-4, traverse each dimension of the normalized feature vector corresponding to each vertex or each edge of the code snippet in the program running graph. If the value of a certain dimension is greater than the preset first threshold, the code snippet is considered to have the corresponding resource consumption feature, and a feature mark is given to it. The code snippet after the feature mark is taken as the hot code and extracted; if the value of each dimension is less than or equal to the preset first threshold, the code snippet is not marked with a feature;

[0021] Step ③-5, clustering all code snippets on each vertex or each edge in the program running graph to obtain multiple clustered code clusters, wherein the specific operation of obtaining a code cluster is: if the Euclidean distance between the normalized feature vectors corresponding to any two code snippets is less than a preset second threshold, the two code snippets are classified into one code cluster;

[0022] Step ③-6, in each code cluster, select the code fragment with the smallest total distance from the feature vectors of other code fragments, and traverse each dimension of the normalized feature vector corresponding to the code fragment. If the value of a certain dimension is greater than the preset first threshold, it is considered that the code fragment has the corresponding resource consumption feature, and all code fragments in the code cluster where the code fragment is located are marked with a feature identifier, and the code fragment after the feature identifier is used as a hot code and extracted; if the value of each dimension is less than or equal to the preset first threshold, all code fragments in the code cluster where the code fragment is located are not marked with a feature. Because the resource features consumed by all code fragments in each code cluster are similar, the code fragment with the smallest total distance from the feature vectors of other code fragments is selected as a representative point in each code cluster. In a code cluster, the resource features consumed by all code fragments in the code cluster can be judged based on the resource features consumed by the representative point, without calculating all code fragments in the code cluster one by one, reducing the amount of calculation and speeding up the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0024] Figure 2 It is a schematic diagram of the program operation diagram in the present invention. DETAILED DESCRIPTION

[0025] The present invention is further described in detail below with reference to the accompanying drawings.

[0026] like Figure 1 As shown, a method for automatically extracting hotspot codes of parallel applications includes the following steps:

[0027] Step ①, by performing static analysis on the parallel application to be tested, the source code of the parallel application to be tested is obtained, and the source code is divided into multiple code fragments according to the resource characteristics consumed by the parallel application to be tested when it is running; wherein static analysis refers to detecting the source code of the application when the application is not running; the resource characteristics include inter-thread communication resources, computing resources and input-output resources; the multiple code fragments include multiple code fragments whose resource characteristics consumed are inter-thread communication resources, multiple code fragments whose resource characteristics consumed are computing resources and multiple code fragments whose resource characteristics consumed are input-output resources;

[0028] Record the start and end positions of MPI (Message Passing Interface) calls in the source code, and use them as code snippets whose resource consumption is characterized as inter-thread communication resources. Use the time of MPI calls as an indicator to measure the inter-thread communication resources consumed.

[0029] Identify and monitor loop structures in source code, and use the consumed resource characteristics as code snippets of computing resources. Use dynamic binary translation technology to intercept program execution flow, calculate the number of instructions consumed by loop execution, and use the number of instructions consumed by the loop as an indicator to measure the consumed computing resources.

[0030] For example: Identify loop structures in source code and find the syntax features of loop structures such as for, while, do-while, etc.

[0031] Identify input and output program fragments in the source code, and use the code fragments of the input and output resources as the consumed resource characteristics, and use the time consumed by the input and output in the input and output program fragments as an indicator to measure the consumed input and output resources;

[0032] An example of identifying input and output program fragments in source code: finding function parameters that are resource identifiers (such as file paths, database connections, etc.) and data containers (such as buffer pointers, etc.), and finding input and output-related library functions (such as fopen, fread, etc.) called in the function body;

[0033] Step ②, marking the running states in order according to the time when the running state of the parallel application to be tested appears, marking the running states as vertices, marking the running process of transferring from one running state to another running state as edges between vertices, and constructing a program running graph; the edges between vertices represent the running process of transferring from one running state to another running state, for example, transferring from one running state to another running state through communication between threads, or transferring from one loop iteration to another loop iteration; wherein the running state includes the memory state and the value of the register;

[0034] Step ③, attach all code snippets to the corresponding vertices and edges in the program running graph, identify the code snippets whose resource consumption characteristics are greater than the set threshold, and extract the identified code snippets as hot codes;

[0035] The vertices of the program execution graph are derived from the call stack information of external function calls, while the edges represent the conversions between these function calls. Connecting the divided code snippets to the corresponding vertices and edges in the program execution graph helps to organize and visualize the execution flow of the program.

[0036] The specific operation process of identifying the code snippets whose resource consumption characteristics are greater than the set threshold and taking the identified code snippets as hot codes and extracting them is as follows:

[0037] Step ③-1, construct a feature vector V for each code snippet, V = (M, C, IO), where M represents the inter-thread communication resources consumed by the code snippet when the parallel application is running, C represents the computing resources consumed by the code snippet when the parallel application is running, and IO represents the input and output resources consumed by the code snippet when the parallel application is running;

[0038] Step ③-2, normalize each dimension of the feature vector of each code snippet to obtain a normalized feature vector, wherein the normalized feature vector corresponding to the i-th code snippet is denoted as V i , n represents the total number of code snippets, M i represents the inter-thread communication resources consumed by the i-th code snippet, C i Indicates the computing resources consumed by the i-th code snippet, IO i Represents the input and output resources consumed by the i-th code snippet;

[0039] Step ③-3, determine whether the total number of code snippets on each vertex or each edge in the program running graph is single or multiple. If it is single, execute step ③-4; if it is multiple, execute step ③-5;

[0040] Step ③-4, traverse each dimension of the normalized feature vector corresponding to each vertex or each edge of the code snippet in the program running graph. If the value of a certain dimension is greater than the preset first threshold, the code snippet is considered to have the corresponding resource consumption feature, and a feature mark is given to it, and the code snippet after the feature mark is taken as a hot code and extracted; if the value of each dimension is less than or equal to the preset first threshold, the code snippet is not marked, and the code snippet is not a hot code; the preset first threshold is determined by the user. For example, if the user wants to filter out code snippets with a resource consumption feature ratio greater than 15%, the first threshold is set to 15%;

[0041] Step ③-5, clustering all code snippets on each vertex or each edge in the program running graph to obtain multiple clustered code clusters, wherein the specific operation of obtaining a code cluster is: if the Euclidean distance between the normalized feature vectors corresponding to any two code snippets is less than a preset second threshold, the two code snippets are classified as one code cluster; the preset second threshold is defined by the user, for example, if the user hopes that the similarity of resource features consumed by two code snippets is greater than 90%, the second threshold is set to 10%;

[0042] Step ③-6, in each code cluster, obtain the distance between the feature vector of each code fragment and the feature vector of other code fragments, select the code fragment with the smallest total distance from the feature vectors of other code fragments, and traverse each dimension in the normalized feature vector corresponding to the code fragment. If the value of a certain dimension is greater than the preset first threshold, it is considered that the code fragment has the corresponding resource consumption feature, and all code fragments in the code cluster where the code fragment is located are marked with a feature identifier, and the code fragment with the feature identifier is used as a hot code and extracted; if the value of each dimension is less than or equal to the preset first threshold, all code fragments in the code cluster where the code fragment is located are not marked with a feature identifier;

[0043] For example: After clustering all the code snippets on vertex X, three code clusters are obtained, namely, code cluster a, code cluster b, and code cluster c. Each code cluster contains some code snippets. Then, in code cluster a, the code snippet with the smallest total distance from the feature vectors of other code snippets is selected, and the dimension where the code snippet M, the dimension where C, and the dimension where IO are located are traversed to determine whether it is greater than the preset first threshold; the same applies to code clusters b and c.

[0044] For example, if the first threshold is set to 15%, and the normalized feature vector corresponding to a code snippet is (0.16, 0.13, 0.155), then the code snippet is considered to be both an MPI snippet and an IO snippet, and is marked with a feature identifier (M, IO).

Claims

1. A method for automatically extracting hotspot codes of parallel applications, characterized in that The following steps are involved: Step ①, by performing static analysis on the parallel application to be tested, the source code of the parallel application to be tested is obtained, and the source code is divided into multiple code fragments according to the resource characteristics consumed by the parallel application to be tested during runtime; Step ②, marking the running states in order according to the time when the parallel application to be tested appears when running, marking the running states as vertices, marking the running process of transferring from one running state to another as edges between vertices, and constructing a program running graph; Step ③, attach all code snippets to the corresponding vertices and edges in the program running graph, identify the code snippets whose resource consumption characteristics are greater than the set threshold, and extract the code snippets after feature identification as hot code.

2. A method for automatically extracting hotspot codes of parallel applications according to claim 1, characterized in that The resource characteristics include inter-thread communication resources, computing resources and input and output resources.

3. A method for automatically extracting hotspot codes of parallel applications according to claim 2, characterized in that In step ③, the code snippets whose resource consumption characteristics are greater than the set threshold are marked, and the marked code snippets are used as hot code and extracted. The specific operation process is as follows: Step ③-1, construct a feature vector V for each code snippet, V = (M, C, IO), where M represents the inter-thread communication resources consumed by the code snippet when the parallel application is running, C represents the computing resources consumed by the code snippet when the parallel application is running, and IO represents the input and output resources consumed by the code snippet when the parallel application is running; Step ③-2, normalize each dimension of the feature vector of each code snippet to obtain a normalized feature vector, wherein the normalized feature vector corresponding to the i-th code snippet is denoted as V i , Indicates the total number of code snippets, M i represents the inter-thread communication resources consumed by the i-th code snippet, C i Indicates the computing resources consumed by the i-th code snippet, IO i Represents the input and output resources consumed by the i-th code snippet; Step ③-3, determine whether the total number of code snippets on each vertex or each edge in the program running graph is single or multiple. If it is single, execute step ③-4; if it is multiple, execute step ③-5; Step ③-4, traverse each dimension of the normalized feature vector corresponding to each vertex or each edge of the code snippet in the program running graph. If the value of a certain dimension is greater than the preset first threshold, the code snippet is considered to have the corresponding resource consumption feature, and a feature mark is given to it. The code snippet after the feature mark is taken as the hot code and extracted; if the value of each dimension is less than or equal to the preset first threshold, the code snippet is not marked with a feature; Step ③-5, clustering all code snippets on each vertex or each edge in the program running graph to obtain multiple clustered code clusters, wherein the specific operation of obtaining a code cluster is: if the Euclidean distance between the normalized feature vectors corresponding to any two code snippets is less than a preset second threshold, the two code snippets are classified into one code cluster; Step ③-6, in each code cluster, select the code snippet with the shortest total distance from the feature vectors of other code snippets, and traverse each dimension of the normalized feature vector corresponding to the code snippet. If the value of a certain dimension is greater than the preset first threshold, the code snippet is considered to have the corresponding resource consumption feature, and all code snippets in the code cluster where the code snippet is located are marked with a feature identifier, and the code snippet after the feature identifier is used as a hot code and extracted; if the value of each dimension is less than or equal to the preset first threshold, all code snippets in the code cluster where the code snippet is located are not marked with a feature.

Citation Information

Patent Citations

  • Wavelet and wavelet packet multi-core parallel computation method based on OpenMP (open multi-processing)

    CN102222019A

  • Hotspot-driven workload feature analysis method and analysis system combining instrumentation and performance event sampling

    CN115328643A

  • Adaptive program parallel division and scheduling method for large-scale supercomputing system

    CN118245181A

  • Software application performance enhancement

    US20100042981A1

  • Method and system for optimizing code for a multi-threaded application

    US20110202907A1