LLVM-based Program Static Analysis Method and System

Through the LLVM-based program static analysis method, IR files are generated and control flow diagrams are constructed, and the basic block transfer probability is calculated. The problem of low accuracy of program performance characteristics and memory reuse distance distribution in the prior art is solved, and efficient program performance analysis is achieved.

CN119883861BActive Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
CN202510352877.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Existing program analysis techniques have problems with low accuracy, low efficiency and limited application in estimating program performance characteristics and memory reuse distance distribution.

Method used

The static analysis method of program based on LLVM is used to generate IR files, build a basic block-level control flow chart, calculate the transfer probability between basic blocks, and use recursive algorithms to calculate the memory access reuse distance distribution, including steps such as IR file generation, reading, control flow chart generation, basic block calculation, static memory tracking and reuse distance distribution calculation.

Benefits of technology

Accurate analysis of program performance characteristics and memory reuse distance distribution is achieved, and analysis efficiency and scope of application are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883861B_ABST
    Figure CN119883861B_ABST
Patent Text Reader

Abstract

The present invention discloses a program static analysis method and system based on LLVM, belonging to the technical field of computer program analysis. The technical problem to be solved is that the existing program analysis technologies have low accuracy, low efficiency and limited application scope in estimating program performance characteristics and memory reuse distance distribution. The method includes: converting a source file into an LLVM IR file with source-level debugging information; traversing modules, functions, basic blocks and instructions in the LLVM IR file to obtain static trace information of a target function; constructing a basic block-level control flow graph of the LLVM IR file and annotating relevant information; identifying and annotating loop information when traversing paths to obtain an execution path with loop annotations, replacing basic blocks in the execution path with corresponding memory access information, and generating a static memory trace with loop annotations; calculating the memory access reuse distance distribution through a recursive algorithm based on the static memory trace with loop annotations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer program analysis, and more specifically to a program static analysis method and system based on LLVM. Background Art

[0002] In order to deeply understand the impact of hardware design changes on system performance and conduct in-depth testing and analysis of software performance when running on hardware, in software program analysis, cache utilization and data access patterns have a significant impact on program performance. Memory reuse distance refers to the time interval between two memory accesses, which has a significant impact on program performance. A smaller reuse distance means more local memory accesses, which can improve cache hit rate, reduce memory access latency, and thus improve the overall performance of the program. On the contrary, a larger reuse distance will lead to cache pollution, increase memory access latency, and affect program performance.

[0003] LLVM is an open-source compiler framework that can provide compiler intermediate representation (IR). The static analysis tool of LLVM can perform code analysis at the IR level to help identify problems in the code. The LLVM framework allows users to write custom passes for static analysis.

[0004] Traditional dynamic analysis methods require running the program and collecting a large amount of memory trace information when obtaining features such as the number of arithmetic operations, memory operations, and the number of basic block executions in the program. This process is extremely time-consuming and space-consuming. Although existing static analysis methods are faster, they have significant deficiencies in terms of accuracy, the ability to handle complex data structures (such as pointers and array references), the effectiveness of branch probability calculation, and the support for multi-function programs.

[0005] The problems of low accuracy, low efficiency, and limited applicability range in existing program analysis techniques for estimating program performance characteristics and memory reuse distance distribution are technical problems that need to be solved. Summary of the Invention

[0006] The technical task of the present invention is to address the above deficiencies and provide a program static analysis method and system based on LLVM to solve the technical problems of low accuracy, low efficiency, and limited applicability range in existing program analysis techniques for estimating program performance characteristics and memory reuse distance distribution.

[0007] In a first aspect, a program static analysis method based on LLVM according to the present invention includes the following steps:

[0008] IR file generation: Process the target program through program instructions to convert the source file into an LLVM IR file with source-level debugging information;

[0009] IR file reading: Taking the LLVM IR file and the function name of the target function as input, the LLVM IR file is parsed by a file reader, and the modules, functions, basic blocks, and instructions in the LLVM IR file are traversed to obtain the static trace information of the target function;

[0010] Control flow graph generation: Based on the static trace information of the target function, calculate the instruction types executed by the basic blocks, determine the successor and predecessor relationships of the basic blocks, and construct a basic block-level control flow graph of the LLVM IR file with the basic blocks of the target function as nodes and the control flow directions between the nodes as edges, and annotate the relevant information, where the relevant information includes the transfer probability between basic blocks, the execution times of basic blocks, and the memory access information;

[0011] Basic block calculation: Determine the transfer probability between basic blocks by analyzing the instructions executed by the basic blocks, and calculate the prior probability of each basic block being executed based on the transfer probability between basic blocks;

[0012] Static memory tracing: Analyze the control flow graph, traverse the paths based on the transfer probability between basic blocks and find the path with the maximum probability, identify and annotate the loop information when traversing the path, obtain the execution path with loop annotations, replace the basic blocks in the execution path with the corresponding memory access information, and generate the static memory tracing with loop annotations;

[0013] Reuse distance distribution calculation: Calculate the memory access reuse distance distribution based on the static memory tracing with loop annotations through a recursive algorithm.

[0014] Preferably, the control flow graph generation includes the following steps:

[0015] Analyze the instruction types executed by each basic block;

[0016] For each basic block, analyze the last instruction in the basic block to determine the predecessor and successor basic blocks of the basic block;

[0017] Identify all possible predecessor basic blocks of each basic block, and record the position of the first instruction in the source program in the successor or predecessor basic block;

[0018] For each basic block, take all possible predecessor basic blocks of the basic block and the position of the first instruction in the source program in the corresponding successor or predecessor basic block as the predecessor and successor information of the basic block, construct the control flow graph of the basic blocks in the target function based on the predecessor and successor information of the basic block, and annotate the relevant information in the control flow graph.

[0019] Preferably, the basic block calculation includes the following steps:

[0020] Analyze the instructions executed by the basic blocks through an offline coverage tool to obtain the transfer probability between basic blocks;

[0021] For each basic block in the control flow graph, a linear balance equation is constructed based on the transition probability from its predecessor basic block to the basic block, the transition probability from the basic block to its successor basic block, and the execution times of each basic block. The linear balance equation is expressed as:

[0022] ,

[0023] where, represents the transition probability from the predecessor basic block to the basic block , represents the transition probability from the basic block to the successor basic block , represents the execution times of the predecessor basic block , represents the execution times of the successor basic block , represents the predecessor basic block, represents the successor basic block;

[0024] The entry basic block of the program is executed only once. For the basic blocks in the control flow graph, homogeneous linear equations and one non - homogeneous equation can be formed. Starting from the entry basic block of the program, the linear balance equation is recursively solved to obtain the execution times of all basic blocks, and then the prior probability of a basic block being executed is calculated. The calculation formula for the prior probability of a basic block being executed is expressed as:

[0025] .

[0026] Preferably, static memory tracking includes the following steps:

[0027] Read the control flow graph and start traversing from the entry basic block. When encountering multiple possible edges, select the edge with the highest transition probability and mark it as visited according to the transition probability between basic blocks. The next time the same basic block is visited, select the unvisited edge with the highest transition probability to find the path with the highest probability;

[0028] After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundary of the loop and the loop control variable, and obtain the execution path with loop annotations;

[0029] Replace the basic blocks in the execution path with memory access information to obtain static memory tracking with loop annotations, so as to facilitate viewing the order and pattern of memory access during loop execution.

[0030] Preferably, the reuse distance distribution calculation includes the following steps:

[0031] For the memory access in the loop part of the program, before the loop starts, record the memory access situation before the loop starts through the created data structure. In the first iteration of the loop, record the memory access address of the current iteration. Starting from the second iteration of the loop, record the memory access address of the current iteration, and compare the memory access address of the second iteration with the memory access address of the first iteration to calculate the reuse distance distribution of the second iteration. For the Nth iteration of the loop, multiply the reuse distance distribution of the second iteration by N - 1 to obtain the reuse distance distribution of the Nth iteration, where N is greater than or equal to 3;

[0032] For the memory access in the non - loop part of the program, each time a memory access occurs, store the accessed memory access address in the created data structure. For each memory access, search for the same memory access address among the previous memory accesses. If the same memory access address is found, calculate the reuse distance distribution between the current access and the previous access. If the same memory access address is not found, set the value of the reuse distance distribution to -1;

[0033] Merge the reuse distance distributions of the loop part and the non - loop part to obtain the memory access reuse distance distribution of the entire program.

[0034] In a second aspect, the present invention provides a program static analysis system based on LLVM, which is used to perform a program static analysis method based on LLVM as described in any item of the first aspect. The system includes an IR file generation module, an IR file reading module, a control flow graph generation module, a basic block calculation module, a static memory tracking module, and a reuse distance distribution calculation module;

[0035] The IR file generation module is used to perform the following: process the target program through program instructions and convert the source file into an LLVM IR file with source - level debugging information;

[0036] The IR file reading module is used to perform the following: take the LLVM IR file and the function name of the target function as inputs, parse the LLVM IR file through a file reader, traverse the modules, functions, basic blocks, and instructions in the LLVM IR file to obtain the static tracking information of the target function;

[0037] The control flow graph generation module is used to perform the following: Based on the static trace information of the target function, calculate the instruction types executed by the basic blocks, determine the successor and predecessor relationships of the basic blocks, construct the basic block-level control flow graph of the LLVM IR file with the basic blocks of the target function as nodes and the control flow directions between the nodes as edges, and annotate relevant information, where the relevant information includes the transfer probability between basic blocks, the execution times of basic blocks, and the memory access information;

[0038] The basic block calculation module is used to perform the following: Determine the transfer probability between basic blocks by analyzing the instructions executed by the basic blocks, and calculate the prior probability of each basic block being executed based on the transfer probability between basic blocks;

[0039] Static memory tracking: Analyze the control flow graph, traverse the paths based on the transfer probability between basic blocks and find the path with the maximum probability, identify and annotate loop information when traversing the paths, obtain the execution paths with loop annotations, replace the basic blocks in the execution paths with the corresponding memory access information, and generate the static memory tracking with loop annotations;

[0040] The reuse distance distribution calculation module is used to perform the following: Calculate the memory access reuse distance distribution based on the static memory tracking with loop annotations through a recursive algorithm.

[0041] Preferably, the control flow graph generation module is used to perform the following operations:

[0042] Analyze the instruction types executed by each basic block;

[0043] For each basic block, analyze the last instruction in the basic block to determine the predecessor and successor basic blocks of the basic block;

[0044] Identify all possible predecessor basic blocks of each basic block, and record the position of the first instruction in the source program in the successor or predecessor basic block;

[0045] For each basic block, use all possible predecessor basic blocks of the basic block and the position of the first instruction in the source program in the corresponding successor or predecessor basic block as the predecessor and successor information of the basic block, construct the control flow graph of the basic blocks in the target function based on the predecessor and successor information of the basic block, and annotate relevant information in the control flow graph.

[0046] Preferably, the basic block calculation module is used to perform the following operations:

[0047] Analyze the instructions executed by the basic blocks through an offline coverage tool to obtain the transfer probability between basic blocks;

[0048] For each basic block in the control flow graph, a linear balance equation is constructed based on the transition probability from its predecessor basic block to the basic block, the transition probability from the basic block to its successor basic block, and the execution times of each basic block. The linear balance equation is expressed as:

[0049] ,

[0050] where, represents the transition probability from the predecessor basic block to the basic block , represents the transition probability from the basic block to the successor basic block , represents the execution times of the predecessor basic block , represents the execution times of the successor basic block , represents the predecessor basic block, represents the successor basic block;

[0051] The entry basic block of the program is executed only once. For the basic blocks in the control flow graph, homogeneous linear equations and a non - homogeneous equation can be formed. Starting from the entry basic block of the program, the linear balance equation is recursively solved to obtain the execution times of all basic blocks, and then the prior probability of a basic block being executed is calculated. The calculation formula for the prior probability of a basic block being executed is expressed as:

[0052] .

[0053] Preferably, the static memory tracking module is used to perform the following operations:

[0054] Read the control flow graph, start traversing from the entry basic block. When encountering multiple possible edges, select the edge with the maximum transition probability and mark it as visited according to the transition probability between basic blocks. The next time the same basic block is visited, select the unvisited edge with the maximum transition probability to find the path with the maximum probability;

[0055] After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundary of the loop and the loop control variable, and obtain the execution path with loop annotations;

[0056] Replace the basic blocks in the execution path with memory access information to obtain a static memory tracking with loop annotations, so as to facilitate viewing the order and pattern of memory access during loop execution.

[0057] Preferably, the reuse distance distribution calculation module is used to perform the following operations:

[0058] For the memory access of the loop part in the program, before the loop starts, record the memory access situation before the loop starts through the created data structure. In the first iteration of the loop, record the memory access address of the current iteration. Starting from the second iteration of the loop, record the memory access address of the current iteration, and compare the memory access address of the second iteration with the memory access address of the first iteration to calculate the reuse distance distribution of the second iteration. For the Nth iteration of the loop, multiply the reuse distance distribution of the second iteration by N - 1 to obtain the reuse distance distribution of the Nth iteration, where N is greater than or equal to 3;

[0059] For the memory access of the non-loop part in the program, each time a memory access occurs, store the accessed memory access address in the created data structure. For each memory access, search for the same memory access address from the previous memory accesses. If the same memory access address is found, calculate the reuse distance distribution between the current access and the previous access. If the same memory access address is not found, set the value of the reuse distance distribution to -1;

[0060] Combine the reuse distance distributions of the loop part and the non-loop part to obtain the memory access reuse distance distribution of the entire program.

[0061] The static program analysis method and system based on LLVM of the present invention have the following advantages: parsing the generated IR file, constructing the basic block-level control flow graph of the IR file, solving the execution times of the basic blocks by solving the linear balance equation of the transition probabilities between the basic blocks, calculating the memory access reuse distance distribution using a recursive algorithm, being able to predict different characteristics of the program, achieving accurate analysis of the application program, and realizing predictive analysis of the application program at the compilation stage. Brief Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] The present invention will be further described below with reference to the drawings.

[0064] Figure 1 It is a flowchart of a static program analysis method based on LLVM for Embodiment 1;

[0065] Figure 2Schematic diagram of a control flow graph with basic blocks and transition probabilities in a program static analysis method based on LLVM in Embodiment 1;

[0066] Figure 3 Schematic diagram of static memory tracking in a program static analysis method based on LLVM in Embodiment 1.

[0067] Among them, Figure 2 In it, BB3 represents basic block 3, BB4 represents basic block 4, BB5 represents basic block 5, for.cond is the conditional judgment part of the for loop, responsible for checking whether the loop should continue to execute; For.end refers to the end flag of a single loop body; in for.cond, there is usually a comparison instruction (such as icmp) to check whether the loop control variable meets the condition to continue the loop; for.body is the loop body part of the for loop, containing the code actually executed in the loop. BB3 main:3:[for.cond] 101 3.479e-06 means that basic block 3 is on the third line of the source code of the main function, is the conditional judgment part of the loop body, 101 means that the loop body where this basic block is located needs to execute 101 times, and 3.479e-06 means the prior probability of this basic block being executed, which is calculated based on the control flow and transition probabilities of the program;

[0068] The edges between basic blocks represent the transition probabilities of different basic blocks, meaning that when the program execution reaches the end of BB3, there is a probability of 0.98999 to then execute BB4, and a probability of 0.01001 to transfer to the end of the loop of BB5 and the function returns. For each basic block, the sum of the probabilities of transferring to other successor basic blocks is 1.

[0069] Figure 3 In it, after the program is compiled, multiple basic blocks are generated according to the main body of the program. BBn refers to the nth basic block; 10-a bb3-bb4 means that the loop in basic block 3 executes 10 times and jumps to basic block 4, and the arrow represents the jump between basic blocks. Specific implementation manner

[0070] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the exemplified embodiments are not intended to limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0071] The embodiments of the present invention provide a program static analysis method and system based on LLVM, which are used to solve the technical problems of low accuracy, low efficiency, and limited application scope existing in the existing program analysis technology in estimating program performance characteristics and memory reuse distance distribution.

[0072] Example 1: A static program analysis method based on LLVM of the present invention includes six steps: IR file generation, IR file reading, control flow graph generation, basic block calculation, static memory tracking, and reuse distance distribution calculation.

[0073] Step S100, IR file generation: Process the target program through program instructions to convert the source file into an LLVM IR file with source-level debugging information.

[0074] In this embodiment, the instruction clang -emit-llvm -S -c source.c -o source.ll is used to convert the source file source.c into the IR intermediate file source.ll. The code is as follows:

[0075] int main () {

[0076] unsigned i, j, k, out, A = 10, arry

[10]

[20] ;

[0077] for (i = 0; i < 10; i++) {

[0078] out += i;

[0079] for (j = 0; j < 20; j++) {

[0080] out = out * j + A;

[0081] for (k = 0; k < 300; k++){

[0082] out += k * k;

[0083] }

[0084] arry[i][j] = out;

[0085] }

[0086] }

[0087] return 0;

[0088] }

[0089] definedso_locali32@main() #0 {

[0090] entry:

[0091] %retval = alloca i32, align 4

[0092] %i = alloca i32, align 4

[0093] %j = alloca i32, align 4

[0094] %k = alloca i32, align 4

[0095] %out = alloca i32, align 4

[0096] store i32 0, i32* %retval, align 4

[0097] store i32 0, i32* %out, align 4

[0098] store i32 0, i32* %i, align 4

[0099] br label %for.cond

[0100] for.cond: ; preds = %for.inc15, %entry

[0101] %0 = load i32, i32* %i, align 4

[0102] %cmp = icmp ult i32 %0, 10

[0103] br i1 %cmp, label %for.body, label %for.end17

[0104] for.body: ; preds = %for.cond

[0105] %1 = load i32, i32* %out, align 4

[0106] %2 = load i32, i32* %i, align 4

[0107] %add = add i32 %1, %2

[0108] store i32 %add, i32* %out, align 4

[0109] store i32 0, i32* %j, align 4

[0110] brlabel%for.cond1

[0111] for.cond1: ; preds = %for.inc12, %for.body

[0112] %3 = loadi32, i32* %j, align4

[0113] %cmp2 = icmpulti32%3, 20

[0114] bri1%cmp2, label%for.body3, label%for.end14

[0115] for.body3: ; preds = %for.cond1

[0116] %4 = loadi32, i32* %out, align4

[0117] %5 = loadi32, i32* %j, align4

[0118] %mul = muli32%4, %5

[0119] %6 = loadi32, i32* @A, align4

[0120] %add4 = addi32%mul, %6

[0121] storei32%add4, i32* %out, align4

[0122] storei320, i32* %k, align4

[0123] brlabel%for.cond5

[0124] for.cond5: ; preds = %for.inc, %for.body3

[0125] %7 = loadi32, i32* %k, align4

[0126] %cmp6 = icmpulti32%7, 300

[0127] bri1%cmp6, label%for.body7, label%for.end

[0128] for.body7: ; preds = %for.cond5

[0129] %8 = load i32, i32* %k, align 4

[0130] %9 = load i32, i32* %k, align 4

[0131] %mul8 = mul i32 %8, %9

[0132] %10 = load i32, i32* %out, align 4

[0133] %add9 = add i32 %10, %mul8

[0134] store i32 %add9, i32* %out, align 4

[0135] br label %for.inc

[0136] for.inc: ; preds = %for.body7

[0137] %11 = load i32, i32* %k, align 4

[0138] %inc = add i32 %11, 1

[0139] store i32 %inc, i32* %k, align 4

[0140] br label %for.cond5

[0141] for.end: ; preds = %for.cond5

[0142] %12 = load i32, i32* %i, align 4

[0143] %idxprom = zext i32 %12 to i64

[0144] %arrayidx = getelementptr inbounds [10 x [20 x i32]], [10 x [20 x i32]]* @arry, i64 0, i64 %idxprom

[0145] %13 = load i32, i32* %j, align 4

[0146] %idxprom10 = zero extend i32 %13 to i64

[0147] %arrayidx11 = getelementptr inbounds [20 x i32], [20 x i32]* %arrayidx, i64 0, i64 %idxprom10

[0148] %14 = load i32, i32* %out, align 4

[0149] store i32 %14, i32* %arrayidx11, align 4

[0150] br label %for.inc12

[0151] for.inc12: ; preds = %for.end

[0152] %15 = load i32, i32* %j, align 4

[0153] %inc13 = add i32 %15, 1

[0154] store i32 %inc13, i32* %j, align 4

[0155] br label %for.cond1

[0156] for.end14: ; preds = %for.cond1

[0157] br label %for.inc15

[0158] for.inc15: ; preds = %for.end14

[0159] %16 = load i32, i32* %i, align 4

[0160] %inc16 = add i32 %16, 1

[0161] store i32 %inc16, i32* %i, align 4

[0162] br label %for.cond

[0163] for.end17: ; preds = %for.cond

[0164] reti320

[0165] }。

[0166] Step S200, LLVM IR file reading: Taking the LLVM IR file and the function name of the target function as inputs, parse the LLVM IR file through a file reader, traverse the modules, functions, basic blocks, and instructions in the LLVM IR file to obtain the static trace information of the target function.

[0167] In this embodiment, the file reader reads the source.ll file, reads information such as entry / for.cond / for.body in the above file, and deeply parses the IR file using the LLVM compiler API, traversing the modules, functions, basic blocks, and instructions therein.

[0168] Step S300, control flow graph generation: Based on the static trace information of the target function, calculate the instruction types executed by the basic blocks, determine the successor and predecessor relationships of the basic blocks, construct a basic block-level control flow graph of the LLVM IR file with the basic blocks of the target function as nodes and the control flow directions between the nodes as edges, and annotate relevant information, where the relevant information includes the transfer probability between basic blocks, the execution times of basic blocks, and memory access information.

[0169] As a specific implementation of control flow graph generation, it includes the following operations:

[0170] (1) Analyze the instruction types executed by each basic block;

[0171] (2) For each basic block, analyze the last instruction in the basic block to determine the predecessor and successor basic blocks of the basic block;

[0172] (3) Identify all possible predecessor basic blocks of each basic block, and record the positions of the first instructions in the successor or predecessor basic blocks in the source program;

[0173] (4) For each basic block, use all possible predecessor basic blocks of the basic block and the positions of the first instructions in the corresponding successor or predecessor basic blocks in the source program as the predecessor and successor information of the basic block, construct the control flow graph of the basic blocks in the target function based on the predecessor and successor information of the basic blocks, and annotate relevant information in the control flow graph.

[0174] In this embodiment, after obtaining the static trace information generated by the file reader through the trace analyzer, the construction of the control flow graph (CFG) is started. First, the types of instructions executed in each basic block are calculated, such as the number of instructions for arithmetic operations, memory load / storage, branches, etc. Then, the successor and predecessor relationships of the basic blocks are determined. Taking the for.cond basic block in source.ll as an example, its last instruction is br and it has two operands. The branch condition is determined according to the comparison instruction (icmp instruction), so as to determine its two successor basic blocks (for.body and for.end15). At the same time, by analyzing the program structure, its predecessor basic block (such as the entry basic block) is determined. After determining the successor and predecessor relationships of all basic blocks, the basic block-level CFG of source.ll is constructed, where the nodes represent basic blocks, the edges represent the control flow direction, and information such as branch probability (determined according to the program logic and input, such as in the for.cond basic block, the branch probability is determined according to the loop condition and the number of iterations) and the execution times of basic blocks (initially unknown and calculated later) is marked.

[0175] Step S400, basic block calculation: Determine the transfer probability between basic blocks by analyzing the instructions executed in the basic blocks, and calculate the prior probability of each basic block being executed based on the transfer probability between basic blocks.

[0176] As a specific implementation of basic block calculation, this step includes the following operations:

[0177] (1) Analyze the instructions executed in the basic blocks through an offline coverage tool to obtain the transfer probability between basic blocks;

[0178] (2) For each basic block in the control flow graph, construct a linear balance equation based on the transfer probability from its predecessor basic block to the basic block, the transfer probability from the basic block to its successor basic block, and the execution times of each basic block. The linear balance equation is expressed as:

[0179] ,

[0180] Among them, represents the transfer probability from the predecessor basic block to the basic block , represents the transfer probability from the basic block to the successor basic block , represents the execution times of the predecessor basic block , represents the execution times of the successor basic block , represents the predecessor basic block, represents the successor basic block;

[0181] (3) The entry basic block of the program is executed only once. For the basic blocks in the control flow graph, homogeneous linear equations and a non-homogeneous equation can be formed. Starting from the entry basic block of the program, the linear balance equations are recursively solved to obtain the execution times of all basic blocks, and then the prior probability of a basic block being executed is calculated. The formula for calculating the prior probability of a basic block being executed is expressed as:

[0182] .

[0183] In this embodiment, the control flow graph is read and traversed starting from the entry basic block. When encountering the branch of the for.cond basic block, the edge with the highest probability is selected according to the previously calculated branch probability (for example, it is determined that the probability of selecting the for.body basic block is greater according to the loop condition and the number of iterations), and it is marked as visited. Continue traversing. When encountering the for.cond basic block again, select the unvisited edge with the highest probability (which may be the for.end15 basic block at this time). In this way, the path with the highest probability is found. During the process of traversing the path, loops (such as for loops) are identified, loop symbols are marked (such as [100~i indicates the start of the loop, ] indicates the end of the loop), the loop boundary (such as the number of loop iterations is 100) and the loop control variable (such as i) are determined, and the execution path with loop annotations is obtained. Finally, the basic blocks in the execution path are replaced with the corresponding memory accesses to generate the static memory trace of the loop annotation.

[0184] Step S500, Static Memory Trace: Analyze the control flow graph, traverse the path based on the transition probability between basic blocks and find the path with the highest probability. When traversing the path, identify and annotate loop information to obtain the execution path with loop annotations, and replace the basic blocks in the execution path with the corresponding memory access information to generate the static memory trace of the loop annotation.

[0185] As a specific implementation of static memory trace, this step includes the following operations:

[0186] (1) Read the control flow graph and start traversing from the entry basic block. When encountering multiple possible edges, select the edge with the highest transition probability according to the transition probability between basic blocks and mark it as visited. When visiting the same basic block again next time, select the unvisited edge with the highest transition probability to find the path with the highest probability;

[0187] (2) After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundary of the loop and the loop control variable, and obtain the execution path with loop annotations;

[0188] (3) Replace the basic blocks in the execution path with memory access information to obtain a static memory trace with loop annotations, which facilitates viewing the order and pattern of memory accesses during loop execution.

[0189] In this embodiment, the control flow graph is analyzed through this step to achieve static memory tracing. Each node in the graph represents a basic block of the target function, the edges represent the direction from one node to another, and the probability of selecting a specific edge is also shown in the graph. First, construct a virtual memory trace. Start traversing from the entry node of the CFG. When multiple possible edges are encountered, select the edge with the highest probability and mark it as visited. The next time the same node is visited, select another unvisited edge with the highest probability to find the path with the highest probability. After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundary of the loop and the loop control variable. For example, for a simple for loop, mark the start and end basic blocks of the loop, as well as information such as the range of change of the loop variable. Generate a static memory trace with loop annotations. Replace the basic blocks in the path with the corresponding memory access information to obtain a static memory trace with loop annotations. In this way, the order and pattern of memory accesses during loop execution can be clearly seen, providing a basis for subsequent analyses such as calculating the reuse distance distribution.

[0190] Step S600, calculating the reuse distance distribution: Calculate the memory access reuse distance distribution based on the static memory trace with loop annotations through a recursive algorithm.

[0191] As a specific implementation of calculating the reuse distance distribution, this step includes the following operations:

[0192] (1) For the memory accesses in the loop part of the program, before the loop starts, record the memory access situation before the loop starts through the created data structure. In the first iteration of the loop, record the memory access address of the current iteration. Starting from the second iteration of the loop, record the memory access address of the current iteration, and compare the memory access address of the second iteration with that of the first iteration to calculate the reuse distance distribution of the second iteration. For the Nth iteration of the loop, multiply the reuse distance distribution of the second iteration by N - 1 to obtain the reuse distance distribution of the Nth iteration, where N is greater than or equal to 3;

[0193] (2) For the memory accesses in the non-loop part of the program, each time a memory access occurs, store the accessed memory access address in the created data structure. For each memory access, search for the same memory access address among the previous memory accesses. If the same memory access address is found, calculate the reuse distance distribution between the current access and the previous access. If the same memory access address is not found, set the value of the reuse distance distribution to -1;

[0194] (3) Combine the reuse distance distributions of the loop part and the non-loop part to obtain the memory access reuse distance distribution of the entire program.

[0195] The method of this embodiment accurately predicts different characteristics of the program and estimates the memory reuse distance of the program by analyzing the LLVM IR intermediate file. The system analyzes the control flow graph of the basic blocks of the generated IR file, solves the execution times of the basic blocks by solving the linear balance equation of the transition probabilities between the basic blocks, and finally calculates the memory reuse distance using a recursive algorithm, thereby realizing the accurate analysis of the application program.

[0196] Embodiment 2: A program static analysis system based on LLVM of the present invention includes an IR file generation module, an IR file reading module, a control flow graph generation module, a basic block calculation module, a static memory tracking module, and a reuse distance distribution calculation module.

[0197] The IR file generation module is used to perform the following: Process the target program through program instructions to convert the source file into an LLVM IR file with source-level debugging information.

[0198] In this embodiment, the instruction clang -emit-llvm -S -c source.c -o source.ll is used to convert the source file source.c into the IR intermediate file source.ll. The code is as follows:

[0199] int main () {

[0200] unsigned i, j, k, out, A = 10, arry

[10]

[20] ;

[0201] for (i = 0; i < 10; i++) {

[0202] out += i;

[0203] for (j = 0; j < 20; j++) {

[0204] out = out * j + A;

[0205] for (k = 0; k < 300; k++){

[0206] out += k * k;

[0207] }

[0208] arry[i][j] = out;

[0209] }

[0210] }

[0211] return 0;

[0212] }

[0213] definedso_locali32@main() #0 {

[0214] entry:

[0215] %retval = allocai32, align4

[0216] %i = allocai32, align4

[0217] %j = allocai32, align4

[0218] %k = allocai32, align4

[0219] %out = allocai32, align4

[0220] storei320, i32* %retval, align4

[0221] storei320, i32* %out, align4

[0222] storei320, i32* %i, align4

[0223] brlabel%for.cond

[0224] for.cond: ; preds = %for.inc15, %entry

[0225] %0 = loadi32, i32* %i, align4

[0226] %cmp = icmpulti32%0, 10

[0227] bri1%cmp, label%for.body, label%for.end17

[0228] for.body: ; preds = %for.cond

[0229] %1 = load i32, i32* %out, align 4

[0230] %2 = load i32, i32* %i, align 4

[0231] %add = add i32 %1, %2

[0232] store i32 %add, i32* %out, align 4

[0233] store i32 0, i32* %j, align 4

[0234] br label %for.cond1

[0235] for.cond1: ; preds = %for.inc12, %for.body

[0236] %3 = load i32, i32* %j, align 4

[0237] %cmp2 = icmp ult i32 %3, 20

[0238] br i1 %cmp2, label %for.body3, label %for.end14

[0239] for.body3: ; preds = %for.cond1

[0240] %4 = load i32, i32* %out, align 4

[0241] %5 = load i32, i32* %j, align 4

[0242] %mul = mul i32 %4, %5

[0243] %6 = load i32, i32* @A, align 4

[0244] %add4 = add i32 %mul, %6

[0245] store i32 %add4, i32* %out, align 4

[0246] store i32 0, i32* %k, align 4

[0247] br label %for.cond5

[0248] for.cond5: ; preds = %for.inc, %for.body3

[0249] %7 = loadi32, i32* %k, align4

[0250] %cmp6 = icmpulti32%7, 300

[0251] bri1%cmp6, label%for.body7, label%for.end

[0252] for.body7: ; preds = %for.cond5

[0253] %8 = loadi32, i32* %k, align4

[0254] %9 = loadi32, i32* %k, align4

[0255] %mul8 = muli32%8, %9

[0256] %10 = loadi32, i32* %out, align4

[0257] %add9 = addi32%10, %mul8

[0258] storei32%add9, i32* %out, align4

[0259] brlabel%for.inc

[0260] for.inc: ; preds = %for.body7

[0261] %11 = loadi32, i32* %k, align4

[0262] %inc = addi32%11, 1

[0263] storei32%inc, i32* %k, align4

[0264] brlabel%for.cond5

[0265] for.end: ; preds = %for.cond5

[0266] %12 = load i32, i32* %i, align 4

[0267] %idxprom = zext i32 %12 to i64

[0268] %arrayidx = getelementptr inbounds [10 x [20 x i32]], [10 x [20 x i32]]* @arry, i64 0, i64 %idxprom

[0269] %13 = load i32, i32* %j, align 4

[0270] %idxprom10 = zext i32 %13 to i64

[0271] %arrayidx11 = getelementptr inbounds [20 x i32], [20 x i32]* %arrayidx, i64 0, i64 %idxprom10

[0272] %14 = load i32, i32* %out, align 4

[0273] store i32 %14, i32* %arrayidx11, align 4

[0274] br label %for.inc12

[0275] for.inc12: ; preds = %for.end

[0276] %15 = load i32, i32* %j, align 4

[0277] %inc13 = add i32 %15, 1

[0278] store i32 %inc13, i32* %j, align 4

[0279] br label %for.cond1

[0280] for.end14: ; preds = %for.cond1

[0281] br label %for.inc15

[0282] for.inc15: ; preds = %for.end14

[0283] %16 = load i32, i32* %i, align 4

[0284] %inc16 = add i32 %16, 1

[0285] store i32 %inc16, i32* %i, align 4

[0286] br label %for.cond

[0287] for.end17: ; preds = %for.cond

[0288] ret i32 0

[0289] }}

[0290] The IR file reading module is used to perform the following: taking the LLVM IR file and the function name of the target function as inputs, parsing the LLVM IR file through a file reader, traversing the modules, functions, basic blocks, and instructions in the LLVM IR file to obtain the static tracing information of the target function.

[0291] In this embodiment, the file reader reads the source.ll file, reads information such as entry / for.cond / for.body in the above file, and deeply parses the IR file using the LLVM compiler API, traversing the modules, functions, basic blocks, and instructions therein.

[0292] The control flow graph generation module is used to perform the following: based on the static tracing information of the target function, calculate the instruction types executed by the basic blocks, determine the successor and predecessor relationships of the basic blocks, construct a basic block-level control flow graph of the LLVM IR file with the basic blocks of the target function as nodes and the control flow directions between the nodes as edges, and annotate relevant information, where the relevant information includes the transfer probability between basic blocks, the execution times of basic blocks, and memory access information.

[0293] As a specific implementation of the control flow graph generation module, this module is used to perform the following operations:

[0294] (1) Analyze the instruction types executed by each basic block;

[0295] (2) For each basic block, analyze the last instruction in the basic block to determine the predecessor and successor basic blocks of the basic block;

[0296] (3) Identify all possible predecessor basic blocks of each basic block, and record the positions of the first instructions in the successor or predecessor basic blocks in the source program;

[0297] For each basic block, use all possible predecessor basic blocks of the basic block and the positions of the first instructions in the corresponding successor or predecessor basic blocks in the source program as the predecessor and successor information of the basic block. Based on the predecessor and successor information of the basic block, construct the control flow graph of the basic blocks in the target function, and mark the relevant information in the control flow graph.

[0298] In this embodiment, after obtaining the static trace information generated by the file reader through the trace analyzer, start constructing the control flow graph (CFG). First, calculate the types of instructions executed by each basic block, such as the number of arithmetic operations, memory load / store, branch and other instructions. Then, determine the successor and predecessor relationships of the basic blocks. Take the for.cond basic block in source.ll as an example. Its last instruction is br and it has two operands. Determine the branch condition according to the comparison instruction (icmp instruction), so as to determine its two successor basic blocks (for.body and for.end15). At the same time, by analyzing the program structure, determine its predecessor basic block (such as the entry basic block). After determining the successor and predecessor relationships of all basic blocks, construct the basic block-level CFG of source.ll, where the nodes represent basic blocks, the edges represent the control flow direction, and mark the branch probability (determined according to the program logic and input. For example, in the for.cond basic block, determine the branch probability according to the loop condition and the number of iterations) and the number of times the basic block is executed (initially unknown and calculated later), etc.

[0299] The basic block calculation module is used to perform the following: Determine the transition probability between basic blocks by analyzing the instructions executed by the basic blocks, and calculate the prior probability of each basic block being executed based on the transition probability between basic blocks.

[0300] As a specific implementation of the basic block calculation module, this module is used to perform the following operations:

[0301] (1) Analyze the instructions executed by the basic blocks through an offline coverage tool to obtain the transition probability between basic blocks;

[0302] (2) For each basic block in the control flow graph, construct a linear balance equation based on the transition probability from its predecessor basic block to the basic block, the transition probability from the basic block to its successor basic block, and the number of times each basic block is executed. The linear balance equation is expressed as:

[0303] ,

[0304] where, represents the transition probability from the predecessor basic block to the basic block , represents the basic block to the successor basic block The transfer probability represents the execution count of the predecessor basic block and represents the execution count of the successor basic block

[0305] (3)The entry basic block of the program is executed only once. For the basic blocks in the control flow graph,

[0306] homogeneous linear equations and a non - homogeneous equation can be formed. Starting from the entry basic block of the program, the linear balance equations are recursively solved to obtain the execution counts of all basic blocks, and then the prior probability of a basic block being executed is calculated. The formula for calculating the prior probability of a basic block being executed is expressed as:

[0307] .

[0308] In this embodiment, the control flow graph is read, and traversal starts from the entry basic block. When encountering a branch of the for.cond basic block, the edge with the highest probability is selected according to the previously calculated branch probability (for example, it is determined that the probability of selecting the for.body basic block is greater based on the loop condition and the number of iterations), and it is marked as visited. Continue traversing. When encountering the for.cond basic block again, select the unvisited edge with the highest probability (which may be the for.end15 basic block at this time). In this way, the path with the highest probability is found. During the process of traversing the path, loops (such as for loops) are identified, loop symbols are marked (such as [100~i indicates the start of the loop, ] indicates the end of the loop), the loop boundary (such as the loop count is 100) and the loop control variable (such as i) are determined, and the execution path with loop annotations is obtained. Finally, the basic blocks in the execution path are replaced with the corresponding memory accesses to generate the static memory trace with loop annotations.

[0308] Static memory trace: Analyze the control flow graph, traverse the path based on the transfer probability between basic blocks and find the path with the highest probability. Identify and mark loop information when traversing the path to obtain the execution path with loop annotations, and replace the basic blocks in the execution path with the corresponding memory access information to generate the static memory trace with loop annotations.

[0309] As a specific implementation of the static memory trace module, this module is used to perform the following operations:

[0310] (1) Read the control flow graph, start traversing from the entry basic block. When encountering multiple possible edges, select the edge with the highest transition probability according to the transition probability between basic blocks and mark it as visited. The next time the same basic block is visited again, select the unvisited edge with the highest transition probability to find the path with the highest probability.

[0311] (2) After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundary of the loop and the loop control variable, and obtain the execution path with loop annotations.

[0312] (3) Replace the basic blocks in the execution path with memory access information to obtain a static memory trace with loop annotations, so as to facilitate viewing the order and pattern of memory accesses during the execution of the loop.

[0313] In this embodiment, the module analyzes the control flow graph to implement static memory tracing. Each node in the graph represents a basic block of the target function, the edge represents the direction from one node to another, and the probability of selecting a specific edge is also shown in the graph. First, construct a virtual memory trace. Start traversing from the entry node of the CFG. When encountering multiple possible edges, select the edge with the highest probability and mark it as visited. The next time the same node is visited again, select another unvisited edge with the highest probability to find the path with the highest probability. After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundary of the loop and the loop control variable. For example, for a simple for loop, mark the start and end basic blocks of the loop, as well as information such as the change range of the loop variable. Generate a static memory trace with loop annotations. Replace the basic blocks in the path with the corresponding memory access information to obtain a static memory trace with loop annotations. In this way, the order and pattern of memory accesses during the execution of the loop can be clearly seen, providing a basis for subsequent analyses such as the calculation of reuse distance distribution.

[0314] The reuse distance distribution calculation module is used to perform the following: calculate the memory access reuse distance distribution based on the static memory trace with loop annotations through a recursive algorithm.

[0315] As a specific implementation of the reuse distance distribution calculation module, this module is used to perform the following operations:

[0316] (1) For the memory access in the loop part of the program, before the loop starts, record the memory access situation before the loop starts through the created data structure. In the first iteration of the loop, record the memory access address of the current iteration. Starting from the second iteration of the loop, record the memory access address of the current iteration, and compare the memory access address of the second iteration with that of the first iteration to calculate the reuse distance distribution of the second iteration. For the Nth iteration of the loop, multiply the reuse distance distribution of the second iteration by N - 1 to obtain the reuse distance distribution of the Nth iteration, where N is greater than or equal to 3.

[0317] (2) For the memory access in the non-loop part of the program, each time a memory access occurs, store the accessed memory access address in the created data structure. For each memory access, search for the same memory access address among the previous memory accesses. If the same memory access address is found, calculate the reuse distance distribution between the current access and the previous access. If the same memory access address is not found, set the value of the reuse distance distribution to -1.

[0318] (3) Combine the reuse distance distributions of the loop part and the non-loop part to obtain the memory access reuse distance distribution of the entire program.

[0319] The system of this embodiment accurately predicts different characteristics of the program and estimates the memory reuse distance of the program through the analysis of the LLVM IR intermediate file. The system analyzes the control flow graph of the basic blocks of the generated IR file, solves the execution times of the basic blocks by solving the linear balance equation of the transition probabilities between the basic blocks, and finally calculates the memory reuse distance using a recursive algorithm, thereby realizing the accurate analysis of the application program.

[0320] The above has introduced in detail the static analysis method and system of the program based on LLVM. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A static program analysis method based on LLVM, characterized in that, It includes the following steps: IR file generation: Process the target program through program instructions to convert the source file into an LLVM IR file with source-level debugging information; IR file reading: Take the LLVM IR file and the function name of the target function as inputs, parse the LLVM IR file through a file reader, traverse the modules, functions, basic blocks, and instructions in the LLVM IR file to obtain the static trace information of the target function; Control flow graph generation: Based on the static trace information of the target function, calculate the instruction types executed by the basic blocks, determine the successor and predecessor relationships of the basic blocks, construct a basic block-level control flow graph of the LLVM IR file with the basic blocks of the target function as nodes and the control flow directions between the nodes as edges, and annotate relevant information, where the relevant information includes the transfer probability between basic blocks, the execution times of basic blocks, and memory access information; Basic block calculation: Determine the transfer probability between basic blocks by analyzing the instructions executed by the basic blocks, and calculate the prior probability of each basic block being executed based on the transfer probability between basic blocks; Static memory tracing: Analyze the control flow graph, traverse the paths based on the transfer probability between basic blocks and find the path with the maximum probability, identify and annotate loop information when traversing the paths to obtain the execution path with loop annotations, replace the basic blocks in the execution path with corresponding memory access information, and generate the static memory tracing with loop annotations; Reuse distance distribution calculation: Calculate the memory access reuse distance distribution based on the static memory tracing with loop annotations through a recursive algorithm.

2. The static program analysis method based on LLVM according to claim 1, wherein The control flow graph generation includes the following steps: Analyze the instruction types executed by each basic block; For each basic block, analyze the last instruction in the basic block to determine the predecessor and successor basic blocks of the basic block; Identify all possible predecessor basic blocks of each basic block, and record the position of the first instruction in the source program in the successor or predecessor basic block; For each basic block, use all possible predecessor basic blocks of the basic block and the position of the first instruction in the source program in the corresponding successor or predecessor basic block as the predecessor and successor information of the basic block, construct the control flow graph of the basic blocks in the target function based on the predecessor and successor information of the basic block, and annotate relevant information in the control flow graph.

3. The static program analysis method based on LLVM according to claim 1, wherein The basic block calculation includes the following steps: Analyze the instructions executed by the basic blocks through an offline coverage tool to obtain the transfer probability between basic blocks; For each basic block in the control flow graph, construct a linear balance equation based on the transfer probability from its predecessor basic block to the basic block, the transfer probability from the basic block to its successor basic block, and the execution times of each basic block. The linear balance equation is expressed as: , Among them, represents the transfer probability from the predecessor basic block to the basic block The transfer probability, represents the basic block to the successor basic block The transfer probability, represents the execution times of the predecessor basic block The execution times, represents the execution times of the successor basic block The execution times, represents the predecessor basic block, represents the successor basic block; The entry basic block of the program is executed only once. For the basic blocks in the control flow graph, homogeneous linear equations and a non-homogeneous equation can be formed. Starting from the entry basic block of the program, the linear balance equations are recursively solved to obtain the execution times of all basic blocks, and then the prior probability of a basic block being executed is calculated. The formula for calculating the prior probability of a basic block being executed is expressed as: 。 4. The static program analysis method based on LLVM according to claim 1, wherein The static memory tracing includes the following steps: Read the control flow graph, start traversing from the entry basic block, when encountering multiple possible edges, select the edge with the maximum transfer probability according to the transfer probability between basic blocks and mark it as visited. The next time the same basic block is visited again, select the unvisited edge with the maximum transfer probability to find the path with the maximum probability; After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundaries of the loop and the loop control variable, and obtain the execution path with loop annotations. Replace the basic blocks in the execution path with memory access information to obtain a static memory trace with loop annotations, which is convenient for viewing the order and pattern of memory accesses during loop execution.

5. The static program analysis method based on LLVM according to claim 1, wherein The reuse distance distribution calculation includes the following steps: For the memory accesses in the loop part of the program, before the loop starts, record the memory access situation before the loop starts through the created data structure. In the first iteration of the loop, record the memory access address of the current iteration. Starting from the second iteration of the loop, record the memory access address of the current iteration, and compare the memory access address of the second iteration with that of the first iteration to calculate the reuse distance distribution of the second iteration. For the Nth iteration of the loop, multiply the reuse distance distribution of the second iteration by N - 1 to obtain the reuse distance distribution of the Nth iteration, where N is greater than or equal to 3. For the memory accesses in the non-loop part of the program, each time a memory access occurs, store the memory access address of the access in the created data structure. For each memory access, search for the same memory access address among the previous memory accesses. If the same memory access address is found, calculate the reuse distance distribution between the current access and the previous access. If the same memory access address is not found, set the value of the reuse distance distribution to -1. Merge the reuse distance distributions of the loop part and the non-loop part to obtain the memory access reuse distance distribution of the entire program.

6. A program static analysis system based on LLVM, characterized in that, A system for a static analysis method of a program based on LLVM as described in any one of claims 1 - 5, the system includes an IR file generation module, an IR file reading module, a control flow graph generation module, a basic block calculation module, a static memory trace module, and a reuse distance distribution calculation module; The IR file generation module is used to perform the following: Process the target program through program instructions to convert the source file into an LLVM IR file with source-level debugging information. The IR file reading module is used to perform the following: Take the LLVM IR file and the function name of the target function as inputs, parse the LLVM IR file through a file reader, traverse the modules, functions, basic blocks, and instructions in the LLVM IR file to obtain the static trace information of the target function. The control flow graph generation module is used to perform the following: Based on the static trace information of the target function, calculate the instruction types executed by the basic blocks, determine the successor and predecessor relationships of the basic blocks, construct a basic block-level control flow graph of the LLVM IR file with the basic blocks of the target function as nodes and the control flow direction between the nodes as edges, and annotate relevant information, where the relevant information includes the transfer probability between basic blocks, the number of times the basic block is executed, and the memory access information. The basic block calculation module is used to perform the following: Determine the transfer probability between basic blocks by analyzing the instructions executed by the basic blocks, and calculate the prior probability of each basic block being executed based on the transfer probability between basic blocks. Static memory tracking: Analyze the control flow graph, traverse the paths based on the transition probabilities between basic blocks and find the path with the maximum probability. Identify and annotate loop information when traversing the path to obtain the execution path with loop annotations. Replace the basic blocks in the execution path with corresponding memory access information to generate the static memory tracking with loop annotations; The reuse distance distribution calculation module is used to perform the following: Based on the static memory tracking with loop annotations, calculate the memory access reuse distance distribution through a recursive algorithm.

7. The static program analysis system based on LLVM according to claim 6, wherein The control flow graph generation module is used to perform the following operations: Analyze the instruction types executed by each basic block; For each basic block, analyze the last instruction in the basic block to determine the predecessor and successor basic blocks of the basic block; Identify all possible predecessor basic blocks of each basic block, and record the positions of the first instructions in the successor or predecessor basic blocks in the source program; For each basic block, use all possible predecessor basic blocks of the basic block and the positions of the first instructions in the corresponding successor or predecessor basic blocks in the source program as the predecessor and successor information of the basic block. Based on the predecessor and successor information of the basic block, construct the control flow graph of the basic block in the target function, and annotate relevant information in the control flow graph.

8. The static program analysis system based on LLVM according to claim 6, wherein The basic block calculation module is used to perform the following operations: Analyze the instructions executed by the basic block through an offline coverage tool to obtain the transition probabilities between basic blocks; For each basic block in the control flow graph, construct a linear balance equation based on the transition probability from its predecessor basic block to the basic block, the transition probability from the basic block to its successor basic block, and the execution times of each basic block. The linear balance equation is expressed as: , Among them, represents the transition probability from the predecessor basic block to the basic block The transition probability, represents the basic block to the successor basic block The transition probability, represents the execution times of the predecessor basic block The execution times, represents the execution times of the successor basic block The execution times, represents the predecessor basic block, represents the successor basic block; The entry basic block of the program is only executed once. For the basic blocks in the control flow graph, homogeneous linear equations and a non-homogeneous equation can be formed. Starting from the entry basic block of the program, the linear balance equations are recursively solved to obtain the execution times of all basic blocks, and then the prior probability of a basic block being executed is calculated. The calculation formula for the prior probability of a basic block being executed is expressed as: 。 9. The static program analysis system based on LLVM according to claim 6, wherein The static memory tracking module is used to perform the following operations: Read the control flow graph, start traversing from the entry basic block. When encountering multiple possible edges, select the edge with the maximum transition probability according to the transition probabilities between basic blocks and mark it as visited. When visiting the same basic block again next time, select the unvisited edge with the maximum transition probability to find the path with the maximum probability; After finding the path, identify the loop structure in the path. If a loop is found, use the information of the loop control variable to annotate the path for the loop, determine the boundaries of the loop and the loop control variable to obtain the execution path with loop annotations; Replace the basic blocks in the execution path with memory access information to obtain the static memory tracking with loop annotations, which is convenient for viewing the order and pattern of memory accesses during loop execution.

10. The static program analysis system based on LLVM according to claim 6, wherein The reuse distance distribution calculation module is used to perform the following operations: For the memory accesses in the loop part of the program, before the loop starts, record the memory access situation before the loop starts through the created data structure. In the first iteration of the loop, record the memory access address of the current iteration. Starting from the second iteration of the loop, record the memory access address of the current iteration, and compare the memory access address of the second iteration with the memory access address of the first iteration to calculate the reuse distance distribution of the second iteration. For the Nth iteration of the loop, multiply the reuse distance distribution of the second iteration by N - 1 to obtain the reuse distance distribution of the Nth iteration, where N is greater than or equal to 3; For the memory access of the non-loop part in the program, each time a memory access occurs, store the accessed memory access address in the created data structure. For each memory access, search for the same memory access address among the previous memory accesses. If the same memory access address is found, calculate the reuse distance distribution between the current access and the previous access. If the same memory access address is not found, set the value of the reuse distance distribution to -1; Merge the reuse distance distributions of the loop part and the non-loop part to obtain the memory access reuse distance distribution of the entire program.

Citation Information

Patent Citations

  • Software vulnerability detection method based on static analysis and dynamic analysis

    CN116049831A

  • Cross-platform tensor program performance prediction method

    CN118426872A

Cited By

  • Running analysis method and system for CAD (Computer Aided Design) path-finding algorithm connection program

    CN121257049A

  • Method and system for run analysis of CAD routing algorithm connection program

    CN121257049B