Incremental defect detection method and system based on function abstract

Through the incremental defect detection method based on function summary, the path explosion problem is solved using symbolic expression graphs and virtual node technology, efficient incremental analysis is achieved, storage and calculation overhead is reduced, and the response efficiency and result consistency of defect detection are improved.

CN120336152APending Publication Date: 2025-07-18XIAMEN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510426434.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The lack of path-sensitive incremental analysis algorithms with high scalability in the prior art leads to path explosion problems, resulting in complex management and storage of intermediate analysis results. Incremental analysis is too expensive and information redundant when processing intermediate states, making it difficult to ensure the consistency and effectiveness of analysis results among different versions.

Method used

The incremental defect detection method based on function summary is adopted, and the path conditions are implicitly stored through symbol expression graph SEG and virtual node technology, combined with differential detection and incremental impact analysis at the IR level, the affected function collection is identified, the path conditions are restored, and the defect detection is performed using the SMT solver.

Benefits of technology

It significantly reduces the storage and computing overhead of path-sensitive analysis, improves the response efficiency of defect detection and the consistency of analysis results, ensures the efficiency and accuracy of incremental analysis, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336152A_ABST
    Figure CN120336152A_ABST
Patent Text Reader

Abstract

According to the incremental defect detection method and system based on the function abstract, a function abstract design step, an SEG-based path condition storage method and a virtual node-based side effect abstract technology are adopted, a complete function call graph is constructed and stored as the function abstract, and the analysis precision and expandability are effectively improved; in the code change detection step, aiming at the differential detection requirement of the program, efficient identification of program version change is realized through an IR structured comparison method; an increment influence analysis step: based on the difference information, identifying an influence range brought by the change, and ensuring high efficiency and accuracy of increment analysis; and an incremental analysis step: recovering a path condition based on the behavior of the function abstract, recovering the complete program of the influenced program, designing a defect detection algorithm based on an SMT solver and the persistent function abstract, and realizing automatic detection and verification of potential defects of the program.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer software, and particularly to an incremental defect detection method and system based on function summaries. Background Art

[0002] Software security issues refer to various security vulnerabilities and risks that may arise during the design, development, deployment, and maintenance of software. These issues include not only defects in the software itself, such as vulnerabilities in the code, configuration errors, etc., but also the risks of external attacks on the software during operation, such as malware, cyberattacks, data leakage, etc. With the popularization of the Internet and the rapid development of information technology, software security issues have gradually become a major challenge faced by enterprises and individuals. To ensure the security of software systems, developers need to adopt various methods to detect and fix potential security hazards, and static analysis is one of the commonly used important means.

[0003] Static analysis is to identify potential security vulnerabilities and code defects by analyzing the source code or binary code itself without running the program. Static analysis tools scan the structure, logic, and possible errors of the code, and can discover and fix problems at an early stage of software development, avoiding potential security risks in the later stage. Compared with other methods, static analysis has significant advantages. First, static analysis can discover problems in advance during the development process because it checks the source code without actually executing the program. This enables developers to identify and fix security issues while writing the code, thereby reducing the later repair cost. Second, static analysis can comprehensively scan the entire code library, detect all code paths, functions, and variables, and provide a more comprehensive vulnerability detection. In addition, static analysis does not depend on the running environment and only requires the source code or binary file, which is applicable to occasions where the running environment cannot be provided.

[0004] In practice, to ensure analysis accuracy while reducing the impact on development efficiency, static analysis is usually integrated into the nightly build process, and a 5- to 10-hour time window is used to deeply analyze the code library. This method can continuously discover potential defects during the development cycle and provide an analysis report the next day to help developers fix problems as early as possible and ensure the security and stability of the software.

[0005] In the actual software development process, the response speed of static analysis is crucial. If the analysis results cannot be provided in a timely manner, the lag in defect reports will have an adverse impact on development efficiency and cost. When developers fix defects, they usually need to re-understand the context of the relevant code. As time goes by, developers may have switched to other tasks, and their familiarity with the original code has decreased, resulting in a more time-consuming and error-prone repair process. In addition, delayed defect reports may also reduce the effectiveness of the testing work already invested. If defects are discovered at a later stage of development, some tests may have been executed based on the vulnerable code, rendering the test results worthless and even requiring additional testing resources for re-verification. Therefore, improving the response speed of static analysis and ensuring that defects can be fed back as early as possible can effectively reduce the cognitive cost of developers while improving the reliability of testing and the overall development efficiency.

[0006] However, between two consecutive versions of the codebase, most of the analyzed code is the same. Therefore, it is very reasonable to adopt incremental analysis and effectively reuse the previous analysis results. For small changes in the codebase, only the reachable parts related to the modification need to be re-analyzed, and the analysis results of those unreachable functions or code segments will not be affected. In addition, the results of incremental analysis must be consistent with the results of full analysis to ensure the accuracy and consistency of the analysis. Currently, multiple static analysis frameworks have integrated incremental analysis tools, which can generate analysis results within an average of several minutes or even seconds, thus significantly expanding the application boundary of static analysis technology in actual software development.

[0007] However, there is currently a lack of highly scalable path-sensitive incremental analysis algorithms. The path explosion problem poses a severe challenge to how to efficiently save and reuse intermediate analysis results. As the number of paths grows, the management and storage of analysis states become complex, and incremental analysis is prone to problems such as excessive overhead and information redundancy when dealing with intermediate states. Therefore, how to design an efficient intermediate result storage mechanism to ensure the consistency and effectiveness of analysis results between different versions is the core problem that urgently needs to be solved in the current field of path-sensitive incremental analysis. Summary of the Invention

[0008] The object of the present invention is to provide an incremental defect detection method and system based on function summaries. By implicitly storing and managing path conditions and using symbolic path conditions to avoid explicit path enumeration, the storage and computational overhead of path-sensitive analysis are significantly reduced, and the maintenance cost of path conditions approaches a linear level. In addition, a method for identifying affected functions based on the dependency relationship of the old version is proposed, which can determine the affected functions without propagating the semantic information of the new version program, improving the response efficiency of defect detection.

[0009] The present invention adopts the following technical solutions:

[0010] On the one hand, an incremental defect detection method based on function summaries includes:

[0011] Function summary design step: Based on the path condition storage method of the symbolic expression graph (SEG) and the side effect summary technology based on virtual nodes, construct a complete function call graph and save it as a function summary.

[0012] Code change detection step: Perform program difference identification at the intermediate representation (IR) level. Adopt a structured comparison method to analyze the syntax structure and semantic information of the program to achieve difference detection, identify and align the syntax units in the program, and understand the control flow and data flow dependencies between the syntax units to achieve the identification of the changed code in the program version; the syntax units include functions, basic blocks, and instructions.

[0013] Incremental impact analysis step: Utilize the cached direct call relationships and, based on the call graph of the old version program, identify the set of all affected functions that need to be re-analyzed due to code changes.

[0014] Incremental analysis step: Based on the affected programs in the set of functions that need to be re-analyzed, construct a symbolic expression graph, combine the saved function summaries, restore the path conditions, and restore the complete programs of the affected programs; encode the behavior of the restored complete programs into a form acceptable to the SMT solver and perform solving to detect potential defects in the programs.

[0015] Preferably, the path condition storage method based on the symbolic expression graph (SEG) specifically includes:

[0016] Use the symbolic expression graph (SEG) to implicitly store path conditions. The SEG uses a linear number of nodes and edges to encode all possible path conditions in the program.

[0017] Preferably, the side effect summary technology based on virtual nodes specifically includes:

[0018] Based on the virtual node method, model the function side effects; for global variables or shared states that may be modified, introduce virtual input nodes and virtual output nodes in the SEG. By establishing edges between the virtual input nodes and the virtual output nodes, depict the state change process of the global variables inside the function, so as to connect the values of the global variables to the virtual input nodes, and at the same time, when the function returns, pass its latest state through the virtual output nodes, thereby completely modeling the side effect behavior of the function.

[0019] Preferably, in the code change detection step, during the IR structured comparison process, a tentative mapping table mechanism is introduced to temporarily record variable pairs whose equivalence cannot be confirmed for the time being; when the control flow or data dependency has not been fully resolved and it is impossible to immediately determine whether two variables match, they will be included in the tentative mapping table as an assumption; as the subsequent analysis progresses, the assumption will be gradually verified. If the verification is successful, the mapping relationship will be converted into a formal confirmation; if the verification fails, it will be marked as a difference.

[0020] Preferably, in the incremental impact analysis step, identify all function sets that need to be re-analyzed due to code changes, specifically including:

[0021] Starting from the set of modified functions ModifiedFunc, initialize the set of affected functions AffectedFunc and the work list, and propagate the impact layer by layer upward along the reverse call mapping table CallerMap; through a while loop, in each round, pop a function f from the work list and traverse all callers Caller of f; if a certain caller is not included in AffectedFunc, it means that it is indirectly affected by the modified function, add it to the result set and add it to the work list to continue the propagation; this process is repeatedly executed until the work list is empty, indicating that all affected paths have been fully traversed; finally, return the complete set of affected functions AffectedFunc, that is, the set of affected functions that need to be re-analyzed.

[0022] Preferably, in the incremental analysis step, restore the path conditions, specifically including:

[0023] Gradually collect data dependency conditions and control dependency conditions along the path, and the combination of the two forms the overall condition describing the path propagation constraint; among them, for the data dependency condition, only need to judge whether two adjacent variables are equal to express the transfer relationship of data between nodes; for the control dependency condition, in addition to judging whether the boolean variable associated with the control dependency edge is consistent with the path label, it is also necessary to recursively obtain the dependency information of these boolean variables themselves to accurately describe the control dependency relationship; to ensure that the path reaches the path start node v1 from the program entry, it is necessary to further collect the control and data dependency information of v1; finally, the union of all dependency conditions constitutes a sufficient condition for the program to pass through the path and execute completely from the entry.

[0024] Preferably, in the incremental analysis step, the method for detecting defects based on SMT solving and function summaries, specifically including:

[0025] The static execution graph of the program is constructed, and then according to the configuration of the source-sink pattern, the source point set and the sink point set are determined; next, starting from each source point, its possible paths are traversed to perform the path search process; during the path search process, when a function call point is encountered, the relevant outputs of the current parameters are obtained, and the local path is connected to the path in the called function, and the path search continues recursively inside the called function; finally, by verifying whether the path conforms to the rules of the extended Dyck-CFL language and whether the path condition is satisfied, it is comprehensively judged whether there are potential defects in this path.

[0026] On the other hand, an incremental defect detection system based on function summaries includes:

[0027] A function summary design module, which is used to construct a complete function call graph based on the path condition storage method of the symbolic expression graph SEG and the side effect summary technology based on virtual nodes, and save it as a function summary;

[0028] A code change detection module, which is used to identify program differences at the intermediate representation IR level, adopt a structured comparison method, analyze the syntax structure and semantic information of the program to achieve difference detection, identify and align the syntax units in the program, and understand the control flow and data flow dependencies between the syntax units to achieve the identification of the changed code of the program version; the syntax units include functions, basic blocks and instructions;

[0029] An incremental impact analysis module, which is used to utilize the cached direct call relationship and based on the call graph of the old version program, identify all affected function sets that need to be re-analyzed due to code changes;

[0030] An incremental analysis module, which is used to construct a symbolic expression graph based on the affected programs in the function set that needs to be re-analyzed, combine the saved function summaries, restore the path conditions, and restore the complete programs of the affected programs; encode the behavior of the restored complete programs into a form acceptable to the SMT solver and solve it to detect potential defects in the programs.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] The present invention proposes a path condition storage method and side effect summary technology based on symbolic expression graphs, effectively improving the accuracy and scalability of analysis; for the differential detection requirements of programs, a code change detection method is designed and implemented, and through the Myers difference algorithm and IR structured comparison method, efficient identification of program version changes is achieved; based on the differential information, an incremental impact analysis step is further constructed to identify the scope of influence brought by the changes, ensuring the efficiency and accuracy of incremental analysis; the incremental analysis step of the present invention includes behavior restoration, behavior encoding, and SMT solving methods based on function summaries, and a defect detection algorithm is designed based on the SMT solver and persistent function summaries to achieve automated detection and verification of potential program defects. Description of the Drawings

[0033] Figure 1 It is a flowchart of the incremental defect detection method based on function summaries according to an embodiment of the present invention;

[0034] Figure 2 It is a schematic diagram of the overall process of the incremental defect detection method based on function summaries according to an embodiment of the present invention;

[0035] Figure 3 It is a SEG graph showing the control dependence relationship between θ1 and θ2 according to an embodiment of the present invention;

[0036] Figure 4 It is an example graph of a function with side effects according to an embodiment of the present invention;

[0037] Figure 5 It is a schematic diagram of the Myers difference grid path of the prior art;

[0038] Figure 6 It is an example graph of the use of the tentative mapping table according to an embodiment of the present invention;

[0039] Figure 7 It is a distribution diagram of the reverse dependence closure of functions in the OpenSSL project according to an embodiment of the present invention. Detailed Embodiments

[0040] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0041] As Figure 1 shown, an incremental defect detection method based on function summaries in this embodiment includes:

[0042] Function summary design step S101: Based on the path condition storage method of the symbolic expression graph (SEG) and the side effect summary technology based on virtual nodes, construct a complete function call graph and save it as a function summary.

[0043] Code change detection step S102: Perform program difference identification at the intermediate representation (IR) level. Adopt a structured comparison method to analyze the syntax structure and semantic information of the program to achieve difference detection, identify and align the syntax units in the program, and understand the control flow and data flow dependencies between the syntax units to achieve the identification of the changed code of the program version; the syntax units include functions, basic blocks, and instructions.

[0044] Incremental impact analysis step S103: Utilize the cached direct call relationships and, based on the call graph of the old version program, identify the set of all affected functions that need to be re-analyzed due to code changes.

[0045] Incremental analysis step S104: Based on the affected programs in the set of functions that need to be re-analyzed, construct a symbolic expression graph, combine with the saved function summary, restore the path conditions, and restore the complete program of the affected programs; encode the behavior of the restored complete program into a form acceptable to the SMT solver and perform the solution to detect potential defects in the program.

[0046] Specifically, as Figure 2 shown, the overall method flow chart of the incremental defect detection method is presented. Its core lies in identifying and isolating code changes to avoid comprehensive static analysis of the entire project, thereby achieving a significant improvement in analysis efficiency. The method locates the incremental change areas in the source code by comparing code versions and, based on the existing historical analysis results, only re-analyzes the affected parts. Since the behavior of the vast majority of functions remains unchanged between the two versions, the analysis results of these codes can be directly reused, effectively reducing the consumption of computing resources while ensuring the consistency and accuracy of the analysis. Through incremental analysis, computationally intensive tasks such as path-sensitive data dependence analysis in the static analysis process can be significantly accelerated, thereby improving the overall analysis efficiency while maintaining the analysis accuracy. Subsequently, by using a difference detection tool to compare different versions of the code, the changed code units are identified, an impact analysis graph of the affected code area is constructed, and the program slices that need to be re-analyzed are determined. Specifically, the present invention adopts static program slicing technology to ensure an accurate division of the impact scope without missing potential dependency relationships. Secondly, for the identified affected areas, path-sensitive static analysis is re-executed to ensure that the analysis results can correctly reflect the latest state of the code and avoid analysis biases introduced by version updates. At the same time, the code areas that have not changed directly reuse the historical analysis data without repeated calculation.

[0047] First, the language model of the program, the Symbolic Expression Graph (SEG), path conditions, feasible paths, and source-sink patterns are defined as follows, providing a theoretical basis for the subsequent analysis process. Subsequently, pointer analysis techniques, including basic semantic transformation and sparse value flow analysis, are elaborated in detail to achieve efficient modeling of the program's data flow and control flow.

[0048] (1) Program language definition.

[0049]

[0050] As described above, the analysis in this paper is formally described through a simple language. The program is in Static Single Assignment (SSA) form. SSA is an intermediate representation form widely used in compiler optimization and static analysis. Its core idea is that each variable in the program is renamed to a new version after each assignment operation, ensuring that each variable is assigned only once. This feature simplifies data flow analysis, enabling compilers and static analysis tools to more easily implement optimizations and program verification, such as constant propagation, dead code elimination

[0051] , common subexpression elimination, etc. When dealing with branches and control flow merges in the program, SSA solves the problem of variable version selection by introducing functions. Specifically, when the program execution flow converges from different branch paths, there may be multiple versions of the same variable. To determine which version of the variable value should be used currently, SSA inserts functions at the control flow merge point. functions select the variable value in the corresponding branch according to the actual execution path of the program and assign it to a new variable version, thus ensuring the single-assignment feature and correct data dependency expression.

[0052] In the φ assignment , is the gating function for each variable vi, which means that variable v is equal to vi if and only if the condition holds. These gating functions can effectively express the selection relationship of variable values under different paths, and their calculation can be completed almost in linear time. This mechanism not only simplifies path condition management in static analysis but also significantly improves analysis efficiency and scalability.

[0053] (2) Symbolic Expression Graph definition

[0054] The Symbolic Expression Graph (SEG) is a directed graph G = (N, E c , E d, C, L), where N, E c , E d and C are defined as follows:

[0055] · N = V ∪ O is a set of nodes, where V represents variables in the program, represents boolean variables in the program. O represents operators in the program.

[0056] · is a set of edges, and each edge represents a data dependency relationship, which contains two types of edges, (v, o) ∈

[0057] V × O represents the data flow from variable v to operator o, that is, variable v is an operand of operator o. (o, v) ∈ O

[0058] × V represents the data flow from operator o to variable v, that is, variable v receives the operation result of operator o.

[0059] · represents the control dependency relationship. If ((v i , v j ), b) ∈ E c , it means that the value of boolean variable b determines whether the data can flow from v i to v j . According to the characteristics of the static single assignment form, each edge can only be associated with one boolean variable. In the present invention, denote E c [v i , v j as the boolean condition corresponding to the control data flowing from v i to v j .

[0060] · Maps a condition to each edge in the graph indicating that this value flow relationship is only valid when the condition holds.

[0061] L maps each data dependency edge to a parenthesis. The value of represents function calls and returns at different call points, corresponding to the left parenthesis and the right parenthesis respectively. For other edges,

[0062] (3) Path condition definition

[0063] For a data flow path in the symbolic expression graph G = (N, E, C, L)

[0064]

[0065] where, Indicating data dependence, the present invention defines the path condition PC(π) as follows:

[0066] The path condition PC(π) represents the constraint conditions that the variable set V needs to satisfy when data flows along the path π. The path condition is divided into two parts: the data dependence condition means that the execution result of one instruction is used by another instruction, that is, the definition-use relationship. The control dependence condition describes the boolean conditions that need to be satisfied for statement execution on the control flow path.

[0067] Generally speaking, the control dependence condition is introduced by the branch statements on the path:

[0068]

[0069] The above formula judges whether there is a control dependence edge for each node. If so, the boolean variable corresponding to the edge in E c needs to be compared with the C corresponding label, and this operation will generate a boolean expression. The conjunction of all the boolean expressions generated by the above operations can obtain the control dependence condition of the path.

[0070] The data dependence condition represents the flow of data and can be directly achieved by the conjunction of a series of equalities:

[0071]

[0072] The conjunction of the data dependence condition and the control dependence condition can obtain the complete path condition:

[0073] PC(π) = PC c (π) ∧ PC d (π)

[0074] (4) Definition of feasible path

[0075] Given the symbolic expression graph G = (N, E, C, L), the path π is feasible if and only if the following two conditions are satisfied simultaneously:

[0076] (4.1) The path condition PC(π) is satisfiable:

[0077] PC(π) is satisfiable (4.2) The label sequence L(π) of the path belongs to the extended Dyck context-free language (Extended Dyck-CFL):

[0078] L(π) ∈ Extended Dyck-CFL

[0079] Among them, Extended Dyck-CFL further ensures the correctness of the call-return matching relationship, that is, procedure calls and returns must follow legal semantic matching.

[0080] (5) Source-sink pattern definition and application

[0081] The source-sink pattern (Source-Sink) is a commonly used analysis pattern widely applied in static defect detection, and its related definitions are as follows:

[0082] · Source-sink reachability refers to given nodes v source , v sink ∈ V, to determine whether there is a feasible path between these two nodes.

[0083] · The source-sink configuration consists of a triple {ΦSoure, ΦSink, σ}. ΦSoure and ΦSink are two propositions, which respectively determine under what circumstances a node will be judged as a source node and a sink node. σ is a boolean constant, which determines how to identify the source-sink model as a vulnerability. If the value of σ is true, then the reachability satisfaction between the source node and the sink node will be identified as a vulnerability; conversely, the non-reachability between the source node and the sink node will be identified as a vulnerability.

[0084] Many common vulnerabilities can be uniformly described and detected using the source-sink model. Under this model, vulnerability detection is formalized as a problem of judging the reachability between specific source nodes (Source) and sink nodes (Sink). Different types of vulnerabilities can be expressed and detected by reasonably configuring the source nodes and sink nodes and combining the reachability conditions.

[0085] For example, information leakage vulnerabilities usually take the read operations of sensitive data (such as user privacy information, passwords, location data) as source nodes, and unauthorized output operations (such as log printing, network sending) as sink nodes. If the source node and the sink node are reachable under the path condition, it means that the data is propagated without authorization, that is, it is determined as an information leakage vulnerability.

[0086] In the Use-After-Free (UAF) vulnerability, the memory release operation (free) can be regarded as the source node, and the subsequent pointer dereference operation is the sink node. If there is path reachability between the release operation and the dereference operation, that is, there is still access to the same memory after the release, it is regarded as having a UAF vulnerability. Conversely, if the path after the release is not reachable, it means that the memory has been safely recycled.

[0087] The modeling of memory leakage vulnerabilities is the opposite. Memory allocation (malloc / new) is the source node, and memory release (free / delete) is the sink node. If there is no reachability on all paths between the allocation point and the release point, it means that the resource has not been released and there is a memory leakage. Such vulnerabilities are determined by analyzing the non-reachability between the source and sink nodes.

[0088] (6) Pointer Analysis

[0089] Pointer analysis is an important research direction in static program analysis, whose goal is to determine the set of memory objects that pointer variables in a program may point to at runtime. Pointer analysis is widely used in fields such as compilation optimization, security detection, memory leak analysis, and program verification. The pointer analysis problem can usually be divided into two sub-problems:

[0090] Points-to Analysis: Determine which memory locations a pointer variable may point to.

[0091] Alias Analysis: Determine whether two pointer expressions may point to the same memory location on a certain execution path.

[0092] (6.1) Basic Semantic Transformations

[0093] The following lists the rules for analyzing basic statements.

[0094]

[0095] Each rule has the form:

[0096]

[0097] Indicates that at the current point, under the pointing environment E, abstract store S, and path condition the statement at program point stmt produces a new pointing environment E′ and / or abstract store S′.

[0098] In these rules, this paper uses E[p7→{...}] and S[o7→{...}] to represent binding the pointer p and memory object o to new abstract values respectively.

[0099] Rule ADDR creates a memory object at the allocation point. Rule COPY updates the pointing environment E of a variable p to the pointing result through another variable q. Since this paper assumes that the code is in SSA (Static Single Assignment) form and each top-level variable has only one definition. Therefore, Rule ADDR and Rule COPY perform strong updates. Rule PHI merges the pointing information at the merge point of multiple paths. Basically, Rule COPY and Rule PHI are obvious, so this paper focuses on Rule STORE and Rule LOAD.

[0100] Rule STORE is used to handle the path condition The store statement *x = q updates the abstract store S accordingly. First, by querying to obtain the set of memory objects that variable x may point to under the path condition . For each protected memory object (π, o) in the set, update its corresponding abstract store entry to hold the value q. According to the traditional singleton-based algorithm, when variable x points to only a single concrete memory object, an indirect strong update can be performed to clear other values stored on that memory object o.

[0101] For the load statement p = *y at program point with path condition , the rule LOAD can be applied. Similar to the rule STORE, first query the memory objects that variable y may point to under the condition , denoted as . Subsequently, extract the values that may be stored in each memory object , denoted as Π π (S(o)). Finally, for each value pair , update the points-to information E to: under the path condition , the points-to set of variable p contains the value v.

[0102] The rules SEQUENCING and BRANCHING handle compound statements. The rule SEQUENC-ING means using the postcondition of the current statement as the precondition of the next statement, and the rule BRANCHING merges the abstract stores and environments of multiple paths.

[0103] (6.2) Online Sparse Value-Flow Analysis

[0104] To alleviate the problem of overly dense value-flow information, huge computational and storage overheads in traditional static analysis, the present invention adopts the Sparse Value-Flow Analysis (SVF) method. This method significantly reduces the analysis cost of redundant data dependencies and path conditions by precisely constructing a Sparse Value-Flow Graph (SVFG) on the SEG. On this basis, efficient abstract transformation rules are introduced to incrementally update the definition-use relations, thereby improving the analysis accuracy and scalability.

[0105] To formally describe this idea, this paper maintains the abstract store S as a set of , where represents the program point Abstract storage at

[0106] The STORE statement rule. The analysis process in this paper is designed based on the SSA form. Therefore, each variable can only be used at the position dominated by its definition point or in its dominance frontier. The dominance relationship is a basic concept in the dominance tree: if node a dominates node b, then all control paths from the program entry to b must pass through a. The dominance frontier refers to the set of nodes that satisfy the following conditions: a node d is not dominated by node a, but there exists a node b such that a dominates b and there is a control edge from b to d. At the dominance frontier position, variable definitions from different control paths usually need to be merged to ensure the single-assignment semantics in the SSA form. Suppose a variable with path constraints (φ, l, v) is written to the memory object o at program node l. As shown in Algorithm 1, the processing first writes this value to the local abstract storage S l (o) (line 2), and then propagates this value to the dominance frontier of node l (lines 3 - 4). It should be noted that it is not necessary to explicitly update the value to the nodes dominated by l. When processing the load statement, it traverses upward along the dominance tree to find visible definitions to ensure semantic correctness and storage consistency.

[0107]

[0108] The LOAD statement rule. As shown in Algorithm 2, for the load statement u = *x at program point , the algorithm will track the values that can be read from the memory object pointed to by x. For each memory object o, it only needs to traverse upward along the dominance tree (see lines 6 to 15) to gradually aggregate the relevant abstract values. During the traversal, the algorithm dynamically maintains the suffix of the path condition (line 14) to determine whether the subsequent nodes can overwrite the stored value at the current position under the current path. The suffix of the path condition is combined with the condition of the variable itself (line 11) to form the complete value reading condition. If a strong update is found at a certain node, it means that the value has been completely overwritten, and the algorithm can terminate the further upward search.

[0109] The following will elaborate on the function summary design step S101 in detail.

[0110] Function Summary is a commonly used technique in program analysis, especially in static analysis. Its main purpose is to summarize the behavior of a function, describe the input, output, and side effects (such as state modification, resource management, etc.) of the function in a concise way, so that when the analysis tool encounters a call to this function, it does not have to repeatedly analyze the inside of the function in depth, but can directly use the summary for reasoning and optimization.

[0111]

[0112] Abstract design often faces two challenges. One challenge is that abstract design needs to balance accuracy and space overhead. The present invention adopts an implicit method to save path conditions in SEG at a linear cost, which effectively avoids the exponential growth of path conditions as the number of branches increases.

[0113] The other is that it is difficult to accurately model function side effects. Function side effects refer to the additional impacts that a function has on the external state or environment of a program during program execution, in addition to returning the function result, such as modifying global variables, modifying parameter data, etc.

[0114] In path-sensitive analysis, since it is necessary to distinguish different paths of program execution, analysis tools usually face the problem of path explosion. To achieve path sensitivity, iSATURN directly encodes each path condition as a Boolean expression and uses a SAT solver to judge the feasibility of the path. However, this encoding method of path conditions has significant redundancy. In different paths, the same literals (i.e., basic propositional units in first-order propositional logic) will be repeatedly saved and processed, resulting in a rapid increase in the consumption of storage space and computing resources. Therefore, this method faces a scalability bottleneck in large-scale program analysis.

[0115] The present invention uses symbolic expressions to implicitly store path conditions. Through the expression method of graph structure, Boolean variables can be reused more effectively, avoiding redundancy problems in the process of encoding path conditions. As Figure 3 shown, when encoding path conditions, if it is necessary to express the existing literal θ1 can be directly reused through the control dependence edge. This structure not only reduces the repeated storage of Boolean variables but also realizes the efficient expression of path conditions by sharing control conditions on the path.

[0116] SEG can encode all possible path conditions in a program by using a linear number of nodes and edges. Compared with the traditional method of encoding path conditions path by path, it greatly improves the scalability of the analysis method. This method effectively alleviates the path explosion problem in path-sensitive analysis and provides a basis for large-scale program analysis.

[0117] In program analysis, function side effects refer to the impacts that a function has on the program state (such as global variables, pointer references, shared memory, etc.) other than its explicit parameters during function execution. Function side effects usually occur in the following situations:

[0118] (1) The function modifies global variables or static variables;

[0119] (2) The function modifies the value of the actual parameter through a pointer or reference;

[0120] (3) The function generates I / O operations or system calls, affecting the state of the external environment.

[0121] Different function side-effect handling strategies will trigger different analysis risks. If function side effects are ignored in static analysis, potential bugs caused by state changes may be missed, reducing the coverage and reliability of vulnerability detection. On the contrary, adopting an overly conservative approach (such as assuming that the function may arbitrarily modify all accessible global states) will significantly reduce the accuracy of the analysis, leading to unnecessary false alarms and path indistinguishability.

[0122] As Figure 4 shown, the incrementCounter function is a typical example with side effects. This function executes counter++ on line 5, directly modifying the state of the global variable counter, which is exactly the manifestation of function side effects. In addition, the function also simply processes the input parameter value and returns the calculation result.

[0123] To resolve the contradictions in the above side-effect analysis, the present invention introduces a virtual node method to model function side effects. For global variables or shared states that may be modified, virtual input nodes and virtual output nodes are introduced in the SEG. Taking counter as an example, when analyzing the incrementCounter function, two nodes, Counter in and Counter out will be created to represent the entry state and exit state of counter respectively.

[0124] In the SEG, by establishing an edge between Counter in and Counter out , the state change process of counter inside the function is characterized. This design enables the value of counter to be connected to the Counter in node in the call graph, and at the same time, when the function returns, its latest state is passed through the Counter out node, thus completely modeling the side-effect behavior of the function. Compared with traditional conservative analysis, this method not only ensures the accurate capture of side-effect information but also avoids unnecessary state pollution, improving the overall reliability and accuracy of static analysis.

[0125] The following will elaborate on the code change detection step S102 in detail.

[0126] The present invention performs program difference identification at the level of Intermediate Representation (IR). This is because program analysis tools usually directly analyze based on IR, and external factors such as configuration files may affect the generation of IR. If only the differences are analyzed at the source code level, the IR changes caused by compilation configurations or optimization options may be ignored, thus affecting the accuracy of difference analysis. Compared with traditional difference methods based on text strings, the difference identification at the IR level requires a comparison oriented to the program structure to determine whether two objects in different IRs are similar or comparable. In this process, the core idea of the Myers difference algorithm is borrowed. By searching for the shortest edit path, the same or similar parts in the IR instruction stream are quickly matched and identified, significantly improving the efficiency and accuracy of difference analysis.

[0127] The Myers difference algorithm is an efficient algorithm for comparing the differences between two sequences, aiming to calculate the Shortest Edit Script (SES) between them. This algorithm realizes the comparison by modeling the difference problem as a path search problem in a two-dimensional grid. The horizontal and vertical axes of the grid represent the indices of the source sequence A and the target sequence B respectively, and the coordinate point (x, y) represents the alignment state of the first x characters of the source sequence A and the first y characters of the target sequence B. The goal of the algorithm is to move from the starting point (0, 0) along the path to the ending point (N, M), where N and M are the lengths of the two sequences respectively. The movement of the path is divided into three operations: a horizontal move (along the x-axis direction) represents a deletion operation, a vertical move (along the y-axis direction) represents an insertion operation, and a diagonal move represents an element match. Only when A[x] = B[y] can a diagonal move be made, and this operation does not increase the edit distance, while insertion and deletion operations each increase the edit distance by one.

[0128] As Figure 5 shows the matching process of the Myers difference algorithm between the source sequence ABCABBA and the target sequence CBABAC. The path of the arrow along the diagonal direction in the figure represents the character matching operation, clearly showing the situation where the characters at the current positions of the two sequences are equal. It should be noted that the horizontal and vertical edges always exist in the actual search graph, corresponding to deletion and insertion operations respectively, but in order to highlight the matching path, only the diagonal arrows are drawn in the figure, and the display of other edges is omitted.

[0129] The core idea of Myers' algorithm lies in adopting the Furthest-Point-First advancement strategy. The algorithm uses the current edit distance D as the iteration level, and each time it explores all the state paths that can be reached within D edit operations. By maintaining a set of variables k representing the current "diagonal" (where k = x - y), the algorithm can effectively describe the paths formed by different combinations of insertions / deletions on the grid. Within each level D, the algorithm advances on all possible k values and preferentially performs consecutive diagonal moves to extend the path as much as possible. This strategy ensures that as many elements as possible are matched at each step, thereby reducing the search space and improving efficiency.

[0130] In the implementation process, the algorithm uses a vector to record the furthest reach point of each diagonal, and the update of this vector follows the recursive advancement rule. By continuously increasing the value of D, the algorithm finally finds a path from the starting point to the ending point, and the edit sequence corresponding to this path is the shortest edit script. The time complexity of Myers' algorithm is O(ND), where N and M are the lengths of the input sequences, and D is the edit distance. In most practical applications, D is usually much smaller than N and M, so the algorithm performs very efficiently. Its space complexity is O(N + M), and through further optimization, it can be reduced to O(D), which makes it very suitable for differential calculations of large-scale texts or files.

[0131] When performing differential calculations, Myers' algorithm relies on a basic premise that the conditions under which two elements can be considered "matched" must be predefined. In traditional text differential scenarios, this matching condition is usually very simple - as long as two strings are exactly the same literally, they are considered equal. However, this strict matching criterion often seems too harsh in intermediate representation (IR) differential analysis and is prone to overlooking code differences that should be recognized as structurally equivalent. Typical scenarios include: the order of basic blocks is adjusted but the control flow graph remains unchanged, and the program execution logic is equivalent; variable names are different, but the data flow and dependency relationships are exactly the same; without violating the SSA constraints, the instruction order is adjusted but the overall calculation logic remains unchanged. Such changes do not affect the program behavior, and it is obviously unreasonable to regard them as substantial differences. Therefore, IR differential analysis needs to introduce a more flexible and semantically aware matching mechanism to avoid missing potential equivalent structures or misreporting meaningless superficial differences.

[0132] The present invention adopts a structured comparison method to achieve differential detection by analyzing the syntax structure and semantic information of the program, rather than simply comparing based on the text content. Compared with traditional text differential analysis, structured comparison can identify and align syntactic units such as functions, basic blocks, and instructions in the program, and understand the control flow and data flow dependencies between them. This method has natural fault tolerance for equivalent transformations (such as variable renaming, instruction reordering), can accurately reveal the substantial differences in program logic or behavior, and significantly reduce the false alarm rate. Through structured analysis, the tool can more accurately align the program structures of different versions, support fine-grained difference detection and high-level semantic analysis.

[0133] The specific implementation includes the following five levels of comparison strategies:

[0134] (1) Module-level comparison first compares the top-level structures of two IR files, including global variables, type definitions, and function lists, to identify newly added, deleted, or modified global variables and functions.

[0135] (2) Function-level comparison for matching functions, further compares their function signatures (name, parameter types, return type), and analyzes the differences inside the function body, focusing on the changes in basic blocks and instructions.

[0136] (3) Basic block-level comparison disassembles the function body into basic blocks and matches them one by one. If the control flow structures of two basic blocks are inconsistent, they will be marked as different.

[0137] (4) Instruction-level comparison based on the matching of basic blocks, compares the instructions one by one. If the operators or operands are different, they will be marked as different. At the same time, changes that are irrelevant to the program semantics are ignored (such as scenarios where variable names are different but have no impact).

[0138] (5) Variable-level comparison delves into the operand level of instructions to compare the equivalence between variables. For function parameters or local variables (SSA values), it determines whether they are in a corresponding relationship through a mapping table; for global variables, it compares the names or determines equivalence according to specific rules; for constants, it recursively compares their values and structures. Variable-level comparison ensures the consistency of the instruction operation objects and guarantees the accuracy of data flow and control flow analysis.

[0139] During the IR structured comparison process, a TentativeValues mechanism is introduced to temporarily record variable pairs whose equivalence cannot be confirmed immediately. When the control flow or data dependencies have not been fully resolved and it is impossible to immediately determine whether two variables match, they will be included in the TentativeValues as an assumption. As the subsequent analysis progresses, these assumptions will be gradually verified: if the verification is successful, the mapping relationship will be converted into a formal confirmation; if the verification fails, they will be marked as different.

[0140] When processing two IR differences as shown in Figure 6 , the analyzer first recognizes that tmp1 = x + y in V2 and a = x + y in V1 are exactly the same in terms of operation type and data dependency. However, due to different variable names, their equivalence cannot be immediately confirmed. Therefore, <tmp1, a> is temporarily stored in the TentativeValues table and awaits further verification. Subsequently, the analyzer continues to process tmp2 = tmp1 * 2 and discovers that its data dependency on tmp1 is consistent with the dependency relationship and calculation logic of b = a * 2 in V1, thus verifying the equivalence of tmp1 and a. Based on this verification, the system promotes <tmp1, a> from the TentativeValues table to a formal mapping relationship and simultaneously establishes a formal mapping for <tmp2, b>. The TentativeValues table, through a delayed decision-making mechanism, ensures that mapping relationships are only confirmed after the analysis of dependency relationships and control flow is completed, effectively avoiding early misjudgments or omissions of potential equivalent structures, thereby enhancing the accuracy and robustness of differential analysis.

[0141] The following will elaborate on the incremental impact analysis step S103 in detail.

[0142] The incremental impact analysis step is used to analyze and determine the set of all functions that need to be re-analyzed due to program modifications. Given the bottom-up design of the analysis framework, the caller needs to conduct its own analysis based on the analysis results of the callee. When a function modification is detected, all direct callers of this function and functions that recursively depend on its analysis results need to trigger the re-analysis process. To efficiently identify recursively dependent functions, the present invention proposes an incremental impact analysis algorithm. This method does not rely on the call relationships of the new version of the program and can accurately determine the set of functions that need to be re-analyzed only based on the call information of the old version, enabling developers to more accurately estimate the time overhead of this incremental analysis earlier.

[0143] S1031, Inference and proof of the invariance of the set of affected function closures

[0144] Program version P original represents the version before the change, and P changed represents the version after the change. Let the set of changed functions be ChangedFunc, that is:

[0145] ChangedFunc = {f | the definition of f has changed in P original and P changed}

[0146] Define C(P, f) to represent the set of all functions that directly call function f in program P. Further define its transitive closure (i.e., all functions on the recursive call path) as C Rec(P, f) satisfies:

[0147] C Rec (P, f) = TransitiveClosure(C(P, f))

[0148] Generalize to the function set S and define as follows:

[0149]

[0150] The present invention hopes to prove that the following equation holds:

[0151] C Rec (P original , ChangedFunc) = C Rec (P changed , ChangedFunc)

[0152] In other words: when the program is modified, the upward call set of the changed function remains unchanged.

[0153] Proof. Assume that there exists a function F ∈ ChangedFunc such that:

[0154] C Rec (P original , F) ≠ C Rec (P changed , F)

[0155] Then there must exist a function Q such that Q exists only in the closure set of one of the versions.

[0156] The situation can be divided into the following two categories:

[0157] · Case 1:

[0158] Q ∈ C Rec (P original , F) and

[0159] At this time, there exists a Q2 ∈ ChangedFunc such that Q ∈ C Rec (P original , Q2), but since Q is no longer in C Rec (P changed , Q2), it indicates that Q no longer calls the changed function through the path of Q2.

[0160] However, as long as Q exists in the closure of any other Q i ∈ ChangedFunc, Q will still be included in the overall C Rec (P changed , ChangedFunc).

[0161] Case 2:

[0162] and Q ∈ C Rec (P changed , F)

[0163] At this time, there exists a Q3 ∈ ChangedFunc such that Q ∈ C Rec (P changed , Q3), and

[0164] Similarly, as long as Q appears in the call closure of another change function in the original version (i.e., Q ∈

[0165] C Rec (P original , Q i )), then it still belongs to the overall C Rec (P original , ChangedFunc). Therefore, even if there are differences in some individual closure sets in different versions, the overall transitive call closure of the set ChangedFunc will not be affected, that is: C Rec (P original , ChangedFunc) = C Rec (P changed , ChangedFunc) is proven.

[0166] S1032, Call Relationship Caching and Affected Function Deduction

[0167] In S1031, it is formally proved that: the call relationships required to identify the set of affected functions can be completely based on the call graph of the old version program without relying on the analysis results of the new version. This conclusion provides a solid theoretical basis for the caching mechanism of call relationships.

[0168] Specifically, limited by the existence of function pointers, the present invention first performs high-precision pointer analysis in the full-scale analysis to obtain accurate function pointer pointing information, and then constructs a complete function call graph. In the incremental analysis, by means of the caching mechanism of call relationships and combined with the property that the call graph remains unchanged during the identification of affected functions, the incremental propagation analysis can be advanced to the pointer analysis stage, thus significantly shortening the response time for identifying affected functions.

[0169] The core advantage of this optimization strategy is that: developers can obtain in advance the set of functions that may be affected by the current code change without waiting for the completion of the full-scale analysis, and can estimate the time-consuming of the incremental analysis. In an integrated development environment with high interactivity requirements, this ability is particularly important and helps to more efficiently play the role of incremental analysis in providing quick feedback.

[0170] The present invention will store the direct call relationships of each function, which requires accessing each call point in the function and querying the variable values involved in the corresponding call point. When a pointer variable is called, it is also necessary to further query the memory content pointed to by the pointer.

[0171]

[0172] Algorithm 3 is used to identify the set of functions that need to be re-analyzed due to changes in function-level incremental analysis. This algorithm improves efficiency by utilizing the cached direct call relationships. The algorithm starts from the set of modified functions ModifiedFunc (line 1), initializes the set of affected functions AffectedFunc and the work list (line 2), and propagates the impact layer by layer along the reverse call map CallerMap. The algorithm uses a while loop (line 3), and in each round, it pops a function f from the work list (line 4) and traverses all the callers Caller of f (line 5). If a certain caller is not yet included in AffectedFunc (line 6), it means that it is indirectly affected by the modified function, and it needs to be added to the result set (line 7) and added to the work list to continue the propagation (line 8). This process is repeated until the work list is empty (line 11), indicating that all affected paths have been fully traversed. Finally, the algorithm returns the complete set of affected functions AffectedFunc (line 12). Since each function is processed at most once, the overall time complexity of the algorithm is linear with the scale of the call graph, which is suitable for efficient impact analysis in large-scale projects.

[0173] The incremental analysis step S104 will be described in detail as follows.

[0174] The incremental analysis step is responsible for constructing a SEG for the affected program part and restoring the complete program behavior of this part in combination with the function summaries saved in the full analysis phase. Subsequently, the restored program behavior is encoded into a form acceptable to the SMT solver and solved.

[0175] Since the principle of constructing a SEG for the affected part is basically the same as that of the full analysis, the present invention will be described from two aspects: the behavior restoration mechanism based on function summaries, and the behavior encoding and SMT solving process.

[0176] S1041, Behavior restoration based on function summaries.

[0177] Algorithm 4 describes how to restore the path condition using the SEG. As mentioned above, the path condition consists of data dependence conditions and control dependence conditions, and the two jointly describe the sufficient conditions for the execution path.

[0178] The algorithm gradually collects data dependence conditions and control dependence conditions along the path (lines 3 - 6). The combination of the two forms the overall condition that describes the path propagation constraints (line 7). Among them, obtaining the data dependence conditions is relatively simple. Just by judging whether two adjacent variables are equal, the transfer relationship of data between nodes can be expressed (line 11). In contrast, calculating the control dependence conditions is more complex. In addition to judging whether the Boolean variables associated with the control dependence edges are consistent with the path labels (line 17), it is also necessary to recursively obtain the dependence information of these Boolean variables themselves to accurately characterize the control dependence relationship (line 18). In addition, to ensure that the path reaches the path start node v1 from the program entry, it is also necessary to further collect the control and data dependence information of v1 (line 8). Finally, the union of all dependence conditions constitutes a sufficient condition for the program to pass through the path from the entry and execute completely.

[0179] S1042, Behavior Encoding and SMT Solving.

[0180] The present invention will adopt an SMT solver as the core tool for constraint solving to support static analysis tasks such as path reachability analysis, program property verification, and potential error detection. By introducing various theoretical reasoning capabilities on the basis of the classical Boolean satisfiability problem (SAT), the SMT solver can efficiently handle complex program state constraint problems. Compared with traditional SAT solvers that can only handle Boolean logic formulas, the SMT solver further expands the supported data types and operation semantics, including theories such as integer and real arithmetic, arrays, bit vectors, strings, and uninterpreted functions, thus providing more powerful expression and solving capabilities for symbolic execution and path condition determination in static analysis.

[0181]

[0182] During the static analysis process, symbolic execution usually generates a large number of logical constraints that describe program states and data dependencies. The SMT solver can efficiently determine whether these constraints are satisfiable, thereby determining whether a program path is reachable, whether a variable assignment meets specific conditions, and whether there are potential defects in the program. Its powerful constraint solving ability provides a solid foundation for tasks such as static path analysis, error detection, and program verification.

[0183] Most current mainstream SMT solvers are implemented based on the DPLL(T) framework, which tightly combines the search in the Boolean layer with the reasoning in the theory layer. Specifically, the SMT solver first performs a Boolean abstraction on the path condition, and the SAT solver searches for feasible assignments at the Boolean level. Subsequently, the theory solver checks the consistency of these Boolean assignments under the corresponding theory. If a conflict is found, it generates a conflict clause and feeds it back to the SAT layer to guide the search to avoid the conflict area, thereby achieving efficient search space pruning and solving acceleration. This solving mechanism significantly improves the performance and accuracy of complex constraint solving in the static analysis process, making it possible to analyze the path reachability of large-scale programs.

[0184] The present invention selects Z3 as the SMT solver to achieve efficient solving of path conditions and related constraints. Z3 is based on a high-performance DPLL(T) framework, supports reasoning with multiple theory combinations, and has incremental solving and interactive query functions. Compared with other solvers, Z3 provides complete language interfaces such as Python, C++, and Java, which are convenient for integration into static analysis tools.

[0185] Due to its flexibility and efficiency, Z3 is widely used in tasks such as symbolic execution, model checking, and formal verification.

[0186]

[0187] Algorithm 5 describes the defect detection algorithm based on SMT solving and summary reuse. The algorithm first constructs a static execution graph (SEG) of the program (line 2), and then determines the source point set and sink point set according to the source-sink pattern configuration (line 3). Next, starting from each source point, it traverses its possible paths and executes the path search process (line 5). During the path search process, when a function call point is encountered, the algorithm obtains the relevant outputs of the current parameters, connects the local path with the paths in the called function, and recursively continues the path search inside the called function (lines 12-17). Finally, the algorithm comprehensively determines whether there are potential defects in the path by verifying whether the path conforms to the extended Dyck-CFL language rules and whether the path condition is satisfied (lines 18-22).

[0188] The following will evaluate the incremental defect detection method based on function summaries proposed by the present invention through empirical experiments. The experiments measure the effectiveness of the method from three aspects: function summary reuse rate, time speedup ratio, and memory savings ratio.

[0189] (1) Experimental settings

[0190] (1.1) Test project set and experimental environment

[0191] Eight test projects (with code sizes ranging from 70KLOC to 891KLOC) were selected from GitHub. They are widely used in the actual production environment, and the code in their GitHub repositories is updated frequently. For each project, 10 historical commit versions were selected as the experimental objects: among them, the first commit used full analysis, and the remaining 9 commits were evaluated using incremental analysis based on the function summaries produced in the previous version.

[0192] All experiments were completed in the same server environment. The operating system of the experimental server is Ubuntu 18.04.6 LTS, equipped with 2 Intel Xeon Gold 6230R processors (a total of 52 cores and 104 threads), with a main frequency of 2.10 GHz, 503GB of memory, a storage space of 3.5TB, and a local mount of 95TB of network storage. Since this paper implements an incremental analysis function based on Pinpoint, the execution results of the native Pinpoint in the full analysis part of the experiment are used as the comparison benchmark. Since this paper implements an incremental analysis function based on Pinpoint, the execution results of the native Pinpoint in the full analysis part of the experiment are used as the comparison benchmark. It should be noted that most of the existing other incremental analysis tools are built on different analysis frameworks and usually do not have path sensitivity in terms of analysis accuracy, making it difficult to reach the analysis accuracy level of this paper, so they are not included in the direct comparison.

[0193] (1.2) Experimental Results

[0194] The following mainly presents the experimental results of three evaluation objectives and demonstrates the effectiveness of the method in this paper through the analysis of experimental data.

[0195] (1.2.1) Function Summary Reuse Rate

[0196] The goal is to evaluate the actual proportion of existing function summaries that can be reused in the incremental analysis scenario. Table 1 counts the function sizes of each project, the number of functions affected by version changes, and the proportion of functions re-analyzed relative to the total number of project functions in the incremental analysis scenario. Specifically:

[0197] Project: Lists the names of all tested projects (openldap, openssl, etc.);

[0198] Total number of functions: Reflects the number of functions included in the overall code size of the project;

[0199] Direct / Transitive Function Count: Represents the number of functions directly affected by changes between selected versions and the number of functions affected transitively through function call relationships. For example, "2 / 7" in openldap means that 2 functions were directly modified in the code change, and another 7 functions were indirectly affected due to the call dependency chain;

[0200] Percentage of Reusable Functions (%): The percentage of functions that can reuse function summaries without being affected by the change in the total number of functions in the project.

[0201] As can be seen from Table 1, in the incremental analysis of different projects, the scope of influence brought about by each version change varies significantly. Among them:

[0202] The proportion of modified functions in openldap and openssl is extremely small, only about 0.1%, indicating that their incremental update scope is usually relatively concentrated;

[0203] The average number of functions affected by each commit in tmux and hwloc is relatively larger. The proportion of re-analyzed functions in the total number of functions reaches 20%-30%, indicating that they have a relatively high dependency propagation in large-scale modification scenarios;

[0204] tinycc, sqlite, and redis are in between the two. The proportion of re-analyzed functions ranges from 5% to 40%. It should be noted that the ratio of "directly modified functions" to "transitively affected functions" can further reflect the complexity of the function call relationships within the project. In Table 1, the direct / transitive influence function ratio of sqlite is as high as 1:70, and in a certain change of tinycc, there is even an extreme case of 4:283. These data indicate that some underlying functions undertake key logic in the project, and once they are modified, it will indirectly affect a large number of upper-layer or same-layer called functions.

[0205] Table 1 Function Summary Reusability Rates for Incremental Analysis of Different Projects

[0206]

[0207] More notably, when using the "propagation - calculation" - style analysis framework under the new version semantics to calculate the reverse - dependency closure, the above - mentioned large - scale propagation phenomenon often leads to unstable response efficiency: for code modifications of similar scale, openldap can complete the reconstruction of the symbolic expression graph (SEG) of affected functions within 20 s due to the small number of transfer - impact functions, while sqlite may take up to 20 min because of the huge scale of indirectly - dependent functions. In contrast, the incremental analysis method based on the old - version dependency - relation summary can complete the calculation of the reverse - dependency closure within seconds, enabling developers to evaluate the time overhead of incremental analysis earlier, fully demonstrating the value of the dependency - relation summary design in terms of practicality and efficiency.

[0208] In summary, the incremental analysis at the function granularity has good adaptability in dealing with code modifications of different scales: for most projects, on average, only about 5% of the functions need to be recalculated or have their analysis updated in the new version, significantly saving the time overhead required for incremental analysis. This result further proves the practicality and efficiency of the function - summary technology in a real - project environment.

[0209] (1.2.2) Time - improvement efficiency

[0210] It aims to measure the time - acceleration effect of incremental analysis compared to full - scale analysis. As shown in Table 2, the time comparison between full - scale analysis and incremental analysis was carried out on 8 test projects, recording the total duration of full - scale analysis, the average duration of incremental analysis, and the resulting speed - up ratio for each project. Among them, since there are very few directly - modified functions in openssl and its transfer impact on other functions is small, the speed - up of incremental analysis compared to full - scale analysis reaches 33 times; the speed - up ratio of openldap also exceeds 17 times. For projects with larger modification scales such as tmux and hwloc, incremental analysis can still provide a speed - up of 2 - 3 times. Generally speaking, the incremental - analysis method at the function granularity shows good time - improvement effects in most scenarios, significantly reducing the workload of repeated calculations on the new version.

[0211] Table 2 Average results and speed - up ratios of incremental analysis for different projects

[0212]

[0213] To further explore the stability of incremental analysis in most cases, the present invention selected all 12,713 functions in the openssl project, deeply analyzed the inter - function dependency relationships, and constructed a reverse - dependency - closure graph.

[0214] Figure 7shows the cumulative distribution function (CDF) of the sizes of the reverse dependency closures of various functions. The horizontal axis represents the closure size (in logarithmic scale), and the vertical axis represents the cumulative distribution ratio. This figure reveals the distribution of the degree of function dependence throughout the project. From Figure 7 , it can be seen that most functions have a small closure size, showing an obvious right-skewed distribution. Statistical data shows that the reverse dependency closure size of 95% of the functions does not exceed 53, and the average closure size is 33.30, indicating that the scope of change impact of the vast majority of functions is relatively limited. On the other hand, the maximum closure size reaches 5490, indicating that there are a small number of core functions with high influence, and their changes may have a cascading impact on a large number of other functions. This result verifies that incremental analysis can achieve significant efficiency improvement by reusing existing analysis results in most cases.

[0215] (1.2.3) Space saving ratio

[0216] is used to evaluate the savings in storage overhead of incremental analysis based on function summaries. Table 3 shows the summary space occupied and the saving ratio of different projects in the full analysis and incremental analysis modes. It can be seen from the table that for openldap and openssl, the volumes of their full analysis summaries are as high as 561MB and 4.3GB respectively, while in the incremental analysis mode, only 10.41MB and 8.26MB are required, with a saving ratio of over 98%. This indicates that for projects with relatively few direct function modifications and relatively sparse dependency relationships, summary reuse at the function granularity can greatly reduce storage and management costs. In contrast, the space saving ratios of tmux and hwloc are relatively low, still in the range of 50% - 60%. This is mainly because their incremental updates involve relatively large changes and mostly involve core underlying function calls, resulting in the need to retain more intermediate data structures in the incremental analysis mode. However, generally speaking, even for projects with a large code size (such as sqlite and redis ), incremental summaries at the function granularity can still bring significant space compression effects (above 80% - 90%), highlighting the feasibility and efficiency of incremental analysis in actual engineering scenarios.

[0217] Table 3 Experimental results of full analysis for different projects

[0218]

[0219] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that the specific embodiments described in this patent are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope protected by the claims of the present invention.

Claims

1. An incremental defect detection method based on function summaries, characterized in that Including: Function summary design steps, a path condition storage method based on the Symbolic Expression Graph (SEG), and a side-effect summary technology based on virtual nodes to construct a complete function call graph and save it as a function summary; Code change detection steps, performing program difference identification at the Intermediate Representation (IR) level, adopting a structured comparison method, analyzing the syntax structure and semantic information of the program to achieve difference detection, identifying and aligning syntax units in the program, and understanding the control flow and data flow dependencies between syntax units to identify code changes in program versions; the syntax units include functions, basic blocks, and instructions; Incremental impact analysis steps, using cached direct call relationships and based on the call graph of the old version program, identifying the set of affected functions that need to be re-analyzed due to code changes; Incremental analysis steps, constructing a Symbolic Expression Graph based on the affected programs in the set of functions that need to be re-analyzed, combining the saved function summary, restoring path conditions, and restoring the complete program of the affected program; encoding the behavior of the restored complete program into a form acceptable to the SMT solver and solving it to detect potential defects in the program.

2. The incremental defect detection method based on function summary according to claim 1, wherein The path condition storage method based on the Symbolic Expression Graph (SEG) specifically includes: Using the Symbolic Expression Graph (SEG) to implicitly store path conditions, where the SEG uses a linear number of nodes and edges to encode all possible path conditions in the program.

3. The incremental defect detection method based on function summary according to claim 1, characterized in that The side-effect summary technology based on virtual nodes specifically includes: Based on the virtual node method, modeling the side effects of functions; for global variables or shared states that may be modified, introducing virtual input nodes and virtual output nodes in the SEG, establishing edges between the virtual input nodes and virtual output nodes to depict the state change process of global variables inside the function, connecting the value of the global variable to the virtual input node, and passing its latest state through the virtual output node when the function returns, thus completely modeling the side-effect behavior of the function.

4. The incremental defect detection method based on function summary according to claim 1, wherein In the code change detection steps, during the IR structured comparison process, a tentative mapping table mechanism is introduced to temporarily record variable pairs whose equivalence cannot be confirmed immediately; when the control flow or data dependency has not been fully resolved and it is impossible to immediately determine whether two variables match, they will be included in the tentative mapping table as an assumption; as the subsequent analysis progresses, the assumption is gradually verified, and if the verification is successful, the mapping relationship is converted to a formal confirmation; if the verification fails, it is marked as a difference.

5. The incremental defect detection method based on function summary according to claim 1, wherein, In the incremental impact analysis steps, identifying the set of functions that need to be re-analyzed due to code changes specifically includes: Starting from the set of modified functions ModifiedFunc, initialize the set of affected functions AffectedFunc and the worklist, and propagate the impact layer by layer upward along the reverse call map CallerMap; through a while loop, in each round, pop a function f from the worklist and traverse all callers Caller that call f; if a certain caller is not yet included in AffectedFunc, it means that it is indirectly affected by the modified function, add it to the result set and add it to the worklist to continue the propagation; this process is repeatedly executed until the worklist is empty, indicating that all affected paths have been fully traversed; finally, return the complete set of affected functions AffectedFunc, that is, the set of affected functions that need to be re-analyzed.

6. The incremental defect detection method based on function summary according to claim 1, wherein In the incremental analysis step, restore the path conditions, specifically including: Gradually collect data dependence conditions and control dependence conditions along the path. The combination of the two forms an overall condition that describes the path propagation constraint; among them, for the data dependence condition, it is only necessary to judge whether two adjacent variables are equal to express the transfer relationship of data between nodes; for the control dependence condition, in addition to judging whether the boolean variable associated with the control dependence edge is consistent with the path label, it is also necessary to recursively obtain the dependence information of these boolean variables themselves to accurately describe the control dependence relationship; to ensure that the path reaches the path start node v1 from the program entry, it is necessary to further collect the control and data dependence information of v1; finally, the union of all dependence conditions constitutes a sufficient condition for the program to pass through the path and execute completely from the entry.

7. The incremental defect detection method based on function summary according to claim 1, wherein In the incremental analysis step, the method for detecting defects based on SMT solving and function summaries specifically includes: Construct a static execution graph of the program, and then determine the set of source points and the set of sink points according to the source-sink mode configuration; next, starting from each source point, traverse its possible paths and execute the path search process; during the path search process, when encountering a function call point, obtain the relevant output of the current parameter, connect the local path with the path in the called function, and recursively continue the path search inside the called function; finally, comprehensively judge whether there are potential defects in the path by verifying whether the path conforms to the extended Dyck-CFL language rules and whether the path conditions are satisfied.

8. An incremental defect detection system based on function summaries, characterized in that, Including: A function summary design module, which is used to construct a complete function call graph based on the path condition storage method of the symbolic expression graph SEG and the side effect summary technology based on virtual nodes, and save it as a function summary; A code change detection module, which is used to identify program differences at the intermediate representation IR level, adopt a structured comparison method, analyze the syntax structure and semantic information of the program to achieve difference detection, identify and align the syntax units in the program, and understand the control flow and data flow dependencies between the syntax units to achieve the identification of the changed code of the program version; the syntax units include functions, basic blocks, and instructions; An incremental impact analysis module, which is used to utilize the cached direct call relationship, based on the call graph of the old version program, to identify the set of all affected functions that need to be re-analyzed due to code changes; An incremental analysis module, which is used to construct a symbolic expression graph based on the affected programs in the set of functions to be re-analyzed, restore the path conditions in combination with the saved function summaries, and restore the complete programs of the affected programs; encode the behaviors of the restored complete programs into a form acceptable to the SMT solver, and perform solving to detect potential defects in the programs.

Citation Information

Cited By

  • Error positioning method and device based on P4 program

    CN120560988A

  • Resource decoupling-oriented domain name system authority engine automatic verification method

    CN120950418A

  • Redundant symbol searching method and device, terminal and medium

    CN121501293A