Static call graph construction system and method based on large language model enhancement

By combining traditional static analysis and large language model methods, a more comprehensive and accurate static call graph is built, which solves the problem that traditional methods cannot parse complex call points, and significantly improves the soundness and accuracy of program analysis.

CN120011199APending Publication Date: 2025-05-16SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510085799.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional rules-based static analysis methods cannot parse complex call points, especially when dealing with multi-file data dependencies, existing methods are difficult to analyze effectively.

Method used

Introduce a large language model, and build a more comprehensive and accurate static call diagram through the combination of traditional call graph analysis module and large language model enhancement analysis module. The specific steps include: the traditional analysis module analyzes the call point. If it is not successful, the call point is input to the large language model for reasoning, and combining static analysis and code context search technology to summarize the analysis results.

Benefits of technology

It effectively improves the soundness of program analysis, solves the limitations of traditional methods when dealing with dynamic language features, and significantly improves the accuracy and comprehensiveness of call graph construction in JavaScript and Python languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011199A_ABST
    Figure CN120011199A_ABST
Patent Text Reader

Abstract

The invention provides a static call graph construction system and method based on large language model enhancement, and the system comprises a traditional call graph analysis module which is used for analyzing a called function according to an input call point, and inputting the corresponding call point into a large language model enhancement analysis module if the analysis is not successful; the big language model enhancement analysis module is used for receiving the calling points which are not successfully analyzed and reasoning by using a big language model to obtain a called function; and summarizing and outputting called functions analyzed by the traditional call graph analysis module and the large language model enhancement analysis module. According to the method, two methods based on data dependence and keyword retrieval are combined, so that the problem of low recall rate in a retrieval enhancement system is solved, and the illusion phenomenon in large language model analysis is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software static analysis, and in particular to a static call graph construction system and method based on large language model enhancement. Background Art

[0002] Call graph is a core concept in static analysis. It represents the calling relationship between functions in a program and is crucial for applications such as optimizing compilers, debugging tools, and security checks. Traditional call graph generation relies on rule-based static analysis methods, which parses the source code and applies predefined data flow analysis rules to determine the calling relationship between functions. However, this method cannot resolve complex call sites.

[0003] Patent document CN117992347A discloses a system and method for generating loop invariants in a program based on a large language model. Although the patent document can generate program invariants, it is difficult to handle complex data dependency situations. For example, when analyzing a certain call point, it is often necessary to extract and integrate related codes from multiple files. This requirement exceeds the processing capabilities of existing methods.

[0004] The present invention aims to solve the problem that traditional rule-based static analysis methods cannot parse complex call points. By introducing a large language model, the call points that cannot be analyzed are input into the LLM for reasoning, and static analysis and code context retrieval technology are combined to achieve a more comprehensive and accurate call graph construction. Summary of the invention

[0005] In view of the defects in the prior art, the object of the present invention is to provide a static call graph construction system and method based on large language model enhancement.

[0006] A static call graph construction system based on large language model enhancement provided by the present invention includes:

[0007] The traditional call graph analysis module is used to analyze the called function according to the input call point. If the analysis fails, the corresponding call point is input into the large language model enhanced analysis module;

[0008] A large language model enhanced analysis module is used to receive call points that failed to be analyzed and use the large language model to perform reasoning to obtain the called function;

[0009] Summarize and output the called functions analyzed by the traditional call graph analysis module and the large language model enhanced analysis module.

[0010] Preferably, the traditional call graph analysis module includes a call graph analyzer based on expert rules, which takes the source code of the analyzed project as input and searches for the called function through the data flow analysis rules defined by experts; if the called function is successfully analyzed, the analysis of the call point is completed and the called function is output.

[0011] Preferably, the large language model enhanced analysis module includes:

[0012] The large language model reasoning module attempts to analyze the function called by the call point based on understanding the program semantics, and performs analysis output through the iterative call prompt unit; the condition for the end of iteration is to analyze the function called by the call point or reach the set maximum number of iterations;

[0013] The code segmentation module is used to segment the source code of the analyzed project into code segments with preset length and semantics;

[0014] The code retrieval module is used to find the context of code elements from the segmented code blocks.

[0015] Preferably, the input of the large language model inference module includes a call point and a set of code blocks; if the provided code block contains all the information required for analysis, the function called by the call point is output; if the provided code block lacks all the information required for analysis, the code element with missing context is output.

[0016] Preferably, the segmentation principle used by the code segmentation module includes scanning the abstract syntax tree of the code from top to bottom, and segmenting at folding points that meet preset requirements; for each code file in the analyzed project, the segmentation module generates a segmentation tree; each block contains code from a specific abstract syntax tree node, and traversing the block tree will obtain all blocks of the file.

[0017] Preferably, the input of the code retrieval module includes questions and documents, the retrieval algorithm is executed, and the output includes retrieval results; the questions include code elements with missing context reported by the large language model inference module; the documents include code blocks after the code segmentation module segments the source code of the analyzed project; the retrieval algorithm includes a data dependency-based retrieval method and a keyword-based retrieval method; the retrieval results include a set of code blocks containing the context of the target code elements.

[0018] Preferably, the prompt unit is used to guide the prompts and processing of the large language model analysis, including a filtering prompt sub-unit, an analysis prompt sub-unit and a summary prompt sub-unit; each time a prompt sub-unit is executed, the large language model is called once; three prompts are executed serially, and the output of the previous execution is input into the next prompt.

[0019] Preferably, the filtering prompt subunit is used to filter out code blocks related to the call point from a given set of code blocks; the analysis prompt subunit is used to guide the large language model to analyze the function called by the call point, fill in the filtering prompt template with specific content, and then provide it to the large language model.

[0020] Preferably, the summary prompt subunit is used to convert the answer of the large language model from a natural language form into a formatted form.

[0021] A static call graph construction method based on large language model enhancement provided by the present invention includes:

[0022] Step S1: receiving the source code of the project to be analyzed;

[0023] Step S2: Analyze the call points in the source code using a traditional call graph analysis module;

[0024] Step S3: If the traditional call graph analysis module fails to analyze the calling function of the call point, the call point is input into the large language model enhanced analysis module;

[0025] Step S4: using the large language model to infer the call points that failed to be analyzed successfully, and obtaining the called functions;

[0026] Step S5: Summarize and output the called functions analyzed by the traditional call graph analysis module and the large language model enhanced analysis module.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] 1. The present invention solves the limitations of traditional program analysis in processing dynamic language characteristics by adopting a method of analyzing program semantics with a large language model, thereby effectively improving the robustness of program analysis. Compared with the CodeQL tool, the present invention improves the reachable node and reachable edge evaluation indicators by 30 and 26 percentage points in the JavaScript language, respectively; and improves the reachable node and reachable edge evaluation indicators by 29 and 24 percentage points in the Python language, respectively.

[0029] 2. The present invention solves the problem of low recall rate in the retrieval enhancement system by combining two methods based on data dependency and keyword retrieval, and significantly reduces the phenomenon of hallucinations in large language model analysis; compared with using only keyword-based methods, the analysis robustness of the present invention is improved by more than 10%.

[0030] 3. The present invention adopts a code block strategy of nested blocks, which effectively solves the problems of large differences in block sizes and incomplete semantics in traditional methods, thereby improving the analysis robustness by 10% to 15%.

[0031] Other beneficial effects of the present invention will be explained in the specific implementation manner through the introduction of specific technical features and technical solutions. Through the introduction of these technical features and technical solutions, those skilled in the art should be able to understand the beneficial technical effects brought about by the technical features and technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0033] Figure 1 It is a system block diagram of the present invention.

[0034] Figure 2 It is a working schematic diagram of the large language model enhanced analysis module of the present invention.

[0035] Figure 3 Schematic diagram of the subunits of the prompt unit of the present invention. DETAILED DESCRIPTION

[0036] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0037] A call graph is a directed graph of function call relationships in a program. Static analysis of call relationships requires analyzing the functions called by each call site in the program.

[0038] The method and device include two components: a traditional call graph analysis module and a large language model enhanced analysis module, such as Figure 1 shown.

[0039] The traditional call graph analysis module is a call graph analyzer based on expert rules. Its input is the source code of the analyzed project. It uses the data flow analysis rules defined by experts to find the called functions and outputs two situations based on whether the called functions are analyzed: if the called functions are analyzed, the analysis of the call point ends; if the called functions are not analyzed, the call point is input to the large language model enhanced analysis module.

[0040] The large language model enhanced analysis module includes three modules: large language model reasoning module, code retrieval module and code segmentation module. Figure 2 shown.

[0041] The large language model inference module attempts to analyze the function called by the call point by understanding the program semantics. Its input is a call point and a set of code blocks, which are then analyzed through iterative call prompt units. The final output results are two: the provided code block contains all the information required for analysis, and the function called by the call point is output; the provided code block lacks all the information required for analysis, and the code elements with missing context are output, such as missing defined variable names and missing implemented function names. The condition for the end of the iteration is to analyze the function called by the call point, or to reach the set maximum number of iterations.

[0042] By adopting a method of analyzing program semantics using a large language model, the present invention solves the limitations of traditional program analysis in dealing with dynamic language characteristics, thereby effectively improving the robustness of program analysis. Compared with the CodeQL tool, the present invention improves the reachable node and reachable edge evaluation indicators by 30 and 26 percentage points in JavaScript language, respectively; and improves the reachable node and reachable edge evaluation indicators by 29 and 24 percentage points in Python language, respectively.

[0043] The prompt unit is a set of prompts and processing units that guide the analysis of the large language model. The design idea of ​​the prompt unit is derived from the chain-of-thought. There are three prompt sub-units: filtering prompt sub-unit, analysis prompt sub-unit and summary prompt sub-unit.

[0044] Each time a prompt sub-unit is executed, the large language model is called once and the output of the large language model is processed. The three prompts are executed in series, and the output of the previous execution is used as the content input to the next prompt, such as Figure 3 shown.

[0045] The filtering hint subunit is responsible for filtering out the code blocks related to the call site from a given set of code blocks. The reason for setting up the filtering hint subunit is that the code retrieval module retrieves multiple code blocks, but only some of these code blocks are related to the analysis of the call site, so irrelevant code blocks are filtered out to reduce the difficulty of analysis and reasoning of the large language model. The input of the filtering hint subunit is a set of code blocks, and the output is a set of filtered code blocks. The filtering hint subunit includes filtering hints and filtering operations. The template of the filtering hint is: given a code block [%CODE%], please list the code block numbers that may be related to the call site [%CALL_SITE%] in the analysis statement [%STATEMENT%]. After filling in the filtering hint template with specific content, it is provided to the large language model. The large language model will report a set of code block numbers. The filtering operation will filter out a set of code blocks based on the numbers and input them into the analysis hint subunit.

[0046] By adopting a code block strategy of nested blocks, the present invention effectively solves the problems of large differences in block sizes and incomplete semantics in traditional methods, thereby improving the analysis robustness by 10% to 15%.

[0047] The analysis hint subunit is responsible for guiding the large language model to analyze the functions called by the call site. The template of the analysis hint is: given a code block [%CODE%], please analyze the functions called by the call site [%CALL_SITE%] in [%STATEMENT%] and explain the key data flow in the analysis. After filling in the filter hint template with specific content, it is provided to the large language model. The output of the large language model will be input into the summary hint subunit.

[0048] The summary prompt subunit is responsible for converting the answer of the large language model from natural language form to formatted form. The template of the summary prompt is: Please summarize the results in JSON. If the called function has been analyzed, please answer in the format {"CALLEE":["FUNC1","FUNC2"]}. Otherwise, please report the code elements with missing context {"MISS_CONTEXT":["QUERY1","QUERY2"]}. These code elements may be variables, methods, global objects, etc.

[0049] The code segmentation module is responsible for segmenting the source code of the analyzed project into code blocks of moderate length and relatively complete semantics, thereby improving the reasoning speed and quality of the large language model reasoning module and the retrieval effect of the code retrieval module. The segmentation principle is similar to the code folding principle in the integrated development environment (IDE). The abstract syntax tree of the code is scanned from top to bottom, and segmented at points suitable for folding (such as object definitions and function bodies) to maintain semantic integrity.

[0050] Specifically, for each code file in the analyzed project, the chunking module generates a chunking tree. Each chunk contains code from a specific abstract syntax tree node, and traversing the chunk tree will get all chunks of the file. The chunking tree is constructed by traversing from the root node of the abstract syntax tree of the file to identify the bracketed abstract syntax tree nodes closest to the root node, such as "class_body", "statement_block", "switch_body", and "objects". Each time such a node is encountered, a chunk is generated. After the number of code tokens contained in each chunk exceeds the preset threshold, the chunk will be further subdivided to keep all chunks of similar size. The code chunk owned by each chunk only contains itself, and the code of its child nodes will be omitted.

[0051] The code retrieval module is responsible for finding the context of code elements from the code blocks of the analyzed project. The code retrieval module is a retrieval tool that includes four key elements: questions, documents, retrieval algorithms, and retrieval results. The input of the code retrieval module is questions and documents, and the retrieval algorithm is executed, and the output is the retrieval result. Among them, the question is the code element with missing context reported by the large language model inference module; the document is the code block after the source code of the analyzed project is segmented by the code segmentation module; the retrieval algorithm includes data dependency-based retrieval methods and keyword-based retrieval methods; the retrieval result is a set of code blocks containing the context of the target code element.

[0052] The data dependency-based retrieval method retrieves related code blocks for code elements with missing context from the perspective of data dependency. The input of the retrieval method is a code element with missing context, and the output is a set of code blocks. The retrieval method is a worklist algorithm, which puts the code elements with missing context into the worklist for initialization. Then a loop is executed, and the condition for the end of the loop is that the worklist is empty. If it is not empty, a code element is popped up each time to find its definition and all references. Then, the data dependency is carefully checked inside and between functions: if the symbolic code elements e and e' are on both sides of the assignment, there is a data dependency between them; between functions, if e is an actual parameter and e' is a formal parameter, there is also a data dependency between them. Therefore, the newly identified code element is added to the worklist for further analysis. In addition, finding definitions and finding references are basic operations of static analysis, and the results can come from a traditional call graph analysis module. The present invention follows the design of the language server protocol and obtains these results from static analysis. This design is easy to integrate with analysis tools that support the language server protocol.

[0053] The keyword-based retrieval method is a traditional BM25 searcher. The input is a code element with missing context, and the output is a set of code blocks. The retrieval method searches on three domains of the code block. The three domains are the description of the code block, the name of the function that appears in the code block, and the entire code in the code block. The specific retrieval process is: first, the score of each code block on the three domains is calculated separately according to the BM25 algorithm; then, the scores of the three domains are weighted and summed with the same weight to obtain the overall score of each code block; finally, the k code blocks with the highest overall score are returned. In the experiment, the value of k is 5.

[0054] By combining the two methods based on data dependency and keyword retrieval, the present invention solves the problem of low recall rate in retrieval enhancement system and significantly reduces the phenomenon of hallucination in large language model analysis. Compared with the method based on keywords alone, the analysis robustness of the present invention is improved by more than 10%.

[0055] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.

[0056] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A static call graph construction system based on large language model enhancement, characterized in that: include: The traditional call graph analysis module is used to analyze the called function according to the input call point. If the analysis fails, the corresponding call point is input into the large language model enhanced analysis module; A large language model enhanced analysis module is used to receive call points that failed to be analyzed and use the large language model to perform reasoning to obtain the called function; Summarize and output the called functions analyzed by the traditional call graph analysis module and the large language model enhanced analysis module.

2. The static call graph construction system based on large language model enhancement according to claim 1, characterized in that: The conventional call graph analysis module includes a call graph analyzer based on expert rules, which takes the source code of the analyzed project as input and searches for called functions through data flow analysis rules defined by experts; If the called function is successfully analyzed, the analysis of the call point ends and the called function is output.

3. The static call graph construction system based on large language model enhancement according to claim 2 is characterized in that: The large language model enhanced analysis module includes: The large language model reasoning module attempts to analyze the function called by the call point based on understanding the program semantics, and performs analysis output through the iterative call prompt unit; the condition for the end of iteration is to analyze the function called by the call point or reach the set maximum number of iterations; The code segmentation module is used to segment the source code of the analyzed project into code segments with preset length and semantics; The code retrieval module is used to find the context of code elements from the segmented code blocks.

4. The static call graph construction system based on large language model enhancement according to claim 3 is characterized in that: The input of the large language model inference module includes a call point and a set of code blocks; if the provided code block contains all the information required for analysis, the function called by the call point is output; if the provided code block lacks all the information required for analysis, the code element with missing context is output.

5. The static call graph construction system based on large language model enhancement according to claim 4, characterized in that: The segmentation principle used by the code segmentation module includes scanning the abstract syntax tree of the code from top to bottom and performing segmentation at the folding points that meet the preset requirements; For each code file in the analyzed project, the chunking module generates a chunk tree; each chunk contains code from a specific abstract syntax tree node, and traversing the chunk tree will get all chunks of the file.

6. The static call graph construction system based on large language model enhancement according to claim 4, characterized in that: The input of the code retrieval module includes questions and documents, and the retrieval algorithm is executed, and the output includes retrieval results; the questions include code elements with missing context reported by the large language model inference module; the documents include code blocks after the source code of the analyzed project is segmented by the code segmentation module; the retrieval algorithm includes a data dependency-based retrieval method and a keyword-based retrieval method; the retrieval results include a set of code blocks containing the context of the target code elements.

7. The static call graph construction system based on large language model enhancement according to claim 3, characterized in that: The prompt unit is used to guide the prompt and processing of the large language model analysis, including a filtering prompt subunit, an analysis prompt subunit and a summary prompt subunit; each time a prompt subunit is executed, the large language model is called once; The three prompts are executed serially, with the output of the previous execution being fed into the next prompt.

8. The static call graph construction system based on large language model enhancement according to claim 7, characterized in that: The filtering prompt subunit is used to filter out the code blocks related to the call point from a given set of code blocks; The analysis prompt subunit is used to guide the large language model to analyze the function called by the call point, and after filling in the filter prompt template with specific content, provide it to the large language model.

9. The static call graph construction system based on large language model enhancement according to claim 7, characterized in that: The summary prompt subunit is used to convert the answer of the large language model from natural language form into formatted form.

10. A method for constructing a static call graph based on large language model enhancement, based on the static call graph construction system based on large language model enhancement according to any one of claims 1 to 9, characterized in that: include: Step S1: receiving the source code of the project to be analyzed; Step S2: Analyze the call points in the source code using a traditional call graph analysis module; Step S3: if the traditional call graph analysis module fails to analyze the calling function of the call point, the call point is input into the large language model enhanced analysis module; Step S4: using the large language model to infer the call points that failed to be analyzed successfully, and obtaining the called functions; Step S5: Summarize and output the called functions analyzed by the traditional call graph analysis module and the large language model enhanced analysis module.

Citation Information

Patent Citations

  • System and method for generating loop invariants in program based on large language model

    CN117992347A

Cited By

  • Key function call feature acquisition method, knowledge retrieval method, equipment and medium

    CN120257218A