System and method for static code analysis using large language model

By simulating pseudocode execution and verifying results using large language models, the challenges of traditional static analysis tools in dealing with complex program logic and high computational costs are solved, achieving more efficient, flexible and accurate static code analysis.

CN120066573AInactive Publication Date: 2025-05-30XIDIAN UNIV HANGZHOU RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510525777.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional static analysis tools have challenges in dealing with complex program logic and high computational costs, and the expertise required to set up and maintain limits their applicability and flexibility.

Method used

Large language models (LLMs) are used for static code analysis, and multiple subunits are coordinated through proxy modules, including pseudo-code execution units and execution specification verification units, to simulate pseudo-code execution and verify results, ensuring the accuracy and security of the analysis.

Benefits of technology

It improves the flexibility and adaptability of static analysis, reduces complex configuration and maintenance requirements, significantly improves the accuracy and reliability of analysis results, and helps developers optimize code more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066573A_ABST
    Figure CN120066573A_ABST
Patent Text Reader

Abstract

The invention discloses a system and a method for performing static code analysis by using a large language model. The system comprises a task information module, a source code module, a pseudo code module, an agent module and a result display module. Static analysis task information and pseudo codes defined by a user and code input of the source code module are received through the agent module. The pseudo-code module is used for storing user-defined pseudo-codes, and the pseudo-codes specify analysis tasks to be executed. And in a set execution specification, simulating and executing the pseudo code by using a large language model, and performing static analysis. And the system simulates and executes the pseudo code through the agent module, verifies an execution result, and then outputs a static analysis result through the result display module. In addition, the system can process intermediate results of pseudo-code execution and provide code optimization suggestions and error repair schemes according to the execution results. By simulating execution and verification of pseudo codes, the system can effectively improve code quality, reduce programming errors and improve development efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software engineering, and specifically relates to a system and method for static code analysis using large language models. Background Art

[0002] Static analysis is a technique for analyzing computer software without executing the program. Traditional static analysis tools such as SonarQube, Fortify, etc., although powerful, mainly analyze known error patterns and defined error rules. With the increasing complexity in the field of software development, traditional tools are unable to handle new code patterns and language features. In addition, the setup and maintenance of these tools often require in-depth expertise, which limits their applicability and flexibility.

[0003] Large language models (LLMs), such as OpenAI's GPT series, have shown strong capabilities in processing natural language in various tasks. In particular, in the field of software engineering, these models have been used to generate code and conduct code reviews. However, the application of LLMs to static code analysis has not been widely explored, especially in performing pseudocode analysis and verification.

[0004] In the field of software analysis, due to the many challenges faced by traditional static analysis techniques, such as complex program logic and high computational costs. To address these issues, static analysis methods combining large language models (LLMs) have begun to emerge. Summary of the Invention

[0005] To solve the deficiencies of the prior art and achieve the purpose of performing static analysis of pseudocode by a large language model and ensuring the accuracy of the analysis through validating these executions, the present invention adopts the following technical solutions:

[0006] A system for static code analysis using a large language model is an innovative and efficient software analysis solution that encompasses multiple functional modules, with the most core being the proxy module. This proxy module plays a crucial coordinating and management role in the entire system. It works closely with multiple subunits, including a source code retrieval unit, a pseudocode execution unit, and an execution specification verification unit.

[0007] Furthermore, the system will obtain pseudocode, source code, and task information from different sources respectively. Among them, pseudocode is a special form of code representation that provides a clear analysis framework for subsequent static analysis tasks; source code is the program code that actually needs to be analyzed, while task information clarifies the specific goals and expected effects of this static analysis. The pseudocode execution unit undertakes the key execution task. It utilizes the capabilities of a powerful large language model to parse the pseudocode. During the parsing process, it will deeply analyze the structure, logic, and various elements of the pseudocode, and execute tasks in a static analysis manner. Specifically, it converts the pseudocode into queries executable by the large language model, and through the internal mechanisms of the large language model, such as complex language understanding and prediction functions, it simulates the execution path of the pseudocode. It should be noted that this simulation is not a real code execution, but relies on the powerful prediction ability of the large language model to infer potential code behaviors and results, so as to provide us with in-depth understanding and analysis of the code without actually running the code.

[0008] While performing the static analysis task, the source code retrieval unit also plays an important role. It will automatically retrieve the source code according to the current analysis task and the provided task information. By retrieving relevant source code fragments, it provides more comprehensive and accurate information support for the pseudocode execution unit, helping the large language model better understand the context and semantics of the code, and thus improving the accuracy of the analysis. The execution specification verification unit is responsible for strictly verifying the entire execution process. It carefully checks the execution results based on a series of pre-set execution specifications to ensure the correctness and security of the operations.

[0009] Furthermore, in the execution specification verification unit, it will compare the analysis results with the expected results. If it is found that the results do not match the expectations, the system will use its powerful analysis ability to mark potential problem areas. These problem areas may involve multiple aspects such as code logic, performance issues, and security vulnerabilities. Once a problem is discovered, the system is highly flexible. It can choose to re-execute the analysis task or optimize the entire analysis process by adjusting the parameters of the large language model, such as adjusting the temperature parameter, maximum length limit, etc., to achieve better analysis results.

[0010] Furthermore, the pseudocode used here is highly targeted and specialized. It is designed for specific static analysis tasks and aims to simulate some complex static analysis processes through large language models, such as backward taint analysis or other static analysis processes. Backward taint analysis can help us trace the flow of data in a program and identify potential security risks or sources of errors. In the execution specification verification unit, more detailed error classification and error attribution operations will be performed using the pseudocode input. Error classification plays an important role in software systems in identifying and differentiating different types of errors. For example, the result of executing the pseudocode is used to determine the responsible function for software crashes, that is, to find the key function or code snippet that causes the software to crash. Error attribution is a process of matching the static analysis results with the actual software repair records. Through this matching, the accuracy of the analysis results can be verified to ensure that the analysis results are consistent with the actual situation, providing a reliable basis for subsequent code repair and optimization.

[0011] To achieve the integrity and systematicness of the system functions, the system also includes a task information module, a pseudocode module, a source code module, and a result output module, which are respectively connected to the proxy module. The task information module, as one of the information input sources of the system, is specifically used to obtain the specific requirements of the analysis task. It collects information such as the analysis target, expected performance indicators, and security requirements from users or other system interfaces to ensure that the entire analysis task has a clear direction and goal. The pseudocode module is responsible for obtaining the pseudocode used to define the specified static analysis task. It can extract the corresponding pseudocode from local storage or other code libraries, providing a clear execution blueprint for the pseudocode execution unit. The main function of the source code module is to obtain the actual source code to ensure the integrity and accuracy of the analysis object.

[0012] The result output module is an important output port of the entire system. It is responsible for outputting the final analysis report and intermediate results. To make the output clearer and more intuitive, the result output module uses a display unit to show users rich information, including the specific location of the problem code, enabling users to quickly locate the code snippets that may have problems; providing possible cause analysis to help users understand the root causes of the problems; and also giving repair suggestions to provide specific guidance for users to optimize the code. Through such comprehensive and detailed output, users can clearly see the code behavior and potential problems, thereby optimizing and improving the code to enhance the quality and performance of the code.

[0013] When the entire analysis process is completed, the system will present the final analysis report and intermediate results to the user through the result display module, providing clear feedback and profound insights to the user. This highly automated and intelligent process not only greatly improves the processing efficiency, making the originally complex and time-consuming static code analysis more efficient and rapid, but also provides a convenient interaction experience for the user, enabling the user to easily grasp the status and problems of the code, providing strong support for code development and maintenance.

[0014] A method for static code analysis using a large language model includes the following steps:

[0015] Step 1: Collect basic static analysis data such as crash reports, performance data, and code change history to provide a basis for subsequent tasks;

[0016] Preprocess the collected data and convert it into a format that can be understood and processed by the large language model, such as a unified crash report format and normalized performance data;

[0017] Step 2: Write or select concise and clear pseudocode that can express the task logic process according to the static analysis goal, such as finding vulnerabilities or optimizing performance;

[0018] Step 3: Input the pseudocode, preprocessed data, and task information into the large language model, and configure its temperature parameter, maximum length, etc.;

[0019] Step 4: The large language model starts to parse the pseudocode, identify variables, operators, control structures, and understand the logical relationships; the large language model simulates the execution of the static analysis task based on the pseudocode logic and parsing results, which involves code traversal, variable calculation tracking, function call analysis, etc.; the source code retrieval unit automatically retrieves relevant source code, provides it as auxiliary information to the large language model, and monitors the execution process and records intermediate results; verify the intermediate results and final results of the execution according to predefined specification standards, check for violations and automatically classify the types of violations; organize and summarize the verified results, remove redundant information, and output the static analysis results in a clear format, such as generating a report, etc.

[0020] Furthermore, in the above Step 1, the process of preprocessing the collected data needs to consider the diversity and complexity of the data. For crash reports from different sources, they may contain different formats and content structures, and various text processing techniques need to be used to unify them into a text format that is easy for the large language model to understand, ensuring that the key information is not lost. For performance data, not only simple normalization processing is required, but also different normalization methods need to be adopted according to different performance metrics (such as response time, throughput, etc.) to better meet the input requirements of the large language model and provide a reliable data basis for subsequent analysis tasks.

[0021] Furthermore, in step 3, the configuration of the large language model is not limited to the temperature parameter and the maximum length limit. It may also involve adjusting other hyperparameters of the language model, such as selecting different optimizers, adjusting the learning rate, setting different training epochs, etc. The adjustment of these parameters will affect the performance of the language model and the quality of the analysis results. In addition, according to different task information and the complexity of the pseudocode, it may be necessary to fine-tune the architecture of the model to better adapt to specific static analysis tasks, so that the model can more accurately parse and execute the pseudocode.

[0022] Furthermore, in step 4, when the large language model parses the pseudocode, in addition to identifying variables, operators, and control structures, it will also parse the comments and special markers in the pseudocode. Comments may contain important hint information, which helps to understand the logic and intention of the code, while special markers may indicate some special processing logic or special cases. This information is very important for accurately simulating the execution of the pseudocode. At the same time, for complex control structures, such as nested loops and multi-layer conditional judgments, the language model needs to carefully handle their logical levels to avoid incorrect parsing and execution.

[0023] Furthermore, in step 4, when the large language model simulates the execution of static analysis tasks, the traversal of the code needs to consider the structure and semantics of the code. It will perform a depth-first or breadth-first traversal of function calls and variable uses according to the dependency relationship and call order of the code to ensure a comprehensive analysis of the code. In terms of calculating and tracking variable values, not only the current value of the variable needs to be considered, but also its changes in different code blocks and the dependency relationship between variables need to be considered to accurately infer the execution result of the code. For the analysis of function calls, the parameter passing, return value, and possible side effects of the function need to be considered to evaluate its impact on the overall program logic.

[0024] Furthermore, in step 4, the automatic retrieval operation of the source code retrieval unit needs to be based on efficient retrieval algorithms and index structures. It will quickly locate relevant code fragments in the source code according to the task information and key elements in the pseudocode. During the retrieval process, not only text matching of the code needs to be considered, but also semantic matching of the code needs to be considered to avoid missing important information. At the same time, the monitoring of the execution process needs to record the timestamp and intermediate results of each operation in detail for subsequent performance analysis and result verification, and the location and time of the problem occurrence can be traced back when an exception occurs.

[0025] Furthermore, in step 4, when performing specification verification on the intermediate results and final results, in addition to checking for violations, the accuracy and integrity of the results, it is also necessary to consider the mutual influence between different types of violations. For complex violation situations, which may involve combinations of multiple violation types, more in-depth analysis is required. In addition, for different analysis tasks, the predefined specification standards may also need to be adjusted dynamically to ensure the effectiveness and accuracy of the verification process.

[0026] Furthermore, in step 4, when sorting and summarizing the verified results, not only unnecessary information should be removed, but also important information should be highlighted and explained. For example, for the potential problems found, their severity and scope of influence should be elaborated in detail to help users better understand the results. When outputting the final static analysis results, according to different user requirements, the report can be generated in different formats, such as HTML format for easy viewing by users, or JSON format for easy invocation and processing by other systems, while ensuring the integrity and consistency of the output information. A series of predefined test cases and verification rules are used to check the output of the language model, including verifying the rationality and consistency of the output results and comparing them with the expected analysis standards.

[0027] The advantages and beneficial effects of the present invention are as follows:

[0028] The system and method for static code analysis using a large language model according to the present invention perform static analysis by using the large language model to execute pseudocode, which not only reduces the complex configuration and maintenance requirements of traditional static analysis tools, but also improves the flexibility and adaptability of static analysis. Through the mechanism of verifying the execution results, the accuracy and reliability of the analysis results are greatly improved, helping developers optimize the code more effectively and reduce potential errors. This systematic and intelligent analysis method can save developers' time and effort and accelerate the development process of high-quality code. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram of the system structure of an embodiment of the present invention.

[0030] Figure 2 is a flowchart of the method of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following details the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention and are not intended to limit the present invention.

[0032] The E&V (Execution and Verification) method is an innovative static analysis method that utilizes the capabilities of LLMs to perform static analysis in a structured and verifiable manner through pseudocode and execution specifications. It involves multiple components, including an Agent, a pseudocode execution component, a source code retrieval component, and an execution specification verification component. Through the collaboration of these components, it enables efficient analysis of the source code of the target program (such as the Linux kernel source code), aiming to improve the accuracy and efficiency of static analysis while reducing the impact of different types of hallucinations (such as context conflict hallucinations, fact conflict hallucinations, and input conflict hallucinations) on the analysis results. In this context, the implementation steps of the E&V method include preparing input data (task information, pseudocode, and target program source code), using LLMs to execute the pseudocode, verifying and iterating the execution process, and finally outputting the analysis results. The entire process controls the workflow through the Agent, dynamically adjusts the components and analysis tasks to adapt to different static analysis requirements, such as crash classification, vulnerability detection, and performance optimization.

[0033] As Figure 1 shown, the present invention proposes a system for static code analysis using large language models, including a task information module, a pseudocode module, a source code module, an Agent module, and a result display module. The Agent module includes a source code retrieval unit, a pseudocode execution unit, an execution specification verification unit, and other tools. By receiving task information, pseudocode, and source code, it uses the core processing module - the Agent - to perform detailed data processing and analysis. The system first obtains the specific requirements of the analysis task from the task information module, and then receives the user-defined pseudocode and the actual source code through the pseudocode module and the source code module respectively. The Agent module is responsible for executing the core functions. It automatically retrieves the source code through the source code retrieval unit, simulates the execution of the pseudocode through the pseudocode execution unit, and verifies the specifications of the execution process through the execution specification verification unit to ensure the correctness and security of the operations. In addition, the Agent can also call other auxiliary tools to improve the analysis efficiency and the accuracy of the results. After the analysis is completed, the system will output the final analysis report and intermediate results through the result display module, providing clear feedback and insights to users to help them understand the code behavior and potential problems, thereby optimizing and improving the code quality. This highly automated and intelligent process greatly improves the processing efficiency and the convenience of user interaction.

[0034] The system can also include a pseudocode optimization module and an execution result feedback module, which, in combination with the verification module and the result display module, respectively perform optimization of the pseudocode, feedback of the execution results, and further dynamic adjustment and optimization.

[0035] Specifically, the pseudocode execution unit converts the pseudocode into queries executable by the large language model, and simulates the pseudocode execution path through the internal mechanism of the large language model, and infers potential code behaviors and results through the prediction function.

[0036] Specifically, the pseudocode specifies the target of the analysis, the expected type of analysis, and the execution parameters; in the execution specification verification unit, if the result does not match the expectation, the system will mark the potential problem area and can choose to re-execute or adjust the model parameters for optimization.

[0037] Specifically, the pseudocode is specific static analysis pseudocode designed to simulate backward taint analysis or other static analysis processes through a large language model; in the execution specification verification unit, the pseudocode input is used for error classification and error attribution to identify and classify errors; error classification includes using the result of pseudocode execution to determine the responsible function for software crashes; error attribution is to match the static analysis results with the actual software patch records.

[0038] As Figure 2 shown, the method for static code analysis using a large language model mainly includes the following steps:

[0039] Step 101: Receive the pseudocode, source code, and task information for specifying the static analysis task. In this step, the system designs a user interface or API that allows the user to input or upload the pseudocode defining the static analysis task. The pseudocode should clearly specify the target of the analysis, the expected type of analysis (such as memory leak check, race condition detection, etc.), and any specific execution parameters. In addition, the system can provide templates and guidelines to help the user correctly write the pseudocode, ensure that it conforms to the parsing standard, and facilitate the correct execution of subsequent steps. The pseudocode should be designed to be concise and easy for the large language model to understand and execute.

[0040] Step 102: Based on the received pseudocode, use the large language model to parse the pseudocode to perform the static analysis task and simulate the execution of the analysis task described by the pseudocode. After receiving the pseudocode, the system provides it as input to a pre-trained large language model (such as GPT-3 or a domain-specific model specially trained). The model will parse the pseudocode and simulate the execution of the described analysis task. The simulation execution includes converting the pseudocode into queries executable by the model and simulating the code execution path through the internal mechanism of the language model. This simulation does not involve actual code execution, but infers potential code behaviors and results through the prediction function of the model. To improve the accuracy and depth of the analysis, the model can combine static analysis algorithms and known programming rule libraries to enhance its prediction ability.

[0041] Further, a specific static analysis pseudocode is used, which is designed to simulate backward taint analysis or other static analysis processes through large language models.

[0042] Step 103: Based on the automatic retrieval of source code performed by the source code retrieval unit and the task information, perform specification verification of the execution process to ensure the accuracy of the analysis results. The key to this step is to verify the results of the simulated execution. The system will use a series of predefined test cases and verification rules to check the output of the language model. This includes verifying the rationality and consistency of the output results and comparing them with the expected analysis criteria. The verification process may include automated unit tests, integration tests, and manual reviews. If the results do not match the expectations, the system will mark the potential problem areas and may choose to re-execute or adjust the model parameters for optimization. Verifying the results of the simulated execution involves using specialized verification algorithms to compare the execution results with the expected behavior to ensure the accuracy of the analysis.

[0043] Further, perform a normalization check on the output of the executed pseudocode to verify whether the output is consistent with the execution specifications.

[0044] Further, the verification step includes performing automatic classification of specific violation types, such as unrecognized execution, incomplete execution, and inconsistent execution.

[0045] Further, use the pseudocode input for error classification and error attribution to identify and classify errors in the software system. Error classification includes using the results of the pseudocode execution to determine the responsible function for software crashes; error attribution further includes matching the static analysis results with the actual software patch records to verify the accuracy of the analysis.

[0046] Step 104: Output the results of the static analysis. Once the accuracy of the results is verified, the system will organize and format the analysis results for reporting to the user. The output may include error reports, lists of performance issues, code improvement suggestions, etc. The system can also provide an interactive interface that allows users to view the detailed analysis results, including the specific location of the problem code, possible cause analysis, and repair suggestions. In addition, to facilitate teamwork and the review process, the system can also support outputting the results to files in various formats or directly integrating them into the existing development environment.

[0047] In the prior art, there is no solution that fully utilizes large language models (LLMs) for static code analysis. The present invention performs static analysis by simulating the execution of pseudocode and has the following advantages compared with existing solutions of traditional static analysis tools such as SonarQube and Coverity:

[0048] 1. Traditional tools usually require complex configuration and frequent rule updates, while the present invention uses the natural language processing capabilities of LLMs to directly understand and execute high-level pseudocode, significantly improving the intelligence of static analysis.

[0049] 2. Traditional tools usually focus on specific programming languages ​​or specific types of errors and have poor adaptability to new programming paradigms or language features. LLMs are trained on multiple languages ​​and tasks and can handle multiple programming languages ​​and different types of analysis tasks. This flexibility allows the same tool to be widely used in multiple development environments and projects.

[0050] 3. Static analysis tools may miss some complex error patterns or generate more false positives, especially when analyzing deep logic errors. LLMs' advanced language understanding capabilities can be used to analyze the intent and potential errors of the code more deeply. At the same time, the accuracy of the analysis is ensured by verifying the execution results of the pseudocode, reducing false positives and missed reports.

[0051] 4. Traditional technologies require users to have a relatively high technical background to effectively use and interpret the analysis results. The present invention provides a more intuitive user interface and interaction method. Users can directly define analysis tasks in pseudocode, which lowers the threshold for use. At the same time, the output analysis results are more user-friendly, making it easier for users to understand and take action.

[0052] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A system for static code analysis using a large language model, including an agent module, characterized in that: The proxy module source code retrieval unit, pseudocode execution unit, and execution specification verification unit obtain pseudocode, source code, and task information, respectively. The pseudocode execution unit parses the pseudocode based on a large language model to perform static analysis and execute tasks, and simulates the analysis tasks described in the pseudocode. Based on the automatic retrieval of the source code performed by the source code retrieval unit and the task information, the execution specification verification unit performs specification verification of the execution process to obtain an execution result.

2. The system for performing static code analysis using a large language model according to claim 1, characterized in that: The pseudocode execution unit converts the pseudocode into a query executable by the large language model, simulates the pseudocode execution path through the internal mechanism of the large language model, and predicts the function to infer potential code behaviors and results.

3. The system for performing static code analysis using a large language model according to claim 1, characterized in that: The pseudocode specifies the analysis objectives, expected analysis types and execution parameters; in the execution specification verification unit, if the results are found to be inconsistent with expectations, the system will mark potential problem areas and can choose to re-execute or adjust model parameters for optimization.

4. The system for performing static code analysis using a large language model according to claim 1, characterized in that: The pseudo code is a specific static analysis pseudo code designed to simulate backward taint analysis or other static analysis processes through a large language model; In the execution specification verification unit, error classification and error attribution are performed using pseudo code input to identify and classify errors; Error classification involves using the results of pseudocode execution to determine the responsible function for the software crash; Bug attribution is the process of matching static analysis results with actual software patch records.

5. The system for performing static code analysis using a large language model according to any one of claims 1 to 4, characterized in that: The system also includes a task information module, a pseudo code module, a source code module, and a result output module respectively connected to the agent module; The task information module is used to obtain specific requirements for the analysis task; The pseudocode module is used to obtain pseudocode for defining a specified static analysis task; The source code module is used to obtain the actual source code; The result output module is used to output the final analysis report and intermediate results.

6. A method based on the system for static code analysis using a large language model according to claim 1, characterized in that The steps include: Step 1: Collect static analysis data; Step 2: Write or select pseudocode according to the static analysis goal; Step 3: Input pseudocode, static analysis data, and acquired task information into the large language model and configure it accordingly; Step 4: The large language model parses the pseudocode and simulates the execution of static analysis tasks based on the pseudocode logic and parsing results; The source code retrieval unit automatically retrieves relevant source code and provides it to the large language model as auxiliary information. It also monitors the execution process and records intermediate results. It verifies the intermediate and final results of the execution according to predefined specifications and standards, checks for violations and automatically classifies the violation types. It outputs static analysis results.

7. The method according to claim 6, characterized in that: In step 1, the collected data is preprocessed and converted into a format that can be understood and processed by a large language model, including a unified crash report format, normalized performance data, and crash reports from different sources are unified into a text format that is easy for a large language model to understand; The performance data is normalized and different normalization methods are used according to different performance indicators.

8. The method according to claim 6, characterized in that: In step 4, the large language model identifies variables, operators, control structures and understands logical relationships when parsing the pseudocode, and also parses comments and special tags in the pseudocode, annotating important prompt information, and special tags are special processing logic and / or special situations.

9. The method according to claim 6, characterized in that: In step 4, when a large language model is simulated to perform a static analysis task, code traversal, variable calculation tracking, and function call analysis are performed. For code traversal, the structure and semantics of the code need to be considered. According to the dependencies and calling order of the code, function calls and variable usage are traversed in a depth-first or breadth-first manner. In terms of variable value calculation and tracking, not only the current value of the variable is considered, but also its changes in different code blocks and the dependencies between variables. For the analysis of function calls, the function's parameter passing, return value, and possible side effects are considered to evaluate its impact on the entire program logic.

10. The method according to claim 6, characterized in that: In step 4, the automatic retrieval operation of the source code retrieval unit quickly locates relevant code fragments in the source code based on the retrieval algorithm and index structure, according to the task information and the key elements in the pseudocode; in the retrieval process, not only the text matching of the code but also the semantic matching of the code is considered. At the same time, the monitoring of the execution process records the timestamp and intermediate results of each operation in detail, which are used for performance analysis and result verification, and for tracing back the location and time of the problem when an exception occurs.

Citation Information

Patent Citations

  • Code analysis method, electronic equipment and computer readable medium

    CN119105799A

  • Code review and optimization method driven by large language model

    CN119512556A