Automatic bug fixing method for sensing dynamic program state based on large model
By obtaining dynamic program status and static information and combining it with the logical reasoning of a large language model, the problem of low vulnerability repair accuracy in existing technologies is solved, and efficient automated vulnerability repair is achieved.
Patent Information
- Application Number
- CN202510805841.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Existing automated vulnerability repair technologies rely on limited and uneven data sets, making them difficult to generalize to a wider range of vulnerability repair tasks. In addition, large language models lack dynamic program state information when understanding the nature of the vulnerability, resulting in low repair accuracy.
By constructing inputs to trigger vulnerabilities, monitoring program running behavior, obtaining crash-free constraints, and utilizing the logical reasoning capabilities of large language models combined with static and dynamic information, program debugging is performed and patches are generated and verified.
It significantly improves the accuracy and effectiveness of vulnerability repair, achieves more precise patch generation, and approaches the repair results of human experts.
Smart Images

Figure CN120705878A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more specifically, to an automatic vulnerability repair method based on large model perception of dynamic program status. Background Art
[0002] Software vulnerabilities are common program defects that attackers can exploit to gain unauthorized access to programs or trigger unexpected program behavior, resulting in financial losses, leaks of important information, and other consequences. According to CVEDetails, the number of vulnerabilities has continued to climb in recent years, with 40,296 new vulnerabilities reported in 2024, and this number is expected to increase further in 2025. This trend indicates that software vulnerabilities have become a significant source of security threats. However, manually remediating vulnerabilities is a challenging task that requires time and expertise. Edgescan reports show that the average remediation cycle for severe vulnerabilities is as long as 65 days, during which time the program remains exposed to potential attack. Therefore, automated vulnerability remediation technology is urgently needed to enhance software security.
[0003] In the past, automated vulnerability repair was a challenging task due to the diverse types of vulnerabilities, complex triggering conditions, difficult verification, and limited code generation capabilities. In recent years, as Transformer-based pre-trained models have demonstrated excellent performance in code understanding and generation, researchers have proposed AI-based automated vulnerability repair methods such as Vrepair and VulRepair. These methods use vulnerability datasets to train models or fine-tune existing pre-trained models, aiming to obtain models that can understand software vulnerabilities and then generate repaired code or patches. Although these methods have made significant progress, their accuracy remains low (usually no more than 25%), which may be attributed to the following factors:
[0004] Dataset Limitations: The effectiveness of existing methods is highly dependent on dataset quality. Models trained on specific datasets often struggle to generalize to a wider range of vulnerability remediation tasks. The quality of existing vulnerability datasets varies widely: VulGen analysis shows that most high-quality datasets are limited in size, while large-scale datasets often contain significant errors and noise, failing to accurately reflect the true nature of vulnerabilities.
[0005] Model limitations: Vulnerability repair is a complex, multi-stage process encompassing fault location, root cause analysis, fix location, and patch generation. Questions remain about whether the parameters of pre-trained models like CodeT5, used by Vrepair and VulRepair, are sufficient to handle this multi-dimensional task.
[0006] In recent years, large language models (LLMs) have demonstrated significant advantages. In the field of program analysis, LLMs demonstrate stronger code understanding and generation capabilities than traditional pre-trained models, and related research has covered areas such as defect detection, logical reasoning, and program debugging. In the vulnerability repair task, Pearce et al. first evaluated the ability of LLMs to generate security patches in zero samples. However, this method only regards LLMs as enhanced pre-trained models, and their performance is limited by the distribution of training data, resulting in poor repair effects for unseen vulnerability types. Due to the lack of targeted training, this method has not yet fully unleashed the potential of LLMs. Although the LLM4CVE framework has established an automated iterative repair process, its reliance on manual input leads to low repair efficiency.
[0007] Furthermore, researchers have proposed using multi-agent collaboration to guide large language models in bug fixing. This approach leverages a toolkit to provide the model with static analysis information of the source code, treating the large language model as an interactive agent rather than simply a code generation tool, thereby fully leveraging its capabilities. However, while static analysis can provide additional information beyond the faulty code snippet, it still lacks some critical details.
[0008] Take the divide-by-zero vulnerability CVE-2016-3623 as an example: in this vulnerability, the variable vertSubSampling is incorrectly set to zero and then propagated through multiple function calls and conditional checks. Using static information alone, it's difficult to precisely locate the zeroing location of vertSubSampling and the path along which the zero value propagates. Similarly, for static analysis challenges like buffer and heap overflows, relying solely on static information makes it difficult for large language models to accurately understand the nature of the vulnerability.
[0009] In order to solve the above problems, it is necessary to propose a vulnerability repair method that can not only analyze the static information of the program, but also obtain key information such as the dynamic state of the program, deeply analyze the root cause of the vulnerability, and make full use of the capabilities of the large language model to accurately complete complex vulnerability repair tasks. Summary of the Invention
[0010] The purpose of the embodiments of the present disclosure is to provide an automatic vulnerability repair method based on a large model to perceive the dynamic program state. The present invention proposes a vulnerability repair method to solve the problem that static analysis can provide additional information beyond the erroneous code fragment but lacks some key details.
[0011] In a general aspect, an automatic vulnerability repair method based on large model perception of dynamic program state is provided.
[0012] Step 1: Run the verification step to verify the existence of the vulnerability and extract the crash-free constraint;
[0013] Step 2: Obtaining the expected state in the repair positioning step, obtaining the expected state based on the no-crash constraint, specifically including obtaining the expected state at the crash site and obtaining the expected state at the potential repair location;
[0014] Step 3: Fix the debugging process in the positioning step, provide an interface set for LLM, obtain the actual status of the program based on program debugging, and obtain static information based on the tool set;
[0015] Step 4: Summarize the debugging information and generate and verify the patch.
[0016] The specific method of verifying the existence of the vulnerability through operation is: constructing an input that can trigger the target vulnerability and submitting it to the program to be analyzed; monitoring the runtime behavior of the program and observing the abnormal response of the program.
[0017] The obtaining of the expected state at the crash point uses the no-crash constraint in step 1 to represent the expected state at the crash point.
[0018] The method of obtaining the expected state of the potential repair location is implemented based on a lightweight guidance method based on prompt engineering. The logical reasoning capability of LLM is used to replace part of the symbolic execution calculation. For a given breakpoint line, the source code of the code line to the crash location is first provided to LLM. Then, the prompt "Think of the constraint and expected state of the program here based on the crash-free constraint. Compare it with the real state of the program to deepen the understanding of the bug" is used to guide LLM to infer the expected state at the breakpoint. LLM outputs the constraint at the breakpoint in a specific format and compares it with the real state during the debugging process.
[0019] The interface set includes: a static information interface for obtaining static information, which enables LLM to deeply understand the code structure and assist it in logical reasoning and breakpoint decision-making; a dynamic interface for obtaining dynamic information, which allows LLM to monitor program status in real time and perform reasoning in combination with crash site data.
[0020] The method for obtaining the true state of a program based on program debugging includes: first, executing the program for the first time with the help of a debugging tool, triggering a program crash to obtain initial error information to guide subsequent debugging; then, based on the call stack information provided by the error report, the LLM intelligently identifies suspicious functions that may cause the crash and initiates an independent debugging session for each suspicious function; finally, based on the expected state information obtained, the LLM retrieves relevant static information and checks the information of specific variables or expressions.
[0021] The specific method of summarizing the debugging information is as follows: when the LLM determines that sufficient information has been collected during the debugging process, it guides it to summarize the information to determine the root cause of the vulnerability and potential repair locations. The debugging information summary S can be formally defined as:
[0022]
[0023] Where: n represents the total number of debugging iterations; ψ i is the expected state prediction at the i-th iteration; Γ i is the actual state observation of the corresponding number; operator Represents the comparative analysis results of LLM on the two; r and l represent the final root cause analysis and repair positioning, respectively. Represents the reasoning and analysis process of LLM.
[0024] The specific method of patch generation and verification is as follows: the patch generation task is handled by an independent agent through a new session to avoid lengthy debugging conversations exceeding the token limit. The patch generation process uses the root cause summary and repair location obtained in the previous step as input. First, the LLM reviews the code at the repair location and evaluates the feasibility of direct repair. If direct repair is not feasible due to positioning deviation, the LLM performs secondary positioning based on root cause analysis. During this process, the LLM continuously checks the code segments involving relevant variables and functions, and gradually evaluates their repairability. When the LLM determines that a certain location is repairable, a patch to be verified is generated; during the operation process, basic syntax errors are checked first, and then the patch is replaced with the source code and the vulnerability exploit program is re-executed. By compiling the patched project and re-running it, it is verified whether the patch eliminates the original vulnerability. If the vulnerability exploit no longer triggers a crash, the process is terminated; otherwise, a new round of iteration is started for other potential repair locations.
[0025] The technical effects to be achieved by the embodiments of the present invention are:
[0026] This invention provides a novel, dynamic, state-aware large language model agent for automated vulnerability remediation. This system obtains the actual operating state through program debugging and derives the expected state based on crash-free constraints. By continuously comparing these two states, the invention provides a deeper understanding of the vulnerability's nature, enabling the generation of more accurate and effective patches. This invention significantly improves the accuracy, effectiveness, and practical usability of patch repairs. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and other objects and features of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings.
[0028] Figure 1 1 is a schematic diagram illustrating an architecture diagram of an automatic vulnerability repair method based on a large model sensing dynamic program state according to an embodiment of the present disclosure;
[0029] Figure 2 is a flowchart diagram illustrating a debugging process according to an embodiment of the present disclosure;
[0030] Figure 3 is a flowchart diagram illustrating a patch generation and verification process according to an embodiment of the present disclosure;
[0031] Figure 4 is a schematic diagram illustrating an example of a patch for vulnerability CVE-2016-5321 according to an embodiment of the present disclosure;
[0032] Figure 5 is a schematic diagram illustrating a process of repairing a vulnerability according to an embodiment of the present invention;
[0033] Figure 6 FIG. 4 is a schematic diagram showing a constraint conversion template according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following detailed description is provided to help the reader gain a comprehensive understanding of the methods, devices and / or systems described herein. However, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear after understanding the disclosure of the present application. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and conciseness, descriptions of features known in the art may be omitted.
[0035] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will become clear after understanding the disclosure of this application.
[0036] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.
[0037] Although terms such as "first," "second," and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are used solely to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Thus, what is referred to as a first member, first component, first region, first layer, or first portion in the examples described herein may also be referred to as a second member, second component, second region, second layer, or second portion without departing from the teachings of the examples.
[0038] In the specification, when an element (such as a layer, region, or substrate) is described as being “on,” “connected to,” or “coupled to” another element, the element may be directly “on,” “connected to,” or “coupled to” the other element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on,” “directly connected to,” or “directly coupled to” another element, there may be no other elements present therebetween.
[0039] The terms used herein are intended only to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular is intended to include the plural. The terms "comprise," "include," and "have" indicate the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0040] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains after understanding the present disclosure. Unless expressly defined otherwise herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.
[0041] Furthermore, in describing the examples, when it is deemed that a detailed description of well-known related structures or functions would cause ambiguous interpretation of the present disclosure, such detailed description will be omitted.
[0042] Figure 1 3 is a schematic diagram illustrating an automatic vulnerability repair method based on large model perception of dynamic program status according to an embodiment of the present disclosure.
[0043] In order to achieve the above-mentioned purpose of the invention, the technical framework adopted by the present invention is as follows Figure 1 shown.
[0044] Step 1: Verify the existence of the vulnerability by running and extract the crash-free constraint;
[0045] 1. Obtain preliminary output and error information by running a proof-of-concept vulnerability. In this phase, researchers construct an input that triggers the target vulnerability (e.g., a malformed file, abnormal network data, or a specific API call) and submit it to the program being analyzed. By monitoring the program's runtime behavior, they can observe abnormal responses, such as crashes, assertion failures, or memory errors. The core goal of this step is to reproduce the vulnerability and capture key information when the program crashes (e.g., signal type, crash address, or exception code). In this step, the present invention uses a sanitizer as a tool to generate abnormal responses. For example, when analyzing an image processing tool with a buffer overflow vulnerability, researchers can submit a deliberately malformed PNG file. After running the program, they may observe a segmentation fault and an operating system crash report (e.g., a core dump in Linux or a Dr. Watson log in Windows). After the program crashes, debugging tools (e.g., GDB, WinDbg, or LLDB) generate a stack trace, which contains the function call sequence and memory state at the time of the crash. The LLM (or researchers) parses this information to identify key functions that may be involved in the vulnerability.
[0046] The extraction of crash-free constraints. The crash-free constraint was first proposed by ExtractFix. Its core idea is to extract the constraint conditions that can avoid program crashes through program analysis techniques and use these constraints to guide patch generation. The core advantage of this technology is that it can avoid patch overfitting, that is, the generated patch can only fix the crashes in specific test cases and cannot be generalized to other input scenarios. ExtractFix relies on memory detection tools (such as AddressSanitizer, Valgrind, or UndefinedBehaviorSanitizer) to collect runtime information (such as out-of-bounds memory access, null pointer dereference, etc.) when the program crashes and extract crash-inducing constraints from it. Subsequently, it generates repair patches that meet the crash-free conditions through constraint solving and program synthesis techniques. ExtractFix proposed a method for extracting crash-free constraints, and this invention uses the same method as ExtractFix to extract crash-free constraints. We simplified the constraint expressions extracted by ExtractFix and designed a set of intuitive templates to convert them into natural language descriptions. For all types of constraints extracted by ExtractFix, this invention designed natural language templates one by one. For example, the constraint condition "start<initial_read" can be converted into the natural language expression "Variable start should be less than variable initial_read", which helps the large model to more accurately understand the crash-free constraint and achieve a precise assessment of the program state. Figure 6 Shows the templates for converting various constraints in this invention.
[0047] Step 2: Obtain the expected state based on the crash-free constraint.
[0048] 1. Obtain the expected state at the crash point. When the program crashes, the crash-free constraint (CFC) can accurately describe the safety conditions that the program should meet at the crash point. For example, if the crash is caused by a buffer overflow, the CFC may require buffer_size >= data_length; if the crash is caused by a null pointer dereference, the CFC may require ptr!= NULL. This invention uses the crash-free constraint CFC extracted in Step 1 to represent the expected state at the crash point.
[0049] 2. Obtain the expected state of the potential repair location. Although CFC can describe the expected state of the crash point, fixing vulnerabilities usually requires imposing constraints at earlier code locations (such as input checks, memory allocation, etc.) rather than just adding protection at the crash point. Developers usually expect to obtain a specific program state at each breakpoint. ExtractFix uses forward symbolic execution to propagate constraints, that is: trace back from the crash point to calculate which variables affect the CFC; try to insert checks at different positions in the program to ensure that the CFC is ultimately satisfied. However, this method has two major problems:
[0050] 1) Incomplete constraint propagation: Symbolic execution can only propagate variable constraints directly related to the CFC, and cannot infer other key program states (such as loop invariants, global configurations, etc.).
[0051] Example: If the CFC is input_len < MAX_LEN, symbolic execution may not be able to infer whether MAX_LEN is correctly set during initialization.
[0052] 2) High performance overhead: Re-executing symbolic execution every time the LLM selects a new debugging target (such as a line of code) has extremely high computational costs and is difficult to apply to large projects.
[0053] Therefore, to solve the above problems, the present invention proposes a lightweight guiding method based on prompt engineering, using the logical reasoning ability of the LLM to replace part of the symbolic execution calculation. The lightweight guiding scheme based on the chain of thought (CoT) technology shows better performance in this problem. Specifically, for a given breakpoint line, first provide the source code from the code line to the crash point to the LLM, and then use the prompt "Think of the constraint and expected state of the program here based on the crash-free constraint. Compare it with the real state of the program to deepen the understanding of the bug." to guide the LLM to infer the expected state at the breakpoint. The LLM outputs the constraints at the breakpoint in a specific format and compares them with the real state during the debugging process. This method can effectively enhance the LLM's awareness of the importance of the current debugging target and significantly reduce irrelevant or inefficient debugging behaviors by strengthening the logical reasoning ability.
[0054] Step 3: Obtain the real state of the program based on program debugging.
[0055] 1. To achieve automatic program debugging, the present invention provides a complete set of interfaces for LLM, covering the two core functions of static code analysis and dynamic runtime monitoring, so that it can obtain necessary debugging information like human developers. The present invention provides the following interfaces for LLM:
[0056] 1) Obtaining static information: The static information interface enables LLM to deeply understand the code structure and assist in logical reasoning and breakpoint decision-making:
[0057] Definition: This interface is used to query the declaration location and type information of a variable, function, or data structure. For example, if LLM finds that buffer_size may cause an overflow, it can use this interface to obtain its definition (e.g., #define buffer_size64) to confirm its actual capacity.
[0058] Summary: Returns the identifier of the function / variable and its associated comments, helping to understand obscure code logic. For example, if a function parameter is named ctx, the comment may reveal that it actually means "network connection context".
[0059] function_body: When LLM needs to analyze the internal logic of a function (such as loop boundaries or conditional branches), it can obtain the source code through this interface.
[0060] get_file_content: Extracts code by line number range and automatically annotates syntax structures (such as highlighting conditional statements). LLM can use this to quickly locate key logic snippets.
[0061] 2) Obtain dynamic information. The dynamic interface allows LLM to monitor program status in real time and perform reasoning based on crash scene data:
[0062] run_program: Calls a debugger (such as GDB) to execute the target program and returns the error type, stack trace, and register status.
[0063] run_to_line: Sets a breakpoint at a suspicious location, pauses execution, and freezes the program state for LLM to inspect variable values or memory layout.
[0064] print_value: Get the real value of a specific variable or expression in the real state.
[0065] 2. Toolset-based debugging process. Figure 2 This article demonstrates the entire process of debugging a program using a large language model (LLM) to determine the actual running status. The debugging process mainly includes the following key steps:
[0066] 1) First, call the run_program interface to execute the program's first run with the help of the debugging tool. The main purpose of this step is to trigger the program to crash, thereby obtaining initial error information to guide subsequent debugging.
[0067] 2) Based on the call stack information provided by the error report, LLM intelligently identifies suspicious functions that may cause the crash and starts a separate debugging session for each suspicious function. In a single debugging session, LLM implements the following operations by calling the run_to_line interface:
[0068] Set a breakpoint within the target call frame.
[0069] Make the program pause at a specific line of code in a suspicious function.
[0070] At this point, you can fully obtain the actual running status of the program at this breakpoint.
[0071] 3) As described in step 2, based on the expected state information, LLM retrieves relevant static information and calls the print_value interface to check the information of a specific variable or expression.
[0072] 3. Static information acquisition based on tool set. The static information acquisition system of the present invention provides LLM with comprehensive code understanding capabilities, enabling it to deeply analyze program structures like experienced developers. These interfaces not only provide basic code information, but also enhance the accuracy of debugging by intelligently integrating multi-source data. Source code and clear line numbers are crucial for LLM to select effective breakpoint locations, so the present invention provides a get_file_content interface to obtain code segments with line number annotations. Since it is difficult to determine the type of a symbol based on its name alone, the definition statement of the symbol is obtained through the definition interface. Although this is sufficient in most cases, when encountering functions with obscure type aliases or parameter names, the definition statement may lose its reference value. Fortunately, developers usually leave explanatory documents in comments. Based on this, the summary interface is designed to generate an analyzed and organized summary description for the symbol by parsing all relevant type aliases and extracting surrounding comments. In addition, a function_body interface is also provided to obtain the complete source code of the function used in the current context.
[0073] Step 4: Patch generation and verification.
[0074] 1. Summary of debug information. When the LLM determines that sufficient information has been collected during the debugging process, we will guide it to summarize the information to determine the root cause of the vulnerability and potential repair locations. The debug information summary S can be formally defined as:
[0075]
[0076] Where: n represents the total number of debugging iterations; ψ i is the expected state prediction at the i-th iteration; Γ i is the actual state observation of the corresponding number; operator Represents the comparative analysis results of LLM on the two; r and l represent the final root cause analysis and repair positioning, respectively, and the symbol This formula indicates that the debugging summary contains two core elements: vulnerability root cause analysis based on the context of the reproduction environment and precise repair location determination.
[0077] 2. Patch generation and verification. The patch generation task is handled by a separate agent through a new session to avoid lengthy debugging conversations exceeding the token limit. Using a new agent to perform this task allows LLM to focus more on this specific goal while reducing redundant information interference. Figure 2 As shown, the patch generation process takes the root cause summary and repair location obtained in the previous step as input. First, LLM will review the code at the repair location and evaluate the feasibility of direct repair. If direct repair is not feasible due to positioning deviation, LLM will perform secondary positioning based on root cause analysis. During this process, LLM continuously checks the code segments involving relevant variables and functions, and gradually evaluates their repairability. When LLM determines that a certain location is repairable, it generates a patch to be verified. The present invention first checks for basic syntax errors (such as mismatched brackets), then replaces the patch with the source code and re-executes the vulnerability exploit program. By compiling the patched project and rerunning it, it will be verified whether the patch has eliminated the original vulnerability. If the vulnerability exploit no longer triggers a crash, the process terminates; otherwise, a new round of iteration will be started for other potential repair locations.
[0078] Complete process example:
[0079] Vulnerability Analysis: Figure 4Shows a typical buffer overflow vulnerability (CVE-2016-5321) existing in the libtiff image processing library. This vulnerability stems from the lack of security checks on a fixed-size array. In the code implementation, the developer defined a pointer array named srcbuffs, whose length is limited by the MAX_SAMPLES macro constant. However, during the subsequent loop processing, the code directly uses the variable s as the array index to access the elements of the srcbuffs array without any boundary check on the value of s. This programming oversight will cause the program to access memory areas outside the allocated range of the array when the value of s exceeds MAX_SAMPLES. According to the C language standard, such out-of-bounds access behavior is undefined behavior and may trigger various abnormal situations in actual operation: at worst, it may cause the program to read incorrect data, and at best, it may lead to memory corruption or even be exploited to execute arbitrary code. The repair solution for this vulnerability is clear. Just adding the boundary check s < MAX_SAMPLES in the loop condition can ensure that the array index is always within the safe range. This case typically demonstrates the necessity of strict boundary checks on fixed-size arrays in C / C++ programming and is also one of the common patterns of many memory security vulnerabilities.
[0080] In the vulnerability analysis experiment, we submitted the function code with security defects to GPT-4 for automated analysis. The model analyzed the memory allocation statement tbuff = (unsigned char
[0081] *)_TIFFmalloc(tilesize+8) produces a seemingly professional analysis report, stating: "This operation dynamically allocates tilesize+8 bytes of heap memory for tbuff. The additional 8 bytes may be used for memory alignment padding or as a safety buffer to prevent potential buffer overflows. However, it is worth noting that neither the code comments nor the context clearly explain the specific purpose of these 8 bytes, and the necessary bounds checking mechanisms are also lacking during subsequent buffer usage." This analysis superficially demonstrates the model's sensitivity to code security, but in reality, it exposes a significant flaw in LLM's program understanding. A thorough analysis of the entire code execution flow reveals a fundamental misunderstanding in GPT-4's judgment. In the real program architecture, the tbuff variable serves only as an intermediary for memory management, its core function being to facilitate the freeing of pointer resources in the srcbuffs array. Crucially, all values in tbuff are strictly derived from pre-initialized elements of the srcbuffs array, which in turn originate from legitimate memory allocations within tbuff itself. This circular reference relationship forms a complete memory safety loop, making buffer overflows physically impossible. Therefore, the so-called "security patches" proposed by GPT-4 are not only unnecessary but also disrupt the original memory management logic. This misjudgment profoundly reveals the limitations of current large language models in understanding program semantics. To systematically evaluate the vulnerability remediation capabilities of LLMs, we designed a rigorous experimental protocol. The research team selected five types of diagnostic information commonly used in software development as auxiliary inputs: 1) vulnerability type classification (such as buffer overflow and null pointer dereference); 2) specific code location of the program crash; 3) runtime error message and exception code; 4) input constraints that trigger the crash; and 5) proof-of-concept (POC) code that reproduces the vulnerability. This information was progressively combined to GPT-4, and its output was quantitatively evaluated on three key dimensions: the accuracy of the vulnerability root cause analysis, the rationality of the remediation strategy, and the correctness of the patch code. Experimental data showed that when provided with any type of information alone, the model performed no better than random guessing. Even when given a complete combination of diagnostic information, the correctness of the patches generated by GPT-4 remained below 15%, and most of the so-called "fixes" introduced new logical errors.
[0082] Figure 5The present invention demonstrates the process of repairing the vulnerability. The entire automated debugging process begins with LLM triggering the execution of the target program through the run_program interface. When the program crashes due to a vulnerability, the system will capture detailed diagnostic information, including but not limited to the specific error type (such as a segmentation fault or stack overflow), the precise memory address that caused the crash, the processor register status, and the complete function call stack information. These raw data are standardized and processed to form a structured report. Based on these runtime information collected in real time, LLM will use its code understanding capabilities to intelligently analyze the call stack, accurately locate the key functions that are most likely to cause the crash and their corresponding source code line numbers, and identify related local variables and parameters. The system then dynamically obtains 50 lines of intelligently annotated code context around the crash point through the get_file_content interface. These codes not only contain syntax highlighting and variable scope markers, but also integrate the data flow relationships and control flow paths obtained by static analysis, building a complete program execution environment model for LLM. During the deep debugging phase, LLM accurately sets breakpoints at key control flow nodes based on its understanding of code semantics. It interacts with the debugger through the run_to_line interface, precisely controlling program execution until it pauses at the target location. It then uses the print_value interface to obtain the memory state and specific values of key variables in real time. For example, in this case, it successfully captured the decisive evidence that the spp variable had a value of 9 and the s variable had a value of 8. The system's built-in constraint analysis engine automatically compares these runtime observations with the safety constraints pre-derived using the ExtractFix algorithm, generating a detailed analysis report with visual difference markers. This report not only clearly identifies existing boundary condition violations but also infers possible anomalous execution paths. Based on these analysis results, LLM initiates multiple iterative debugging sessions, tracing the historical path of suspicious variable assignments, analyzing their initialization and possible sources of contamination, and verifying the extent to which external input parameters influence program state. This process continuously refines its understanding of program behavior, ultimately building a comprehensive cognitive model of the vulnerability's root cause. During the repair solution generation phase, LLM uses its code generation capabilities to comprehensively consider the design intent of the original code (derived through analysis of function naming, comments, and call relationships), constraint violations during actual runtime, the valid value range of related variables, and the project's coding standards. The generated patch code not only adds necessary security check conditions, but also ensures that the original business logic invariants are not destroyed, while maintaining consistency in the code style.To ensure the quality of the fix, the system automatically triggers a complete verification pipeline, including static code style checking (to ensure compliance with the project's coding standards), regression test suite execution (to verify that existing functionality is intact), and targeted fuzz testing (specifically stress testing the fixed vulnerability). This series of automated verification measures fully reflects the quality assurance thought process of professional software developers. This innovative approach, which deeply integrates dynamic debugging information with static code analysis, abstracts low-level debugging details into high-level semantics understandable to LLM through a carefully designed tool chain interface. This effectively overcomes the limitations of traditional zero-shot learning models, which lack domain knowledge, enabling LLM to demonstrate practical value close to that of human experts in real-world software vulnerability diagnosis and repair scenarios, providing a new technical path for automated software maintenance.
[0083] While some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that changes may be made to these embodiments without departing from the principles and spirit of the disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. An automatic vulnerability repair method based on large model perception of dynamic program status, characterized in that: Step 1: Run the verification step to verify the existence of the vulnerability and extract the crash-free constraint; Step 2: Obtaining the expected state in the repair positioning step, obtaining the expected state based on the no-crash constraint, specifically including obtaining the expected state at the crash site and obtaining the expected state at the potential repair location; Step 3: Fix the debugging process in the positioning step, provide an interface set for LLM, obtain the actual status of the program based on program debugging, and obtain static information based on the tool set; Step 4: Summarize the debugging information and generate and verify the patch.
2. The automatic vulnerability repair method based on large model perception of dynamic program status according to claim 1, characterized in that: The specific method of verifying the existence of the vulnerability through operation is: constructing an input that can trigger the target vulnerability and submitting it to the program to be analyzed; monitoring the runtime behavior of the program and observing the abnormal response of the program.
3. The automatic vulnerability repair method based on large model perception of dynamic program status according to claim 2, characterized in that: The obtaining of the expected state at the crash point uses the no-crash constraint in step 1 to represent the expected state at the crash point.
4. The automatic vulnerability repair method based on large model perception of dynamic program status as described in claim 3 is characterized in that: The method of obtaining the expected state of the potential repair location is implemented based on a lightweight guidance method based on prompt engineering. The logical reasoning capability of LLM is used to replace part of the symbolic execution calculation. For a given breakpoint line, the source code of the code line to the crash location is first provided to LLM. Then, the prompt "Think of the constraint and expected state of the program here based on the crash-free constraint. Compare it with the real state of the program to deepen the understanding of the bug" is used to guide LLM to infer the expected state at the breakpoint. LLM outputs the constraint at the breakpoint in a specific format and compares it with the real state during the debugging process.
5. The automatic vulnerability repair method based on large model perception of dynamic program status according to claim 4 is characterized in that: The interface set includes: a static information interface for obtaining static information, which enables LLM to deeply understand the code structure and assist it in logical reasoning and breakpoint decision-making; a dynamic interface for obtaining dynamic information, which allows LLM to monitor program status in real time and perform reasoning in combination with crash site data.
6. The automatic vulnerability repair method based on large model perception of dynamic program status according to claim 5, characterized in that: The method for obtaining the true state of a program based on program debugging includes: first, executing the program for the first time with the help of a debugging tool, triggering a program crash to obtain initial error information to guide subsequent debugging; then, based on the call stack information provided by the error report, the LLM intelligently identifies suspicious functions that may cause the crash and initiates an independent debugging session for each suspicious function; finally, based on the expected state information obtained, the LLM retrieves relevant static information and checks the information of specific variables or expressions.
7. The automatic vulnerability repair method based on large model perception of dynamic program status according to claim 6 is characterized in that: The specific method of summarizing the debugging information is as follows: when the LLM determines that sufficient information has been collected during the debugging process, it guides it to summarize the information to determine the root cause of the vulnerability and potential repair locations. The debugging information summary S can be formally defined as: Where: n represents the total number of debugging iterations; ψ i is the expected state prediction at the i-th iteration; Γ i is the actual state observation of the corresponding number; operator Represents the comparative analysis results of LLM on the two; r and l represent the final root cause analysis and repair location, respectively. Represents the reasoning and analysis process of LLM.
8. The automatic vulnerability repair method based on large model perception of dynamic program status according to claim 7 is characterized in that: The specific method for patch generation and verification is as follows: the patch generation task is handled by an independent agent through a new session to avoid lengthy debugging conversations exceeding the token limit. The patch generation process uses the root cause summary and repair location obtained in the previous step as input. First, the LLM reviews the code at the repair location and evaluates the feasibility of a direct repair. If a direct repair is not feasible due to positioning deviation, the LLM performs a secondary positioning based on the root cause analysis. During this process, the LLM continuously checks the code segments involving relevant variables and functions, gradually evaluating their repairability. When the LLM determines that a location is repairable, it generates a patch for verification. During the run, it first checks for basic syntax errors, then replaces the patch with the source code and re-executes the exploit. By compiling the patched project and re-running it, it verifies whether the patch eliminates the original vulnerability. If the exploit no longer triggers a crash, the process terminates. Otherwise, a new iteration will be initiated for other potential repair locations.
Citation Information
Cited By
Automatic program repairing method combining executable invariant and differential signal
CN121255253A
AI-based SaaS system vulnerability automatic repair method and system
CN121579265A