Debugger testing method based on large language model

Through the debugger testing method based on large language model, test cases with diverse and complex C++ syntax structure are generated, which solves the problem of insufficient debugger testing capabilities in the existing technology, and improves the debugger test coverage and defect detection capabilities.

CN120216370APending Publication Date: 2025-06-27DALIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510296985.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing debugger testing methods have limited effectiveness in improving debugger testing capabilities. The main reason is that these methods focus on changing the program's control flow and data flow, and cannot fully detect the debugger's accuracy and completeness of debugging information.

Method used

Using a debugger testing method based on a large language model, high-quality test cases are automatically generated by constructing random function regulations, syntax variations and combined compilation strategies. These test cases are compileable and executable, and cover rich C++ syntax features and deep function call structures to simulate complex software debugging scenarios.

Benefits of technology

Improve the debugger's test coverage and defect detection capabilities, ensuring that the debugger can maintain accurate functional performance when facing complex debugging information, thereby more comprehensively evaluating its reliability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216370A_ABST
    Figure CN120216370A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of software testing, and relates to a debugger testing method based on a large language model. The method comprises the steps that firstly, a random function protocol is generated, the random function protocol is constructed by a statement template and semantic constraints, the generated random function protocol serves as a cue word to be input into a large language model, and a function template meeting description of the random function protocol is generated; thirdly, compiling the generated function template, and if compiling of the function template fails, regenerating a function protocol and performing iterative correction until codes meet grammar and semantic constraint requirements; and if the function template is successfully compiled, reserving the function template for subsequent use. And secondly, performing grammar variation on the successfully compiled function template, and outputting a source file after the grammar variation. And combining and compiling the source file output after variation with the corresponding header file and the test case. And generating an executable file after combination and compilation. And finally, inputting the executable file into a debugger for execution, and detecting whether a bug exists or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software testing, and relates to a debugger testing method based on a large language model, which can be used to automatically test a debugger. Background Art

[0002] Debuggers play an indispensable role in modern software development. They enable developers to obtain controllable observation and feedback during program execution, thereby effectively diagnosing and fixing defects in the code. Currently, GDB (GNU Debugger) and LLDB (Low-Level Debugger) are two widely used debugging tools, occupying key positions in the GNU and LLVM development ecosystems respectively. These debuggers usually have functions such as breakpoint setting, step-by-step execution, variable tracking, memory inspection, and stack backtrace. Through these functions, developers can precisely control the execution flow of the program, and then efficiently discover problems, analyze program behavior, and optimize performance. With the rapid expansion of software system scale and the increasing complexity of the architecture, the role of debuggers becomes even more important. Although debuggers are crucial in software development, their correctness is difficult to fully guarantee. Debuggers rely on debugging information generated by compilers, such as DWARF or PDB formats, which map the relationship between source code and machine code. If a debugger makes an error in parsing this information, it may lead to incorrect variable values, incomplete function call stacks, or even loss of key information, thus affecting program debugging. To improve the reliability of debuggers, they must be rigorously tested.

[0003] Existing debugger testing methods usually rely on automated test case generation tools, such as CSmith and YarpGen. These tools are widely used to generate random test cases to evaluate the performance of debuggers. CSmith is a test tool for randomly generated C language programs, which tests the capabilities of compilers and debuggers by generating programs with random control flow and data flow; YarpGen is mainly used to generate software test cases with specific structures and complexities. To improve the test coverage and diversity, existing methods combine the test cases generated by these two tools and use a coverage-guided mutation method to mutate the generated seed programs. Specifically, the mutation method simulates various program states by introducing fine-grained mutation operations (such as adjusting control flow, modifying data flow, and changing syntax structures), thereby increasing the probability of the debugger encountering boundary cases and potential errors. By mutating the generated seed programs, the complexity and diversity of test cases can be effectively improved, enabling the debugger to be verified in a wider range of program scenarios, and thus improving the reliability and accuracy of the debugger when dealing with complex programs.

[0004] Existing testing methods have limited effectiveness in enhancing the testing capabilities of debuggers. The reason may be that these methods mainly focus on changing the control flow and data flow of programs, and this type of mutation is usually used to evaluate the effectiveness of compiler optimization processes, rather than the debugger's handling of the accuracy and integrity of debugging information. The core task of a debugger is to accurately map the relationship between source code and machine code based on the debugging information generated by the compiler. If the debugger fails to correctly parse this information, it may lead to the presentation of incorrect program behaviors or misjudgments during the debugging process. Therefore, simple control flow and data flow mutations are not sufficient to comprehensively detect the capabilities of debuggers. To effectively evaluate the performance of debuggers, test cases should cover a richer set of syntactic structures, not only testing common execution paths but also including code segments that can generate complex debugging information. In addition, most of the differences between debuggers come from different interpretations of function inlining, and deep function calls are significantly meaningful for testing debuggers. Such diverse test cases can prompt the debugger to maintain accurate functional performance when faced with more complex debugging information, thus more comprehensively evaluating its reliability and stability. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a testing method for debuggers based on large language models. This method is specifically tailored to the characteristics of debuggers and automatically generates high-quality test cases that can be used to test debuggers, in order to improve the test coverage rate and defect detection capabilities of debuggers. The method of the present invention combines the code generation capabilities of large models and ensures that the generated test cases are compilable, executable, and cover rich C++ syntactic features and deep function call structures by constructing random function specifications, syntactic mutations, and combined compilation strategies, so as to simulate complex software debugging scenarios.

[0006] A testing method for debuggers based on large language models, the specific steps are as follows:

[0007] Step (1): Construct a random function specification through statement templates and semantic constraints, and use the constructed random function specification as a prompt to input into a large language model to generate a corresponding function template. Subsequently, the generated function template is subjected to compilation verification. If the compilation is successful, the function template is retained for subsequent use; if the function template fails to compile, a new random function specification is generated and iteratively corrected until the compilation is successful.

[0008] Step (2): Based on the function template that successfully passed compilation in step (1), perform semantic mutation processing and write the mutated function template into the source file.

[0009] Step (3): Combine and compile the source file with the corresponding header files and test cases to generate an executable file.

[0010] Step (4): Run the generated executable file and test it on a debugger to verify the correctness of the debugging information and detect whether there are bugs.

[0011] Further, step (1) specifically includes the following steps:

[0012] 1-1) Generate a random function specification. The function specification is composed of four basic statement templates (declaration, assignment, condition, loop). Randomly select one of them and include a randomly generated function signature, including the function name, return type, parameter types, and the number of parameters. Adding the corresponding template parameters to the statement template forms a complete prompt statement. The template parameters are the types and names of variables.

[0013] 1-2) During the generation process, to avoid semantic errors such as undefined variables and illegal access when the large model implements the function, introduce semantic constraints, that is, record the data types and their scopes of variables during the generation process, and maintain the current context file environment. The declaration statement adds new variable information to the current context. The remaining statements select variables with consistent types from the current environment as the parameters of the statement template to avoid semantic errors caused by types. When the next statement prompt to be generated leaves a certain scope, the variables declared in that scope are deleted from the context environment to avoid using undefined variables.

[0014] 1-3) The constructed random function specification is used as a prompt to input into the large language model, and the function template generated by the large language model is compiled and verified. If the compilation is successful, the function template is retained for subsequent use; if the function template compilation fails, a new random function specification is generated and iteratively corrected. This process will continue to iterate until all function templates can pass the compilation verification.

[0015] Further, step (2) specifically includes the following steps:

[0016] 2-1) Extract C++ language features from the official C++ standard document and build a mutation operator library based on the extracted C++ language features. This operator library covers a variety of advanced C++ syntax features, including but not limited to constexpr, move semantics, Lambda expressions, and structured bindings, etc. An operator is a specific prompt, such as "use alignas and alignof", which is often used for memory alignment to let the large model modify the function so that the function applies the C++ syntax specified by the prompt.

[0017] 2-2) The system randomly selects mutation operators from the mutation operator library and combines them with the mutation prompt template to generate mutation prompts and input them into the large language model, so that the large model mutates the function according to the mutation prompts to generate a mutated function.

[0018] 2-3) Compile and verify the mutated function. If the compilation is successful, write the mutated function template into the source file. If the compilation fails, select other mutation operators from the mutation operator library, input them into the large language model again to generate a new function, and then perform compilation verification again. If successful, continue to select operators for mutation. The parameter k for the number of mutations can be set, that is, apply the compilation mutation operator k times to each function.

[0019] Further, step (3) specifically includes the following steps:

[0020] 3-1) Based on the generated function, construct a complete project structure. The project includes multiple header files, source files, and test case files, and is stored in the corresponding source folder, header folder, and test case folder.

[0021] 3-2) Randomly select a source file from the source code folder, and extract the corresponding header files and test cases from the header folder and test case folder according to the relevance of the source file.

[0022] 3-3) Combine the obtained source file with its corresponding header files and test cases, and perform compilation to generate an executable file. The generated executable file is used as the executable test case for the debugger.

[0023] Compared with the prior art, the present invention has the following advantages and effects:

[0024] The present invention proposes a debugger testing method based on a large language model, which can efficiently and automatically test the debugger. This method uses the large language model to generate test cases with diverse and complex C++ syntax structures. These test cases can not only cover various common program behaviors, but also simulate complex execution paths and debugging information to ensure the effective testing of the debugger's performance when facing long function call chains. To improve the accuracy and success rate of test case generation, the present invention introduces templates with semantic constraints to generate test cases that conform to specific function specifications. This method ensures the logical consistency of the generated functions and effectively avoids possible syntax or functional errors during the generation process, thereby improving the quality and reliability of the testing. Description of the Drawings

[0025] Figure 1 It is a schematic flowchart of a debugger testing method based on a large language model according to the present invention.

[0026] Figure 2 It is a sub-flowchart of function template generation in a debugger testing method based on a large language model according to the present invention.

[0027] Figure 3 It is a process sub - graph of semantic mutation in a debugger testing method based on a large - language model of the present invention.

[0028] Figure 4 It is a process sub - graph of combined compilation in a debugger testing method based on a large - language model of the present invention. Detailed implementation manners

[0029] The method of the present invention will be described in detail below in combination with the accompanying drawings, technical solutions, and embodiments.

[0030] As Figure 1 shown, a debugger testing method based on a large - language model of the present invention is carried out according to the following process: First, generate a random function specification. The random function specification is constructed by a statement template and semantic constraints. The generated random function specification is used as a prompt word and input into the large - language model, and the large - language model generates a function template that conforms to the description of the random function specification. Subsequently, compile the generated function template. If the function template compilation fails, regenerate the function specification and perform iterative correction until the code meets the syntax and semantic constraint requirements; if the function template compilation is successful, retain it for subsequent use. Second, perform syntax mutation on the successfully compiled function template, and output the source file after syntax mutation. Then, perform combined compilation on the source file output after mutation, the corresponding header file, and test cases. After combined compilation, generate an executable file. Finally, input the generated executable file into the debugger for execution to detect whether there are bugs.

[0031] The implementation details of each process will be described in detail below in combination with specific examples. The detailed implementation manners are as follows:

[0032] (1) As Figure 2As shown in the figure, the first step is to generate a random function specification, that is, to construct a function signature through a randomization strategy and fill in the statement template to form a complete function specification. In this embodiment, first, a function signature is randomly generated, such as "template<typename T1> unsigned char func3(T1 p_0, double p_1)". This function is a template function, whose return type is unsigned char, and the parameters include the template parameter T1 and the double-type parameter p_1. During the process of generating the function body, the system needs to fill in the statement prompt words, which are composed of the statement template and the template parameters. The template parameters correspond to the types and names of variables. To ensure semantic correctness and suppress the "hallucination" phenomenon of the large language model during code generation, the present invention adopts a context variable environment management mechanism to dynamically maintain the valid variable information within the current scope. When the func3 function is generated, its parameters p_0 and p_1 are immediately added to the current context environment, and at the same time, T1 is also added to the set of optional types to ensure that the variable types generated subsequently meet the constraints.

[0033] When entering a new code block, the system first selects a declaration statement template to ensure that the variable initialization in the code block meets the requirements of the function return type. For example, in the function body of func3, the prompt of the first declaration statement template is "Declare 2 variables, the variables are 'unsigned char var1, unsigned char var2', and initialize.", that is, both variables var1 and var2 are initialized with the return type unsigned char of func3. After the declaration is completed, the system selects a statement template suitable for the current context from the statement templates of the assignment type, conditional type, and loop type. In this embodiment, func3 selects an assignment statement template. The system generates an assignment statement "Assignment statement, use variables var4->member_2, p_1 to assign value to variable var4->member_2" according to the existing variables in the current context, where both p_1 and var4->member_2 are of double type, meeting the type matching requirements. In addition, func3 also selects an assignment statement template based on function call "Declare 'unsigned long int var5' inited by function call func2(double var4->member_1, double var4->member_1)". In this statement, var5 is initialized by func2. The return type of func2 needs to match the type of var5, and its parameters need to be filled with variables that meet the double type from the context. When generating a conditional statement, func3 selects a conditional statement template "If statement, the conditional expression uses 1 operator". Since the conditional statement and the loop statement introduce new code blocks, it is necessary to reselect the declaration statement template and declare new variables in this code block, such as [unsigned char var10, unsigned char var11, unsigned char var12], and then add these variables to the current context. During the execution of the code block, the system repeats the above steps until all statement templates are filled and the function body reaches the expected length. When leaving the code block, the method of the present invention will automatically remove the variables declared inside the code block from the context to maintain the correctness of the variable scope. After the above steps are processed, the final complete function specification is as follows:

[0034]

[0035]

[0036] This function specification describes how to implement this function. Based on this description, the large language model gives the corresponding implementation, which is as follows:

[0037]

[0038]

[0039] After the function specification is generated, the generated function needs to be compiled and checked. If the compilation fails, the function specification is regenerated, and the large language model performs the function implementation until the generated code can pass the compilation correctly. In this embodiment, the generated func3 successfully passes the compilation, and there is no need to regenerate the specification.

[0040] (2) As Figure 3 shown, after obtaining a compilable function implementation, the system enters the syntax mutation stage to enhance the syntax diversity of test cases. The present invention constructs a mutation operator library, which contains various C++ syntax features to ensure that test cases cover rich code structures. In this embodiment, the system randomly selects a Lambda expression from the mutation operator library as the mutation operator and constructs the corresponding mutation prompt words. Specifically, the generated mutation prompt word sentence is as follows:

[0041] "Here is a C++ function for you. Please modify it and ensure that the function signature remains unchanged. You can modify the semantics of the function and insert statements. Modify the function to use lambda expressions. Only return the code."

[0042] This prompt word requires the large language model to modify the internal implementation of the function, introduce Lambda expressions, and allow appropriate semantic adjustments while keeping the function signature unchanged. Subsequently, the large language model mutates the code of func3 according to the mutation prompt word and generates the mutated function implementation. The specific mutation results are as follows:

[0043]

[0044]

[0045] Iteratively perform repeated random selection of mutation operators to endow the function with more complex semantics and syntax. After each mutation, it is necessary to verify whether the program can be compiled successfully.

[0046] (3) As Figure 4 shown, after completing the syntax mutation, the method of the present invention enters the combined compilation stage to build a complete test project and generate executable test cases. The system first organizes the file structure of the test project, which includes three main parts: the header folder, the source folder, and the test case folder.

[0047] Header folder (include): Stores all header files, including container.h, prog0.h, prog1.h, prog2.h, and prog3.h. Among them, container.h is mainly used to store type information, while the remaining header files contain function declarations and function templates to ensure that different source files can reference the corresponding interfaces and make function calls.

[0048] Source folders (src_0, src_1, src_2): Contain multiple source files. Each source folder stores prog0.cpp, prog1.cpp, prog2.cpp, and prog3.cpp, which respectively correspond to the header files in the header folder. This structural design allows the system to flexibly select combinations of different source files during compilation, thereby generating different test cases and enhancing the diversity of testing.

[0049] Test case folder (test_case): Contains multiple test cases, and each test case contains a main function as the entry point of the program. In this embodiment, the test case folder stores 10 test case files, and each file is responsible for calling prog3.h and the functions hierarchically called by it to ensure that the generated executable file has complete test logic.

[0050] During the compilation stage, the method of the present invention adopts a random combination compilation strategy, selects source files from different source folders for compilation to form diverse executable test cases. The following is a specific compilation instruction in this embodiment:

[0051] "clang++ -I include src_0 / prog0.cpp src_2 / prog1.cpp src_2 / prog2.cpp src_1 / prog3.cpp test_case1.cpp -O1 -o test_case1 -g"

[0052] The above compilation command selects prog0.cpp, prog1.cpp, prog2.cpp and prog3.cpp from different source folders src_0, src_1 and src_2, and links them with test_case1.cpp, and finally generates the executable file test_case1 with optimization level -O1 and debug information -g. This process continues until 10 different executable files are generated to ensure the diversity of test cases and extensive coverage of debugger features.

[0053] (4) During the test execution phase, the executable file is debugged using a debugging script. The outputs of gdb and lldb are compared. The debugging output is shown below:

[0054] 2089GDB:func9(long long,double):4var241={'t':'short','v':1},LLDB:func9(long long,double):4var241=None

[0055] 2090GDB:func9(long long,double):4var242={'t':'short','v':2},LLDB:func9(long long,double):4var242=None

[0056] 2095GDB:func9(long long,double):12p_1=None,LLDB:func9(long long,double):12p_1={'t':'double','v':'300.000'}

[0057] 2096GDB:unsigned long func11<long long,long double> (long long,longdouble,short):21p_1={'t':'long double','v':'0.000'},LLDB:func11<long long,long double> (long long,long double,short):21p_1={'t':'double','v':'300.000'}

[0058] Inconsistent output between gdb and lldb may come from many aspects, such as inconsistent interpretation of debugging information or inconsistent source code lines where instructions are located. In order to identify bugs from inconsistencies, manual confirmation and identification is required.

Claims

1. A debugger testing method based on a large language model, characterized in that: The specific steps are as follows: Step (1) constructing a random function specification through a statement template and a semantic constraint, and inputting the constructed random function specification as a prompt word into a large language model to generate a corresponding function template; then, the generated function template is compiled and verified, and if the compilation is successful, the function template is retained for subsequent use; if the function template compilation fails, the random function specification is regenerated and iteratively corrected until the compilation is successful; Step (2) performs semantic mutation processing based on the function template successfully compiled in step (1), and writes the mutated function template into the source file; Step (3) compiles the source file with the corresponding header file and test case to generate an executable file; Step (4) verifies the correctness of the debugging information and detects whether there are any bugs by running the generated executable file and testing it on a debugger.

2. The debugger testing method based on a large language model according to claim 1, characterized in that: Step (1) specifically includes the following steps: 1-1) Generate a random function specification. The function specification consists of four basic statement templates. The basic statement templates include declaration, assignment, condition and loop. Randomly select one of them and include a randomly generated function signature, including function name, return type, parameter type and number of parameters. The statement template plus the corresponding template parameter is a complete prompt word statement. The template parameter is the type and name of the variable. 1-2) During the generation process, semantic constraints are introduced, that is, the data types and scopes of variables in the generation process are recorded to maintain the current context file environment; declaration statements will add new variable information to the current context environment, and the remaining statements will select variables of the same type from the current environment as parameters of the statement template; if the prompt word of the next statement to be generated leaves a certain scope, the variables declared in the scope will be deleted from the context environment; 1-3) The constructed random function specification is input into the large language model as a prompt word, and the function template generated by the large language model is compiled and verified; if the compilation is successful, the function template is retained for subsequent use; if the function template compilation fails, the random function specification is regenerated and iteratively corrected; this process will continue to iterate until all function templates can pass the compilation verification.

3. The debugger testing method based on a large language model according to claim 1, characterized in that: Step (2) specifically includes the following steps: 2-1) Extract C++ language features from the C++ official standard document, and build a mutation operator library based on the extracted C++ language features; 2-2) The system randomly selects a mutation operator from the mutation operator library, and combines it with the mutation prompt word template to generate a mutation prompt word and input it into the large language model, so that the large model mutates the function according to the mutation prompt word to generate the mutated function; 2-3) Compile and verify the mutated function; if the compilation is successful, write the mutated function template into the source file; if the compilation fails, select other mutation operators in the mutation operator library and re-enter the large language model to generate a new function and compile and verify again; if successful, continue to select operators for mutation and set the parameter k for the number of mutations, that is, apply the compilation mutation operator to each function.

4. The debugger testing method based on a large language model according to claim 1, characterized in that: Step (3) specifically includes the following steps: 3-1) Build a complete project structure based on the generated functions. The project contains multiple header files, source files and test case files, and stores them in the corresponding source folders, header folders and test case folders; 3-2) Randomly select a source file from the source code folder, and extract the corresponding header file and test case from the header folder and test case folder according to the relevance of the source file; 3-3) Combining the obtained source file with its corresponding header file and test case, and compiling to generate an executable file, and generating the executable file as an executable test case for the debugger.

Citation Information

Cited By

  • Automatic software debugging method, system and device based on large language model

    CN121579328A