Unit test intelligent generation method guided by using structured seed use case

Through the structured seed use case guidance method, combined with static analysis and large-scale model generation technology, the problems of low unit test generation efficiency and insufficient coverage in the existing technology are solved, and efficient and high-quality unit test generation is achieved.

CN119988237AInactive Publication Date: 2025-05-13BEIJING XUANYU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510457455.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is inefficient and insufficient coverage when automated unit test generation, especially when facing C and complex programming languages.

Method used

Using the structured seed use case-oriented method, we parse the tested functions through static analysis, construct seed use cases of context and preset structures, use big models to generate test cases, and convert them into standard test code through rule-based methods to optimize test coverage.

Benefits of technology

It improves the efficiency of automatic generation of unit tests, enhances test coverage and quality, can meet requirements in actual development, and achieve a high level of execution pass rate and coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988237A_ABST
    Figure CN119988237A_ABST
Patent Text Reader

Abstract

A unit test intelligent generation method guided by using a structured seed case belongs to the technical field of software testing, and comprises the following steps: analyzing a tested function, constructing a context of the tested function, and constructing a seed case of a preset structure according to the tested function; according to the context of the tested function, the code block and the seed case, a cue word is constructed, the cue word is input into the large model to generate a test case of the preset structure, and the test case is converted into a test code of a preset standard; and executing the test code to obtain a test coverage rate, and determining an optimization demand according to the test coverage rate condition. According to the method, the interface data of the tested function is analyzed, the structured seed case is constructed for the tested function and used for guiding the large model to generate the structured test case, the structured test case is converted into the test code by using a rule-based method, the problem that the error probability is high when the test code is directly generated by the large model is solved, and the test efficiency is improved. And the expression capability of the large model in the unit test is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a unit test intelligent generation method guided by structured seed use cases, belonging to the technical field of software testing. Background Art

[0002] Unit testing plays a vital role in ensuring software correctness. It helps developers identify potential problems in the early stages of program development, effectively reducing the occurrence of defects and thus reducing development costs. However, manually creating and maintaining unit test cases can be laborious and time-consuming.

[0003] To address this problem, researchers have proposed various methods to automate the unit test generation process. Traditional unit testing tools adopt search-based or constraint-based techniques. Although these methods can generate test cases to enhance software coverage, they have several key limitations in terms of test readability, scenario coverage, and assertion quality, which increases the complexity of maintenance. Deep learning-based methods can learn from real-world focus methods and generate test cases that are very similar to the developer's writing style.

[0004] Although existing Large Language Model (LLM)-based methods can sometimes generate unit tests that are well explained, easy to understand, and have a certain degree of accuracy and coverage, there are still obvious limitations and defects. First, existing LLM-based test case generation works are all targeted at specific languages, such as Java and Python. However, there is a lack of dedicated support for the C programming language, which is widely used in key real-world applications such as operating systems and network devices. The performance of LLM on different language tasks varies greatly, which makes it difficult for these methods to be extended to the C language. Second, when faced with development tasks in real scenarios, existing works still face problems such as low compilation success rate, low execution pass rate, and low coverage of generated test cases, which are difficult to meet actual needs. The fundamental limitations of existing LLM-driven unit test generation methods mainly come from their reliance on LLM for code generation. These methods are hindered by the model's ability to generate programs in specific programming languages. Specifically, when these languages ​​exhibit complex characteristics and test cases require more advanced functions, the constraints become obvious. Summary of the invention

[0005] The technical problem solved by the present invention is: to overcome the shortcomings of the prior art, provide a method for intelligently generating unit tests guided by structured seed use cases, solve the current problems of low efficiency and insufficient coverage of automated unit test generation, improve the efficiency of automatic unit test generation, and enhance the testing effect.

[0006] The technical solution of the present invention is: in the first aspect, a method for intelligently generating unit tests guided by structured seed use cases, comprising: Parse the function under test, build the context of the function under test, and build a seed use case with a preset structure based on the function under test; Construct prompt words according to the context and code blocks of the preset seed test case and the function under test, input the prompt words into the large model to generate the test case of the preset structure, and convert the test case into the test code of the preset standard; Execute the test code to obtain test coverage, determine optimization requirements based on the test coverage, and if the coverage of the tested function is lower than a preset value, adjust the prompt word to generate the test case again.

[0007] Furthermore, the method for parsing the tested function is a static analysis method, including: extracting the tested function dependencies, code blocks and interface data of the tested function, and organizing these three parts into the tested function context; the dependencies include the code segments required to separately compile the tested function, including type declarations, global variable definitions, preprocessor instructions and function declarations involved in the implementation of the tested function; the interface data includes all data exposed by the tested function and related to the test, including formal parameters, global variables, return values, type definition data and functions called in the Focal method.

[0008] Furthermore, the static analysis method includes: Extract the macro definition pi, external type shape, global variable area and function calarea called by the target method of the tested function; Split the tested function code segment from the tested function C file and combine it with the dependencies to create a standalone compilable program; Analyze the parameter list, return value, global variables and their type information of the tested function; Analyze the called function calarea and collect interface data.

[0009] Furthermore, the preset structure includes test input, test output and stub function; the test input and test output include expressions and values, indicating test data that need to be assigned in the test; the stub function includes function name, expression and value, indicating the assignment of a certain test data in the called function.

[0010] Furthermore, the preset standard test code is a C test code.

[0011] Furthermore, the preset value is 80%.

[0012] In a second aspect, a unit test intelligent generation system guided by structured seed use cases includes: The first module parses the function under test, builds the context of the function under test, and builds a seed use case with a preset structure based on the function under test; The second module constructs prompt words according to the context and code blocks of the preset seed case and the function under test, inputs the prompt words into the large model to generate the test case of the preset structure, and converts the test case into the test code of the preset standard; The third module executes the test code to obtain the test coverage, and determines the optimization requirements according to the test coverage. If the coverage of the tested function is lower than the preset value, the prompt word is adjusted to generate the test case again.

[0013] Furthermore, the method for parsing the tested function is a static analysis method, including: extracting the tested function dependency, code block and interface data of the tested function, and organizing these three parts into the tested function context; the dependency includes the code segment required for separately compiling the tested function, including type declarations, global variable definitions, preprocessor instructions and function declarations involved in the implementation of the tested function; the interface data includes all data exposed by the tested function and related to the test, including formal parameters, global variables, return values, type definition data and functions called in the focal method; The static analysis method comprises: Extract the macro definition pi, external type shape, global variable area and function calarea called by the target method of the tested function; Split the tested function code segment from the tested function C file and combine it with the dependencies to create a standalone compilable program; Analyze the parameter list, return value, global variables and their type information of the tested function; Analyze the called function calarea and collect interface data; The preset structure includes test input, test output and pile function; the test input and test output include expressions and values, which represent test data that need to be assigned in the test; the pile function includes function name, expression and value, which represents the assignment of a certain test data in the called function; The preset standard test code is the C test code; The preset value is 80%.

[0014] In a third aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for intelligently generating unit tests guided by structured seed use cases are implemented.

[0015] In a fourth aspect, a device for intelligently generating unit tests guided by structured seed use cases comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of a method for intelligently generating unit tests guided by structured seed use cases when executing the computer program.

[0016] The advantages of the present invention compared with the prior art are: The present invention uses an efficient unit test intelligent generation method based on structured seed case guidance. By analyzing the interface data of the tested function, a structured seed case is constructed for the tested function to guide the large model to generate structured test cases. The structured test cases are converted into test codes using a rule-based method, which solves the problem of high error probability when the large model directly generates test codes and improves the performance of the large model in unit testing. The present invention can guide LLM to generate high-quality unit test cases. This method has been verified on multiple cases, achieving an execution pass rate of 97.73%, and the line coverage and branch coverage rates are 85.32% and 75.26% respectively, which can meet the requirements of unit testing in actual development. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings: Figure 1 A schematic diagram of a flow chart of a method for intelligently generating unit test cases guided by a structured seed case provided in an embodiment of the present application; Figure 2 An example diagram of constructing a context for a function under test provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] In order to better understand the above technical scheme, the technical scheme of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical scheme of the present invention, rather than limitations on the technical scheme of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0019] The following is a further detailed description of a method for intelligently generating unit tests guided by structured seed use cases provided by an embodiment of the present invention in conjunction with the accompanying drawings of the specification. The specific implementation method may include: Parse the function under test, build the context of the function under test, and build a seed use case with a preset structure based on the function under test; Construct prompt words according to the context, code blocks and seed cases of the function under test, input the prompt words into the large model to generate the test cases of the preset structure, and convert the test cases into test codes of preset standards; Execute the test code to obtain test coverage, determine optimization requirements based on the test coverage, and if the coverage of the tested function is lower than a preset value, adjust the prompt word to generate the test case again.

[0020] In the solution provided by the embodiment of the present invention, a structured seed test case is constructed for each tested function to guide LLM to generate high-quality unit test cases, thereby improving the unit test efficiency.

[0021] This method is mainly divided into three parts: 1) constructing the context information of the function under test and generating structured seed cases for the interface; 2) test case generation and rule-based test code generation; 3) test execution and feedback-based test optimization. The main steps are as follows: (1) Perform static analysis on the function under test and build a context database Use static analysis methods to parse the tested function, extract the tested function dependencies, code blocks and interface data of the tested function, and organize these three parts into a tested function context database file.

[0022] Define test cases as structured patterns. Each test case consists of test input, test output, and stub functions. Input and output consist of expressions and values, indicating test data that needs to be assigned in the test. Stub functions consist of function names, expressions, and values, indicating the assignment of values ​​to certain test data in the called function.

[0023] (2) Test case generation and rule-based test code generation The prompt words for generating test cases consist of the context of the function under test, the code block, and the seed case. The prompt words use a one-shot prompt method, using the seed case as an example to demonstrate the standard test case format to the LLM.

[0024] This method uses a rule-based code conversion method to convert structured test cases into C test code. The test code is divided into three parts: data preparation, test execution, and result verification.

[0025] (3) Test execution and feedback optimization During the test execution phase, the test code is first compiled and run, and the optimization requirements are determined based on the test coverage. If the coverage of the tested function is less than 80%, an optimization is performed to collect relevant information to guide LLM to generate test cases again.

[0026] The following is a further detailed description of a method for intelligently generating unit test cases guided by structured seed test cases provided in an embodiment of the present application in conjunction with the accompanying drawings. The specific implementation of the method may include the following steps (the method flow is as follows: Figure 1 shown): Step 1: Context Building LLM completes specific tasks based on prompt words. Good prompt words can help LLM focus on key issues and reduce interference factors. Therefore, extracting the context of the function under test and providing relevant information of the function under test to LLM can help LLM generate unit tests better.

[0027] Use static analysis methods to parse the tested function, extract the tested function dependencies, code blocks and interface data of the tested function, and organize these three parts into a tested function context database file.

[0028] Dependencies refer to the code segments required to compile the function under test separately. These include type declarations, global variable definitions, preprocessor directives, and function declarations involved in the implementation of the function under test. Once all dependencies are obtained, the function under test can be compiled independently.

[0029] Interface data refers to all test-related data exposed by the tested function, including formal parameters, global variables, return values, type definition data, and functions called in the Focal method.

[0030] exist Figure 2 In the example of context building, we show that first, we extract the macro definition pi, the external type shape, the global variable area, and the function calarea called by the target method. Then, we split the Focal method code segment from the entire C file and combine it with the dependencies to create an independent compilable program. Finally, we analyze the parameter list, return value, global variables, and their type information of the circle function. A similar analysis is performed on the called function calarea to collect the relevant interface data.

[0031] After obtaining the context of the function under test, a structured seed test case is constructed using an interface-oriented approach. The pattern of the structured test case is defined as follows: TestCase={Inputs,outputs,Stubs} Inputs = { (ExprInputi,Valuei), (ExprInput2,Value2), … } Outputs = { (Exproutputi,Valuei), (Exproutput2,Value2), … } Stubs = { (StubFunctioni,ExprStubi,Valuesi), (StubFunction2,ExprStub2,Values2) … } Each test case consists of inputs, outputs, and stubs, which represent test input, test output, and stub functions respectively. Inputs and outputs consist of expressions and values, indicating the test data that needs to be assigned in the test. Stub functions consist of function names, expressions, and values, indicating the assignment of a test data in the called function.

[0032] Figure 2 Example of structured seeds generated by the circle function in : { "cases":[{ "inputs":[{ "expr": "radius", "value": 0.0 },{ "expr": "shape->area", "value": 0.0 },{ "expr": "shape->perimeter", "value": 0.0 },{ "expr": "area", "value": 0 }], "stubs":[{ "funcName": "calArea", "expr": "returnvalue", "value": 0.0 }], "outputs":[{ "expr": "shape->area", "value": 0.0 },{ "expr": "shape->perimeter", "value": 0.0 }] }] } First, scan the interface data of the parameters radius, shape, and global variable area. Use this data as test input. From the callee data in the interface, you can see that the focus method circle calls the calArea function. Based on the analysis of calArea, it is found that only its return value affects circle, so only the return value example of calArea needs to be included in the stubs. Finally, it is found that the members of shape are modified during the test. Since this parameter is a pointer type, output examples of the area and perimeter members of shape are provided. In addition, a default value is generated for each expression based on its type. These default values ​​can help LLM better understand the data type of the assignment target.

[0033] Step 2: Unit test generation The prompt words for generating test cases are: Please generate some test cases for the following C function. Iwillgive you a test case example, please continue to generate. Meet the following requirements. 1.Try to cover all branches. 2.Use the stub function to simulate the return value of the called function. I will provide you with all the data that the called function may change. 3.Each test case outputs in a standard JsON format. {{ context}} {{ focal method}} {{ seed case}} The prompt consists of the context of the function under test, a code block, and a seed case. The prompt explicitly states that mock data needs to be generated for the called function. In the seed case, data that may be changed by the called function is provided, allowing the LLM to determine the assignment based on the test branch. A one-shot prompting approach is adopted to demonstrate the standard test case format to the LLM using the seed case as an example. The seed case explicitly identifies the objects that must be assigned during testing and the expected test output format. The key to this strategy is to provide the LLM with a clear template to ensure that the generated test cases accurately match the expected input and output formats. In addition, by providing the model with sufficiently detailed context and examples, the generated test cases show higher consistency and reliability. This approach improves the accuracy of test data generated by the LLM and minimizes ambiguous or uncertain results.

[0034] Once the structured test cases are generated, a rule-based code transformation approach is used to convert these cases into C test code. Each test case corresponds to two main components in the test code: the stub function and the test case function. The test case function itself consists of three parts: data preparation, function execution, and result verification. The structured seed example generated by the circle function described above shows the contents of the structured test case on the left, while the generated test code is shown on the right. For example, lines 2–4 of the test code show the calArea stub function generated from the stubs section of the structured test case. This function returns 3.14, simulating the stub behavior specified in the test case. Lines 12–15 in the test code set the input variables as shown in the input section of the structured test case. Lines 19–20 represent the result verification section, which corresponds to the output section of the structured test case, where the ASSERT_FLOAT_EQ macro checks whether shape->area is 8.14 and shape->perimeter is 6.18 after the function is executed.

[0035] Step 3: Test execution and feedback optimization In the test run and optimization phase, first compile and run the test, collect the compilation and runtime results, as well as the line and branch coverage metrics. Based on the results of the coverage analysis, determine the need for optimization. Specifically, only focus methods with coverage below 80% are optimized once to achieve good test efficiency and reduce costs. Then, collect information about uncovered branches and analyze whether true or false conditions lead to the lack of coverage. Integrate the uncovered branch code and related information, and provide feedback to LLM to guide it to regenerate test cases targeting the branch. The optimization tips are as follows: Well done. But the following branches are not covered. I willprovideyou with the uncovered conditions. Please generate testcases for theuncovered branches. The test cases you designneed to be able to reach thisbranch and cover this branch. Donot output test cases that are repeated with the previous ones. 1. if (a>b): true condition uncover 2. if (x==1): false condition uncover In the prompt, the LLM is required not to generate duplicate test cases and is allowed to independently determine the path conditions that must be met to reach the branch. If the generated test cases do not improve the coverage, they are discarded; otherwise, the extended test cases are added to the test suite.

[0036] The present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer executes Figure 1 The method described.

[0037] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.

[0038] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0039] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0040] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0041] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

[0042] The contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.

Claims

1. A method for intelligently generating unit tests guided by structured seed cases, characterized in that: include: Parse the function under test, build the context of the function under test, and build a seed use case with a preset structure based on the function under test; Construct prompt words according to the context and code blocks of the preset seed test case and the function under test, input the prompt words into the large model to generate the test case of the preset structure, and convert the test case into the test code of the preset standard; Execute the test code to obtain test coverage, determine optimization requirements based on the test coverage, and if the coverage of the tested function is lower than a preset value, adjust the prompt word to generate the test case again.

2. The method for intelligently generating unit tests guided by structured seed use cases according to claim 1, characterized in that: The method for parsing the tested function is a static analysis method, including: extracting the tested function dependency, code block and interface data of the tested function, and organizing these three parts into the tested function context; the dependency includes the code segment required for separately compiling the tested function, including type declarations, global variable definitions, preprocessor instructions and function declarations involved in the implementation of the tested function; the interface data includes all data exposed by the tested function and related to the test, including formal parameters, global variables, return values, type definition data and functions called in the Focal method.

3. The method for intelligently generating unit tests guided by structured seed use cases according to claim 2, characterized in that: The static analysis method comprises: Extract the macro definition pi, external type shape, global variable area and function calarea called by the target method of the tested function; Split the tested function code segment from the tested function C file and combine it with the dependencies to create a standalone compilable program; Analyze the parameter list, return value, global variables and their type information of the tested function; Analyze the called function calarea and collect interface data.

4. The method for intelligently generating unit tests guided by structured seed use cases according to claim 1, characterized in that: The preset structure includes test input, test output and stub function; the test input and test output include expressions and values, indicating test data that need to be assigned in the test; the stub function includes function name, expression and value, indicating the assignment of a certain test data in the called function.

5. The method for intelligently generating unit tests guided by structured seed use cases according to claim 1, characterized in that: The preset standard test code is the C test code.

6. The method for intelligently generating unit tests guided by structured seed use cases according to claim 1, characterized in that: The preset value is 80%.

7. A unit test intelligent generation system guided by structured seed cases, characterized in that: include: The first module parses the function under test, builds the context of the function under test, and builds a seed use case with a preset structure based on the function under test; The second module constructs prompt words according to the context and code blocks of the preset seed case and the function under test, inputs the prompt words into the large model to generate the test case of the preset structure, and converts the test case into the test code of the preset standard; The third module executes the test code to obtain the test coverage, and determines the optimization requirements according to the test coverage. If the coverage of the tested function is lower than the preset value, the prompt word is adjusted to generate the test case again.

8. The unit test intelligent generation system guided by structured seed use cases according to claim 7, characterized in that: The method for parsing the tested function is a static analysis method, including: extracting the tested function dependency, code block and interface data of the tested function, and organizing these three parts into the tested function context; the dependency includes the code segment required for separately compiling the tested function, including type declarations, global variable definitions, preprocessor instructions and function declarations involved in the implementation of the tested function; the interface data includes all data exposed by the tested function and related to the test, including formal parameters, global variables, return values, type definition data and functions called in the Focal method.

9. The unit test intelligent generation system guided by structured seed use cases according to claim 8, characterized in that: The static analysis method comprises: Extract the macro definition pi, external type shape, global variable area and function calarea called by the target method of the tested function; Split the tested function code segment from the tested function C file and combine it with the dependencies to create a standalone compilable program; Analyze the parameter list, return value, global variables and their type information of the tested function; Analyze the called function calarea and collect interface data.

10. The unit test intelligent generation system guided by structured seed use cases according to claim 7, characterized in that: The preset structure includes test input, test output and stub function; the test input and test output include expressions and values, indicating test data that need to be assigned in the test; the stub function includes function name, expression and value, indicating the assignment of a certain test data in the called function.

11. The unit test intelligent generation system guided by structured seed use cases according to claim 7, characterized in that: The preset standard test code is the C test code.

12. The unit test intelligent generation system guided by structured seed use cases according to claim 7, characterized in that: The preset value is 80%.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

14. A unit test intelligent generation device guided by structured seed cases, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Test method of C code software and readable storage medium

    CN113360410A

  • Fuzzy testing method for security business background system interface

    CN118349444A

  • Unit test case generation system based on large language model

    CN119105965A

  • System kernel fuzzy test seed generation method and system based on large language model

    CN119645882A

Cited By

  • Test case generation method and device, equipment, storage medium and program product

    CN121636361A