Unit test instant automatic updating method based on large language model
By employing a collaborative approach combining large language models and static analysis tools, unit tests are automatically updated, resolving the inconsistency between test cases and production code in existing technologies. This enables efficient test code updates and verification, improving test correctness and coverage.
Patent Information
- Application Number
- CN202511165251.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-28
AI Technical Summary
Existing unit testing methods struggle to keep up with frequent software system iterations and code changes, leading to inconsistencies between test cases and production code, resulting in compilation errors and runtime exceptions. Furthermore, existing technologies rely on manual maintenance, which is inefficient and prone to delays in fixing issues, making them unsuitable for complex scenarios and for verifying test code output from large language models.
We employ a collaborative approach based on large language models and static analysis tools. Through context collection, natural language prompts, iterative repair, and a strategy of minimizing modifications, we automatically update unit tests, including a context collection module, a prompt generation module, a test generation module, an execution and verification module, and an iterative repair module, to ensure the correctness and coverage of the test code.
It enables real-time automatic updates of unit tests, improves the pass rate and coverage of test code, enhances test quality, can cope with test obsolescence issues caused by production code changes, and reduces manual intervention and errors.
Smart Images

Figure CN121029607A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the fields of deep learning and software engineering, and in particular to a unit test instant automatic updating method based on a large language model. BACKGROUND
[0002] In modern software development processes, unit testing is an important foundation for ensuring software quality and system stability. By verifying the functions of each component in the code base, unit testing helps to identify potential regression problems in the early stages of the development cycle, significantly reducing the risk of software defects propagating to subsequent stages. However, as software systems continue to evolve and iterate frequently, existing unit tests may not be updated in a timely manner, leading to deviations from the structure or behavior of the production code, and thus causing compilation errors, runtime exceptions, and other problems. Outdated test cases cannot effectively cover the newly added or changed code logic, which may result in a decrease in test coverage and the failure to identify vulnerabilities in a timely manner. Therefore, ensuring the consistency and timeliness between test cases and the code under test is crucial to maintaining test effectiveness and overall software quality. In this regard, current mainstream practices still rely on manual maintenance of test code, which is not only inefficient, but also prone to problems such as repair lag, update omission, and loss of test intent in actual operation.
[0003] Recently, scholars have explored the use of large language models to automatically repair failed unit tests to improve the efficiency and automation of test maintenance. However, these techniques mainly focus on handling test case failures caused by explicit changes such as code interface signature changes and syntax structure adjustments, and still show obvious deficiencies in complex scenarios that require test enhancement to cover newly added business logic or behavior semantic changes. Existing methods generally rely on rule-based context extraction strategies to obtain the information needed to generate tests from the original code, but such strategies are often difficult to adapt to diverse code structures, difficult to extend to complex code changes, and have obvious limitations in handling large or deep changes. Moreover, although most tests generated by large language models are grammatically correct, there are still a large number of compilation or execution errors, and existing methods generally lack mechanisms for verifying and optimizing test code output by large language models, which requires manual debugging and limits the effectiveness of automated test updates. SUMMARY
[0004] The present application aims to address the shortcomings of existing test update techniques and provides an innovative solution for implementing unit test instant automatic updating based on a large language model according to code changes.
[0005] The purpose of the present application is achieved by the following technical solution: a unit test instant automatic updating method based on a large language model, comprising the following steps: (1) Collecting context information related to test updates using a context collector based on a large language model and a static analysis tool; (2) Building natural language prompts for guiding the large language model to generate test update code according to the obtained context information; (3) Guiding the large language model to generate test update code according to the obtained prompts, preparing for the execution of the test update code, including importing the required dependencies for new additions; (4) Compiling and executing the test update code; (5) If there is a compilation failure or runtime test failure, analyzing the error cause, collecting error-related context information, and providing feedback to the large language model to guide test case repair; (6) Guiding the large language model to repair the failed test update code, re-executing step (4) to enter a new round of iteration until the test passes or the iteration number exceeds the set upper limit; (7) If the iteration number exceeds the set upper limit, instructing the large language model to generate a basic test using a minimization conservative modification strategy.
[0006] Further, the step (1) is specifically: (1.1) Input the change of the focus method and the corresponding old test code, use the large language model to analyze the change of the focus method, use the thought chain structure to build prompts to enhance reasoning ability, automatically identify the required related methods and classes, and output the identification results in a pre-defined data structure format; (1.2) Use the static analysis tool to locate the definitions of related methods and classes, and extract related code snippets; (1.3) Use the large language model again to filter out irrelevant content for test updates; (1.4) Collect test class fields, i.e. all class-level variables defined inside the test class, as supplements.
[0007] Further, the step (2) is specifically: a. Build a structured prompt template, which includes the change of the focus method, the old test code, and the context information related to the focus method; b. Set the role of the large language model to the corresponding language expert role; c. Provide a multi-step semantic guidance prompt information to guide the large language model to test updates to the large language model; d. Further instruct the large language model to identify new dependencies and only generate corresponding import statements for new dependencies to avoid repeated imports.
[0008] Further, the multi-step semantic guidance prompt information includes updating the original test method based on the change of the focus method, verifying whether the updated test logic covers the new function or behavior change, generating corresponding new test logic if the focus method introduces new functions, generating simulation objects or using default values if the focus method adds new parameters, and only repairing the original test method if the core function of the focus method does not change.
[0009] Further, the step (5) is specifically: (5.1) If there is a compilation error, the error type and error location information are extracted from the compiler error information, and the context information required for error repair is located using a static analysis tool, including error-related definitions and code fragments; then a corresponding repair prompt is generated to help the large language model better repair the test; (5.2) If there is a runtime test error, the error assertion is located from the test report and the context information is extracted, including the expected value and actual value of the assertion, and the line number where the failed assertion is located; then a prompt containing the corresponding assertion code and context information is constructed to guide the large language model to repair the test.
[0010] The application also provides a large language model-based unit test instant automatic updating system for implementing the above method, comprising: A context collection module for collecting context information related to test updates using a context collector based on a large language model and a static analysis tool; A prompt generation module for constructing natural language prompts for guiding the large language model to generate test update code according to the context information; A test generation module for guiding the large language model to generate test update code according to the prompts; An execution and verification module for compiling and executing the test update code; if there is a compilation failure or a runtime test failure, the error cause is analyzed, error-related context information is collected, and feedback is provided to the large language model to guide test case repair; An iterative repair module for guiding the large language model to repair the failed test update code, re-executing the execution and verification module to enter a new round of iteration until the test passes or the number of iterations exceeds the set upper limit; A minimal modification module for instructing the large language model to generate basic tests using a minimal conservative modification strategy.
[0011] The application also provides an electronic device comprising a memory and a processor, wherein the memory is coupled to the processor; the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned large language model-based unit test instant automatic updating method.
[0012] The application further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the method.
[0013] The application has the beneficial effects that: the large language model is used to fully and accurately analyze and collect context information, guide the test update process, and improve the correctness of the generated test through the introduction of an iterative optimization mechanism, which can not only repair invalid tests, but also enhance test quality by verifying new logic to improve the pass rate and coverage of test code, and better cope with the problem of test obsolescence caused by production code changes. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0015] Figure 1 is a general diagram of the method of the application; Figure 2 is a flowchart of the method of the application, which collects test update related components through a large language model and a static analysis tool to ensure the correctness, relevance and executability of the test update generated by the large language model. DETAILED DESCRIPTION
[0016] The exemplary embodiments will be described in detail herein with reference to the drawings. Unless otherwise indicated, the same numbers on different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the application. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the application as detailed in the appended claims.
[0017] The application realizes the instant automatic update of unit testing based on a large language model through a large language model and a static program analysis tool (such as a language server protocol-based tool).
[0018] As shown in Figure 1 The embodiment of the application provides a method for instant automatic update of unit testing based on a large language model, which comprises the following steps: (1) Collecting context information related to test updates using a context collector based on a large language model and a static analysis tool, which is adjusted from existing technology, and cooperates with static analysis tools and large language models to avoid redundant context compared to using a large language model alone, and to enhance the understanding of semantics and change logic compared to using a static analysis tool alone. The context information includes test update related components and test class fields obtained as shown in Figure 2 The specific steps are as follows: (1.1) Input the change of the focus method and the old test code, use the large language model to analyze the change of the focus method, use the thought chain structure to construct the prompt to enhance the reasoning ability, and automatically identify the required related methods and classes, and require the identification results to be output in JSON data structure format; (1.2) Use a static analysis tool (such as a language server protocol-based tool, Language Server) to locate the definitions of related methods and classes, and extract related code fragments by parsing the code into a syntax analysis tree; (1.3) Use the large language model again to filter out irrelevant content to test updates; (1.4) Collect test class fields, i.e. all class-level variables defined inside the test class, as a supplement, which guides the large language model to avoid unnecessary data simulation of defined variables, thereby greatly reducing errors caused by additional test logic.
[0019] (2) According to the obtained context information, construct a natural language prompt for guiding the large language model to generate test update code, and design a structured prompt template in Markdown format, which includes the change of the focus method, the original test method, and the collected context information. The prompt template first sets the role of the large model as the corresponding language expert role (set to Java expert in this embodiment), and then guides the model to complete the update of the original test method and introduce necessary dependencies (such as external classes, tool libraries, etc. that may be added after the test code is updated). For the part of updating the original test method, the prompt provides five clear and operable instructions to help the large model repair and optimize the test method according to the adjustment of the core method; further instruct the large language model to identify the added dependencies, and only generate the corresponding import statement for the added dependencies to avoid repeated imports. The prompt information includes updating the original test method based on the change of the focus method, verifying whether the updated test logic covers the new function or behavior change, generating corresponding new test logic if the focus method introduces new functions, generating simulation objects or using default values if the focus method adds parameters, and only repairing the original test method if the core function of the focus method does not change.
[0020] (3) According to the obtained prompt, guide the large language model to generate test update code, parse the generated Java code, extract the updated test method and import statement, and integrate the two components into the test class, so as to ensure that the generated test code is complete and executable; (4) Compile and execute the test update code; (5) If there is a compilation failure or runtime test failure, analyze the error reason and collect error-related context information, and provide feedback to the large language model to guide the test case repair, including the following sub-steps: (5.1) If there is a compilation error, extract error type, error location, etc. From the compiler error information, and use static analysis tools to locate the context information required to fix the error, such as error-related definitions and code snippets, etc. Then generate the corresponding repair prompt to help the large language model better repair the test.
[0021] (5.2) If there is a runtime test error, locate the error assertion from the test report and extract the context information, including the expected value and actual value of the assertion, and the line number of the failed assertion. Then build a prompt containing the corresponding assertion code and context information to guide the large language model to repair the test.
[0022] (6) Guide the large language model to repair the failed test update code, and re-execute step (4) to enter a new round of iteration until the test passes or the iteration number exceeds the set upper limit (the iteration number set upper limit is generally set to 2, which can cover most common error repairs, and avoid the model entering an error loop); (7) If the iteration number exceeds the set upper limit, instruct the large language model to generate a basic test using a minimal conservative modification strategy (such as assigning a default value to a new parameter), which ensures that the test method can at least be successfully compiled and run.
[0023] Example 1 The present application carries out two evaluations on the proposed large language model-based unit test instant automatic update method: 1) comparison with the existing baseline scheme in terms of compilation, test pass rate and coverage, and 2) verification of the robustness of the method to different large language models.
[0024] The specific implementation and settings of the method, as well as the specific experiments of comparison with the baseline scheme and exploration of robustness, are described below.
[0025] The method realizes and sets: in view of the existing test repair data set (such as TARBENCH), the main focus is on repairing failed test cases, ignoring the limitation of ignoring those test cases that can pass the test but still need to enhance the test logic, and the present application constructs and uses a new data set. The present application collects open source Java projects by calling the GitHub open interface, and constructs a new data set named UPDATES4J. In order to ensure that the test cases can be successfully compiled and generate test coverage report, the projects supporting Maven construction method and compatible with JaCoCo code coverage tool are selected. The constructed UPDATES4J data set covers the actual scene of the cooperative evolution of production code and test code in the real development process, not only contains the repair samples of failed test cases, but also covers the test logic of the test cases that do not cause test failure but still need to be enhanced. The data set contains 165 repaired test cases and 30 enhanced un-repaired test cases, which is close to the distribution in actual development in terms of data proportion. In addition, UPDATES4J also contains 92 method signature changes and 103 method internal logic modifications, so as to more comprehensively reflect the influence of code evolution on test cases. In terms of model calling, the present application uses LangChain framework combined with large language model interface to construct prompt template. Specifically, the large language model Llama-3.3-70B-Instruct is deployed on the local by vLLM inference engine, and GPT-4.1 and DeepSeek-V3 are called through their respective official APIs. The temperature parameter of the language model is uniformly set to 0.1, each model is evaluated for three rounds, and the result with the best performance is retained for final analysis. The static code analysis tool used in the method is Language Server, and the maximum number of repair iteration is set to 2.
[0026] Performance comparison with baseline scheme: Table 1 shows the comparison of the present application with the baseline scheme, using UPDATES4J as the test data set, and the large language model using DeepSeek-V3 model, and the baseline scheme is SYNTER and NAIVELLM also using large language model, using compilation pass rate, test pass rate, branch coverage and line coverage to evaluate performance. Specifically, the coverage index is calculated based on the average line or branch coverage of all test cases in the data set, regardless of whether the test result passes or not. The experimental results show that the present application achieves a compilation pass rate of 94.4%, a test pass rate of 84.6%, a branch coverage of 45.8%, and a line coverage of 68.3%, which significantly surpasses NAIVELLM and SYNTER in each index. This reflects the practical value of the method.
[0027] Table 1: Performance comparison of the present application and the baseline scheme Verify the robustness of the method to different large language models: Table 2 shows the robustness of the application using different large language models, using UPDATES4J as the test dataset, and using the method to evaluate the compilation pass rate and test pass rate on Llama-3.3-70B, GPT-4.1, and DeepSeek-V3 respectively. The experimental results show that the method using different large language models has a high compilation pass rate and test pass rate, indicating that the application has good robustness to different large language models, and using stronger large language models may bring certain performance improvement.
[0028] Table 2: Robustness of the application using different large language models The application also provides a large language model-based unit test instant automatic updating system for implementing the above method, comprising: A context collection module for collecting context information related to test updates using a context collector based on a large language model and a static analysis tool; A prompt generation module for constructing natural language prompts for guiding the large language model to generate test update code according to the context information; A test generation module for guiding the large language model to generate test update code according to the prompts; An execution and verification module for compiling and executing the test update code; if there is a compilation failure or runtime test failure, analyze the error cause, collect error-related context information, and provide feedback to the large language model to guide test case repair; An iterative repair module for guiding the large language model to repair the failed test update code, re-executing the execution and verification module to enter a new round of iteration until the test passes or the number of iterations exceeds the set upper limit; A minimal modification module for instructing the large language model to generate basic tests using a minimal conservative modification strategy.
[0029] The application also provides an electronic device comprising a memory and a processor, the memory being coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to realize the large language model-based unit test instant automatic updating method.
[0030] The application also provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to realize the large language model-based unit test instant automatic updating method.
[0031] The computer readable storage medium can be an internal storage unit of any of the aforementioned data processing capable devices, such as a hard disk or a memory. The computer readable storage medium can also be any of the aforementioned data processing capable devices, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can also include both an internal storage unit of any of the aforementioned data processing capable devices and an external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the aforementioned data processing capable devices, and can also be used to temporarily store data that has been output or will be output.
[0032] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and all of them should be covered in the scope of the claims of the present application.
Claims
1. A method for real-time automatic updating of unit tests based on a large language model, characterized in that, Includes the following steps: (1) Use a context collector based on a large language model and static analysis tools to collect context information related to test updates; (2) Construct natural language prompts based on the obtained context information to guide the large language model in generating test update code; (3) Based on the prompts obtained, guide the large language model to generate test update code and prepare for the execution of test update code, including importing the newly added required dependencies; (4) Compile and execute the test update code; (5) If there is a compilation failure or a runtime test failure, analyze the cause of the error, collect error-related context information, and provide feedback to the large language model to guide the repair of test cases; (6) Guide the large language model to fix the failed test update code, and re-execute step (4) to enter a new round of iteration until the test passes or the number of iterations exceeds the set limit; (7) If the number of iterations exceeds the set upper limit, the large language model is instructed to adopt a minimal conservative modification strategy to generate basic tests.
2. The method for real-time automatic updating of unit tests based on a large language model according to claim 1, characterized in that, The specific steps (1) are as follows: (1.1) Input the changes to the focus method and the corresponding old test code, analyze the changes to the focus method using a large language model, construct hints using a thought chain structure to enhance reasoning ability, automatically identify the relevant methods and classes required, and output the identification results in a predefined data structure format; (1.2) Use static analysis tools to locate the definitions of relevant methods and classes, and extract relevant code snippets; (1.3) Use the large language model again to filter out content that is irrelevant to the test update; (1.4) Collect all class-level variables defined in the test class as supplementary data.
3. The method for real-time automatic updating of unit tests based on a large language model according to claim 1, characterized in that, Step (2) specifically involves: a. Build a structured hint template that includes changes to the focus method, old test code, and contextual information related to the focus method; b. Assign the large language model role to the corresponding language expert role; c. Provide the large language model with multi-step semantic guidance prompts to guide the testing and updating of the large language model; d. Further instruct the large language model to identify new dependencies and generate corresponding import statements only for new dependencies to avoid duplicate imports.
4. The method for real-time automatic updating of unit tests based on a large language model according to claim 3, characterized in that, The multi-step semantic guidance prompts include updating the original test method based on changes to the focus method, verifying whether the updated test logic covers new functions or behavioral changes, generating corresponding new test logic if the focus method introduces new functions, generating mock objects or using default values if the focus method adds new parameters, and fixing only the original test method if the core function of the focus method has not changed.
5. The method for real-time automatic updating of unit tests based on a large language model according to claim 1, characterized in that, Step (5) specifically involves: (5.1) If there is a compilation error, extract the error type and error location information from the compiler error information, and use static analysis tools to locate the context information required to fix the error, including error-related definitions and code snippets; Then, corresponding repair prompts are generated to help the large language model to better perform repair tests; (5.2) If there is a test error during runtime, locate the erroneous assertion from the test report and extract its context information, including the expected value and actual value of the assertion, as well as the line number where the failed assertion is located; Then, a hint containing the corresponding assertion code and context information is built to guide the large language model repair test.
6. A real-time automatic update system for unit tests based on a large language model, implementing the method as described in claim 1, characterized in that, include: The context collection module is used to collect context information related to test updates using a context collector based on a large language model and static analysis tools; The suggestion generation module is used to build natural language suggestions based on context information to guide the large language model in generating test update code; The test generation module is used to generate test update code based on prompts from the large language model; The execution and verification module is used to compile and execute test update code; if there is a compilation failure or the runtime test fails, it analyzes the cause of the error, collects error-related context information, and provides feedback to the large language model to guide the repair of test cases; The iterative repair module is used to guide the large language model to repair the failed test update code, re-execute the execution and verification module to enter a new round of iteration, until the test passes or the number of iterations exceeds the set limit; The Minimize Modifications module instructs the large language model to adopt a minimal, conservative modification strategy when generating basic tests.
7. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement a unit test instant automatic update method based on a large language model as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a real-time automatic update method for unit tests based on a large language model as described in any one of claims 1-5.