Test case generation method and system based on multi-dimensional feedback and causal attribution

By employing a test case generation method based on multidimensional feedback and causal attribution, and utilizing a large language model and mutation score optimization generation strategy, this approach addresses the issues of single feedback and inefficient iteration in existing technologies. It achieves high-quality and efficient test case generation, thereby improving test quality and defect detection capabilities.

CN120909945APending Publication Date: 2025-11-07HAIER CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511228791.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies provide only one dimension of feedback when generating test cases and lack intelligent iteration strategies, resulting in limited improvement in test quality. Furthermore, they lack the ability to effectively evaluate generated test cases and delve into deeper code logic.

Method used

A test case generation method based on multidimensional feedback and causal attribution is adopted. Test cases are generated, compiled and executed through a large language model, and a quality assessment report is generated. The generation strategy is optimized by using mutation scores and strategy summaries until the quality target is achieved.

Benefits of technology

It achieves high-quality test case generation with high defect detection capability, improves the efficiency and quality of test case generation, solves the problems of insufficient feedback and inefficient iteration in traditional methods, and enhances the intelligence and adaptability of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909945A_ABST
    Figure CN120909945A_ABST
Patent Text Reader

Abstract

According to the test case generation method and system based on multi-dimensional feedback and causal attribution, from generation of an initial test case to repeated generation based on a strategy abstract until the test case reaches the standard, a generation strategy is adjusted according to the logic characteristics of a source code to be tested and the weak link of the current test case, so that the test case generation efficiency is improved. Developers / testers can be helped to clearly understand the thinking process of the automatic test system, and the problem of'black box 'in the traditional automatic test generation process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field related to test case generation, and particularly relates to a test case generation method and system based on multi-dimensional feedback and causal attribution. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] Software testing is the core link to ensure software quality, and the writing of test cases accounts for a considerable part of the development and testing cost. In recent years, with the breakthrough of large language models in the field of code generation, the use of large language models to automatically generate test cases has become a research hotspot.

[0004] However, the current use of large language models to automatically generate test cases at least has the following defects: 1. Single and superficial feedback dimension: Existing automated test generation systems, such as preliminary applications based on large language models, usually only rely on single code coverage such as line coverage and branch coverage as feedback indicators. When the large language model generates a set of test cases, the system executes them and calculates the coverage, and then simply feeds back the un-covered code line number or branch information to the large language model. This feedback method is "phenomenal" rather than "semantic", it tells the model "where it hasn't been tested", but cannot tell the model "why it hasn't been tested" and "how it should be tested". For example, the large language model has difficulty understanding the business logic or abnormal scenarios behind the "45th line else branch", resulting in limited improvement in the quality of subsequent test generation, and easily falling into local optimization or random attempts.

[0005] 2. Insufficient test quality evaluation: Test suites that only focus on coverage may have the "insecticide paradox" problem, i.e. test cases can cover the code, but cannot find the hidden defects in the code. For example, a test may execute the a>b line of code, but its assertion part does not effectively verify the correct behavior when a is really greater than b. Existing technologies generally lack effective evaluation of the "defect discovery ability" of generated test cases.

[0006] 3. Lack of intelligent iteration strategy: The existing iterative generation process is usually mechanical. The system directly lists the un-covered code as part of the prompt, requiring the large language model to supplement in the next round. This approach lacks macro planning of the test strategy, resulting in low efficiency of the generation process and difficulty in reaching deep code logic that requires complex preconditions to trigger.

[0007] In summary, how to provide deep, multi-dimensional and understandable feedback to the large language model so that the large language model can generate high-quality test cases with high defect detection ability is a problem to be solved. SUMMARY

[0008] In order to overcome the above-mentioned deficiencies of the prior art, the present application provides a test case generation method and system based on multi-dimensional feedback and causal attribution, which adjusts the generation strategy according to the logical characteristics of the source code to be tested and the weak link of the current test case, and can generate test cases with high quality and high defect detection capability.

[0009] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions: In a first aspect, the present application provides a test case generation method based on multi-dimensional feedback and causal attribution, comprising: For the source code to be tested, a large language model is used to generate test cases, the generated test cases are compiled and executed, and execution results are obtained; Based on the source code to be tested and the execution results, a quality evaluation report is generated; wherein the quality evaluation report includes a mutation score evaluating the test case; The key data corresponding to the quality evaluation report and the source code to be tested are filled into a preset prompt template to obtain a target prompt, and based on the target prompt, a large language model is used to analyze the fundamental weaknesses of the generated test cases, and a strategy summary is obtained; According to the strategy summary and the source code to be tested, the large language model is used to generate test cases again, and the above steps are repeated until test cases reaching the quality target are obtained.

[0010] In a second aspect, the present application provides a test case generation system based on multi-dimensional feedback and causal attribution, comprising: An execution module configured to: for the source code to be tested, a large language model is used to generate test cases, the generated test cases are compiled and executed, and execution results are obtained; A quality evaluation module configured to: based on the source code to be tested and the execution results, a quality evaluation report is generated; wherein the quality evaluation report includes a mutation score evaluating the test case; An analysis module configured to: the key data corresponding to the quality evaluation report and the source code to be tested are filled into a preset prompt template to obtain a target prompt, and based on the target prompt, a large language model is used to analyze the fundamental weaknesses of the generated test cases, and a strategy summary is obtained; A generation module configured to: according to the strategy summary and the source code to be tested, the large language model is used to generate test cases again, and the above steps are repeated until test cases reaching the quality target are obtained.

[0011] In a third aspect, the present application provides an electronic device comprising a memory and a processor, and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the method of the first aspect is completed.

[0012] In a fourth aspect, the present application provides a computer readable storage medium for storing computer instructions, when the computer instructions are executed by the processor, the method of the first aspect is completed.

[0013] In a fifth aspect, the present application provides a computer program product comprising a computer program, when the computer program is executed by the processor, the method of the first aspect is completed.

[0014] The above one or more technical solutions have the following beneficial effects: In the present application, from generating the initial test case, to repeatedly generating based on the policy summary, until reaching the standard, according to the logical characteristics of the source code to be tested and the weak link of the current test case, the generation strategy is adjusted, which can help the development / test personnel to clearly understand the thinking process of the automatic test system, and solves the problem of "black box" in the traditional automatic test generation process.

[0015] In the present application, "mutation score" is introduced as the core index of the quality evaluation report, and the mutation score is generated by the mutation test tool for the source code to be tested, and then calculated according to the test case execution result. This makes the generated test case not only pursue the traditional code coverage, but also focus on "the effectiveness of discovering defects", which can avoid covering the code but cannot detect the defects, and the finally output test case is far superior to the traditional scheme which only focuses on coverage in ensuring software quality.

[0016] In the present application, the complex test analysis result is converted into natural language strategy which is easy for the large language model to understand, which makes each iteration of the large language model be "intelligent optimization" with clear goal, rather than "blind trial and error", thereby greatly accelerating the convergence speed of reaching the high-quality test goal.

[0017] The advantages of the additional aspects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0018] The drawings accompanying the specification of the present application serve to provide further understanding of the present application, the illustrative embodiments of the present application and the description thereof serve to explain the present application, and do not constitute an improper limitation on the present application.

[0019] Figure 1 The flow chart of the test case generation method based on multi-dimensional feedback and causal attribution in the embodiment one of the present application. DETAILED DESCRIPTION

[0020] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0021] It should be noted that the terms used herein are only intended to describe specific embodiments and are not intended to limit exemplary embodiments according to the present application.

[0022] In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0023] Embodiment one The present embodiment discloses a test case generation method based on multi-dimensional feedback and causal attribution, comprising: For the source code to be tested, a large language model is used to generate test cases, the generated test cases are compiled and executed, and execution results are obtained; Based on the source code to be tested and the execution results, a quality evaluation report is generated; wherein the quality evaluation report includes a mutation score for evaluating the test cases; The key data corresponding to the quality evaluation report and the source code to be tested are filled into a preset prompt template to obtain a target prompt, and the large language model is guided based on the target prompt to analyze the root weakness of the generated test cases, and a strategy summary is obtained; According to the strategy summary and the source code to be tested, the large language model is used to generate test cases again, and the above steps are repeated until test cases that meet the quality target are obtained.

[0024] The present embodiment, from generating initial test cases, to repeatedly generating based on the strategy summary, until reaching the standard, according to the logical characteristics of the source code to be tested and the weak links of the current test cases, adjusting the generation strategy, can help the development / test personnel to clearly understand the thinking process of the automatic test system, and solve the problem of "black box" in the traditional automatic test generation process.

[0025] The test case generation method based on multi-dimensional feedback and causal attribution proposed in the present embodiment will be described in detail as follows: Step 1: For the source code to be tested, a large language model is used to generate test cases, the generated test cases are compiled and executed, and execution results are obtained.

[0026] Specifically, the large language model can adopt GPT-4 or CodeLlama, and the test cases are generated by the large language model through the source code to be tested .

[0027] ​Step 2: Based on the source code to be tested and the execution result, generate a quality evaluation report; wherein the quality evaluation report includes the mutation score of the evaluation test case.

[0028] In this embodiment, the test case is compiled and executed , and the execution result is collected , the execution result includes the pass / fail status of the test, logs, etc.

[0029] In this embodiment, according to the source code and the execution result , multi-dimensional quality analysis is performed, and a structured quality report is output.

[0030] Specifically, for a test case , its quality can be defined as a vector:

[0031] Wherein: are the traditional line coverage and branch coverage, respectively, is the mutation score, is the assertion density.

[0032] The mutation score is a key indicator of the "defect discovery ability" of the test case. First, use a mutation testing tool such as Pitest for Java to mutate the source code to generate a set of mutation bodies with minor defects . Then execute the test case , determine the detection ability of the generated test case for each mutant in the set of minor defect mutants by judging the execution result of the test case, and determine the mutation score.

[0033] is defined as:

[0034] Wherein, is the set of mutants successfully "killed" (i.e. causing the test to fail) by the test case . is the total mutant set, refers to the complete set of all mutants generated by the mutation testing tool such as Pitest according to the preset rules after modifying the original code, that is , this set is the basis for calculating the mutation score. is the set of equivalent mutants, i.e. the set of mutants whose code logic has not changed. A high means that the test suite's assertions are very effective at catching subtle changes in the code.

[0035] assertion density Measures the ratio of effective assertions in the test cases to the number of lines of test code, to prevent generating "empty runs" of tests.

[0036] quality report Contains the concrete values of the aforementioned metrics, as well as the list of uncovered branches and the list of not-killed mutants.

[0037] The specific flow of this embodiment is as follows: 1. Test execution: First, the generated test cases are executed.

[0038] 2. Quality analysis (QAM): Subsequently, the execution process and results are analyzed in depth, a process that usually calls two core tools: Obtain "uncovered branches information": Use a code coverage tool (e.g. JaCoCo for Java); when running the test cases, it monitors which lines and branches of code are executed. At the end of the run, a detailed report is generated that precisely lists all the branches that were not executed, this list is what is called "uncovered branches information". Contains the "uncovered branches list".

[0039] Obtain "not-killed mutants list": Use a mutation test such as Pitest for Java. First, multiple "mutants" of the source code are created. Then, the existing test suite is used to run these mutants. If a mutant is run and no test case fails, then this mutant is considered "alive" (i.e. "not-killed"). The mutation test tool outputs a report that contains all the mutants that survived, i.e. the "not-killed mutants list".

[0040] In summary, these two key lists are the core output of the quality analysis, transforming the raw results of the test execution into structured data that can be analyzed further. This detailed report that contains both lists is the multi-dimensional quality report .

[0041] Quality vector Q(T) is a summary report that macroscopically evaluates the quality of the test suite with several core numbers such as 80% branch coverage, 75% mutation score, etc. Quality vector Q(T) is used to determine whether the quality target is reached, and the specific reasons that lead to the unsatisfactory summary numbers, i.e. process details and raw data, are transformed into LLM-understandable, instructive strategies that guide LLM to improve.

[0042] Step 3: Fill in the key data corresponding to the quality assessment report and the source code to be tested into the preset prompt template to obtain the target prompt. Based on the target prompt, guide the large language model to analyze the fundamental weaknesses of the generated test cases, and obtain a strategy summary.

[0043] In this embodiment, a complex Prompt is constructed and input into a large language model. The large language model used in step 3 plays a different role from the large language model used in step 1.

[0044] The Prompt contains: source code

[0045] uncovered branch information a list of surviving mutants, as well as the source code location and mutation type corresponding to these mutants.

[0046] In this embodiment, through the target Prompt, the large language model is required to perform the following reasoning tasks, such as a thought chain CoT: semantic clustering: associating uncovered branches and surviving mutants semantically. For example, the large language model may find that "an uncovered catch(IOException) branch" and "a surviving mutant about file handle not closed" both point to "file IO exception handling" as a weak link.

[0047] root cause inference: infer the root cause of these test weaknesses. For example, infer that "the current test data are all valid, existing file paths, and never simulate the scenario of file nonexistence or no permission to access".

[0048] strategy summary generation: based on the above inference, the large language model is required to generate a concise, clear, and actionable natural language "strategy summary" .

[0049] The format of the strategy summary is strictly constrained, usually containing three parts: [weak link]: highly summarize the current main weakness of the test suite. (e.g., "insufficient handling of negative and zero values for user input.")

[0050] [Action Proposal]: Provide specific, next-round test scenarios to be generated. (e.g., "Please design test cases using -10, 0 as input, and verify if the function can throw an IllegalArgumentException as expected.")

[0051] [Context Hint]: Provide necessary context or code snippets to help the LLM better locate. (e.g., "Focus on the if (input>0) condition in the calculate_rate function.").

[0052] Specifically, the target Prompt guides the large language model to perform the following reasoning tasks. This process is usually done through a carefully designed, structured single Prompt containing "Chain of Thought" (CoT) instructions, so that the LLM can maintain a comprehensive understanding of all contexts such as source code, test vulnerabilities, etc. in a complete call, ensuring the coherence of reasoning.

[0053] 1. Implementation of Chain of Thought (CoT): Guided reasoning within a single Prompt.

[0054] Chain of Thought CoT does not refer to the linking of multiple Prompts, but a special Prompting technique that explicitly instructs the large language model to perform step-by-step, logical analysis and reasoning before giving the final answer. This carefully designed Prompt will contain the following parts: Role Setting (Role Setting): instruct the LLM to play a role. Context Input (Context Input): provide all necessary original data. Task Instruction (Task Instruction): contains CoT instructions, requires the LLM to complete the reasoning step by step, and outputs the final result in a specific format.

[0055] 2. Structure example of unified Prompt.

[0056] The following is a specific Prompt template example that integrates all data and instructions: # ROLE: You are an expert software quality assurance engineer and a seniortest architect. Your task is to analyze a given set of testing weaknesses anddevise a strategic plan to address them. # CONTEXT: Here is the relevant information for your analysis. [Source Code S]: --- / / Java code for a file utility class public class FileProcessor { public void processFile(String filePath) throws IOException { File file = new File(filePath); if (!file.exists()) { throw new IllegalArgumentException("File does not exist."); } / / ... complex logic for processing file ... FileReader reader = null; try { reader = new FileReader(file); / / ... read file content ... } catch (IOException e) { / / This is line 52 throw new IOException("Failed to read file.", e); } finally { / / ... potential bug here, reader might not be closed if constructorfails ... } } } --- [Uncovered Branches]: --- - FileProcessor.java, Line 52: The catch(IOException) block was notexecuted. --- [Un-killed Mutants]: --- - MUTANT_ID_123: - Location: FileProcessor.java, Line 52 - Type: Method Call Removal - Description: The original code `throw new IOException(...)` wasreplaced with ` / / do nothing`. - Status: SURVIVED (This means no test failed when the exception wasswallowed). --- # TASK: Based on the provided CONTEXT, perform the following reasoning stepsand then generate a final "Strategy Summary". Reasoning Steps (Chain-of-Thought): 1. Semantic Clustering: Analyze the [Uncovered Branches] and [Un-killed Mutants]lists. 确定这些项目中是否有任何语义相关。 将它们分组为集群,并为每个集群的主题提供高级描述。 2. 根本原因推断: 对于每个识别出的集群,推断当前测试套件不足的根本原因。这不是关于代码为什么错误,而是关于测试为什么未能检测到弱点。 根据你的推断,描述当前缺少的特定测试场景或测试数据类型。 策略制定:根据你的根本原因推断,为下一轮测试生成制定一个具体且可操作的计划。 最终输出: 完成推理后,仅以下列结构化格式提供最终输出。 [Weakness]: {A concise summary of the main weakness identified in your reasoning} [Actionable Suggestions]: {Specific, actionable instructions for generating new test cases} [Context hints]: {Relevant code snippets or context to help focus the test generation} 3. Correlation between parts of the results When the LLM receives the above Prompt, its "thinking" process is as follows: Step 1: Semantic Clustering. LLM's internal reasoning: "I see an uncovered branch is catch(IOException). At the same time, I see a surviving mutant is also in the same line, whose role is to let IOException be swallowed (not thrown). These two issues are obviously related. They both point to the same theme." LLM's intermediate result (part of CoT): Clustering: {Correlation Items: [Uncovered Branch: Line 52], [Surviving Mutant: MUTANT_ID_123]}. Theme: Insufficient test coverage of exception handling paths for file IO operations.

[0057] Step 2: Root Cause Inference. LLM's internal reasoning: "In the last step, I identified the weakness as 'file IO exception handling'. Why didn't the existing tests cover this scenario? Because all the test cases provided filePaths that were probably valid, readable files. The tests never simulated real scenarios that would trigger IOException, such as the file becoming unreadable during reading or the disk being full. Therefore, the root cause is that the test scenarios are monotonous and lack simulation of IO exception boundary conditions." LLM's intermediate result (part of CoT): Root Cause: Current test data only covers the "success path" of file handling, completely without providing negative test scenarios that would trigger java.io.IOException.

[0058] Step 3: Generation of Strategy Summary (Constrained by Inference). Internal inference of LLM: "Since the root cause is the lack of IO exception scenarios, my strategy must explicitly require the generation of such scenarios. I need to convert this inference result into formatted instructions." How to constrain: The strict output format of the TASK section serves as a powerful constraint. The LLM is forced to fill in the conclusions it has drawn in the "root cause inference" into the three fields of [Weak Link], [Action Proposal], and [Context Hint]. Final output of LLM (formatted): [Weak Link]: The exception handling path test for file IO operations is severely insufficient. [Action Proposal]: Please design a test case that simulates the scenario of IOException occurring during file reading. For example, you can try to operate a locked file or a file with insufficient permissions. The assertion must verify whether the processFile method correctly captures and throws IOException upwards. [Context Hint]: Focus on the catch (IOException e) code block at line 52 in FileProcessor.java.

[0059] The above effects are mainly achieved through sophisticated Prompt Engineering. These large models have learned the ability to code, test, logically reason, and follow instructions in their vast training data. Through the structured instructions of CoT, these existing capabilities can be effectively "stimulated" and guided to complete complex analysis tasks.

[0060] Step 4: According to the strategy summary and the source code to be tested, use the large language model to generate test cases again, repeat the above steps until the test cases that meet the quality target are obtained.

[0061] In this embodiment, the strategy summary and the source code to be tested are input into the large language model again to generate test cases, and steps 2-3 are repeated until the preset quality target or the number of iterations is reached, and finally the optimized test suite is output .

[0062] This embodiment introduces mutation testing as the core feedback indicator, and the generated test cases not only pursue high coverage, but also pursue high defect detection capability. This makes the final output of the test suite far superior to traditional methods in ensuring software quality, significantly improving the quality and defect detection capability of test cases.

[0063] The embodiment can convert complex test analysis results into natural language strategies that are easy for large language models to understand through innovative "causal attribution". This makes each iteration of the large language model a "intelligent optimization" with a clear goal, rather than "blind exploration", greatly accelerating the convergence speed to the high-quality test goal, and greatly improving the intelligence and efficiency of test generation.

[0064] The embodiment constructs a complete automatic closed-loop scheme, automatically adjusts its generation strategy according to the characteristics of the code to be tested and the weak links of the current test suite, and exhibits high adaptability. Without human intervention, it can continuously output high-quality test results, realizing self-evolution and self-adaptation of the test generation process.

[0065] In the embodiment, the "strategy summary" itself is a high-level summary of the current test status, providing excellent interpretability for developers and testers, helping them understand the "thinking process" and work focus of the automated test system, and enhancing the interpretability of the generation process.

[0066] The embodiment effectively solves the core pain points in existing LLM test generation techniques through its unique multi-dimensional feedback and causal attribution mechanism, providing an innovative and practical technical solution for the field of automated and intelligent software testing.

[0067] Embodiment two The purpose of the embodiment is to provide a test case generation system based on multi-dimensional feedback and causal attribution, which includes: The execution module is configured to: for the source code to be tested, generate test cases using a large language model, compile and execute the generated test cases, and obtain execution results; The quality evaluation module is configured to: based on the source code to be tested and the execution results, generate a quality evaluation report; wherein the quality evaluation report includes a mutation score for evaluating the test cases; The analysis module is configured to: fill the key data in the quality evaluation report and the source code to be tested into a preset prompt template to obtain a target prompt, guide the large language model to analyze the root weakness of the generated test cases based on the target prompt, and obtain a strategy summary; The generation module is configured to: according to the strategy summary and the source code to be tested, generate test cases again using a large language model, repeat the above steps, and obtain test cases that meet the quality target.

[0068] In more embodiments, there are also provided: An electronic device includes a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in embodiment one is completed. For brevity, this will not be repeated here.

[0069] It should be understood that, in this embodiment, the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0070] The memory can include read-only memory and random access memory, and provide instructions and data to the processor, and a portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.

[0071] A computer readable storage medium for storing computer instructions, which are executed by a processor to complete the method described in embodiment one.

[0072] The method in embodiment one can be directly embodied as a hardware processor to complete, or be completed by a combination of hardware and software modules in the processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers or other mature storage media in the art. The storage medium is located in the memory, and the processor reads information in the memory to complete the steps of the above method in combination with the hardware. To avoid repetition, it will not be described in detail here.

[0073] A computer program product comprising a computer program, which, when executed by a processor, implements the method described in embodiment one.

[0074] The present application also provides at least one computer program product tangibly stored on a non-transitory computer readable storage medium. The computer program product includes computer executable instructions, such as instructions included in program modules, which are executed by devices on real or virtual processors of the target to perform processes / methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. In various embodiments, the functions of the program modules can be combined or divided as desired. Machine executable instructions for program modules can be executed within a local or distributed device. In a distributed device, program modules can be located in local and remote storage media.

[0075] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages. The computer program code can execute entirely on a computer, a special purpose computer, or other programmable apparatus to produce the functions / acts specified in the flow diagrams and / or block diagrams. The program code can execute entirely on a computer, a special purpose computer, or other programmable apparatus, as a stand-alone software package, partly on the computer and partly on a remote computer, or entirely on the remote computer or server.

[0076] In the context of the present application, the computer program code or related data can be carried by any suitable carrier to enable the device, apparatus or processor to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, and the like. Examples of signals can include electrical, optical, radio, sound or other forms of propagated signals, such as carrier waves, infrared signals, and the like.

[0077] Those skilled in the art can realize that the units and algorithm steps of the examples described in conjunction with the present embodiments can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0078] The above describes the specific embodiments of the present application in conjunction with the accompanying drawings, but is not a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications or variations made by those skilled in the art on the basis of the technical solutions of the present application without inventive labor are still within the scope of protection of the present application.

Claims

1. A method for test case generation based on multidimensional feedback and causal attribution, characterized in that, The method comprises the following steps: For the source code to be tested, a large language model is used to generate test cases, the generated test cases are compiled and executed, and execution results are obtained; Based on the source code to be tested and the execution results, a quality evaluation report is generated; wherein the quality evaluation report includes a mutation score evaluating the test cases; The key data corresponding to the quality evaluation report and the source code to be tested are filled into a preset prompt template to obtain a target prompt, and based on the target prompt, a large language model is guided to analyze the fundamental weaknesses of the generated test cases, and a strategy summary is obtained; According to the strategy summary and the source code to be tested, the large language model is used to generate test cases again, and the above steps are repeated until test cases reaching the quality target are obtained.

2. The test case generation method based on multi-dimensional feedback and causal attribution of claim 1, wherein, The mutation score is determined by performing a mutation operation on the source code to be tested to obtain a set of micro-defect mutants, executing the generated test cases, and determining the mutation score by judging the detection ability of the generated test cases on each mutant in the set of micro-defect mutants through the test case execution results.

3. The test case generation method based on multi-dimensional feedback and causal attribution of claim 1, wherein, The target prompt includes the source code to be tested, uncovered branch information, a list of non-killed mutants, and the source code location and mutation type corresponding to the non-killed mutants.

4. The method for test case generation based on multi-dimensional feedback and causal attribution as claimed in claim 1 wherein, Based on the target prompt, the large language model is guided to analyze the fundamental weaknesses of the generated test cases, and a strategy summary is obtained, which specifically comprises: Semantically associating the uncovered branches and the non-killed mutants; According to the semantic association result, the root cause of the test weakness is inferred; According to the inference result of the root cause, a strategy summary including weak links, action suggestions and context prompts is generated.

5. The multi-dimensional feedback and causal attribution based test case generation method of any one of claims 1-4, wherein, The quality evaluation report also includes line coverage, branch coverage and assertion density.

6. The multi-dimensional feedback and causal attribution based test case generation method of any one of claims 1-4, wherein, Based on the target prompt, the CoT is used to guide the large language model to reason to generate the fundamental weaknesses of the test cases, and the strategy summary is obtained.

7. A test case generation system based on multidimensional feedback and causal attribution, characterized in that, The method comprises the following steps: An execution module is configured to: for the source code to be tested, a large language model is used to generate test cases, the generated test cases are compiled and executed, and execution results are obtained; A quality evaluation module is configured to: based on the source code to be tested and the execution results, a quality evaluation report is generated; wherein the quality evaluation report includes a mutation score evaluating the test cases; An analysis module is configured to: the key data corresponding to the quality evaluation report and the source code to be tested are filled into a preset prompt template to obtain a target prompt, and based on the target prompt, a large language model is guided to analyze the fundamental weaknesses of the generated test cases, and a strategy summary is obtained; A generation module is configured to: according to the strategy summary and the source code to be tested, the large language model is used to generate test cases again, and the above steps are repeated until test cases reaching the quality target are obtained.

8. An electronic device, comprising: The computer program product comprises a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the method of any one of claims 1-6 is completed.

9. A computer-readable storage medium, characterized in that, The computer program product is used to store computer instructions, and when the computer instructions are executed by the processor, the method of any one of claims 1-6 is completed.

10. A computer program product, characterised in that, A computer program comprising computer program elements which, when executed by a processor, perform the method according to any one of claims 1-6. A computer program comprising computer program elements which, when executed by a processor, perform the method according to any one of claims 1-6.