Software defect reproduction test case generation method and system based on retrieval enhancement generation

By combining retrieval-enhanced generation and large language models, the problem of insufficient accuracy in generating defect reproduction test cases in software testing is solved, efficient and automated defect location and repair are achieved, and the efficiency and quality assurance of project development are improved.

CN120670306APending Publication Date: 2025-09-19YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510777451.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies lack accuracy in generating defect reproduction test cases in software testing, making it difficult to efficiently locate defects, affecting project development cycles and quality assurance.

Method used

The Retrieval Augmented Generation (RAG) technology is combined with a large language model. By preprocessing and vectorizing historical defect reports, a retrieval database is established. Similarity retrieval is used to obtain relevant reports. A prompt template is constructed to guide the large language model to generate defect reproduction test cases with a test method structure, and the effectiveness is verified.

Benefits of technology

It significantly improves the accuracy and efficiency of generating defect reproduction test cases, reduces labor costs, ensures that the generated test cases have good executable and reproducibility, and improves the reliability of the test and engineering applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670306A_ABST
    Figure CN120670306A_ABST
Patent Text Reader

Abstract

The invention discloses a software defect reproduction test case generation method and system based on retrieval enhancement generation, and the method comprises the steps: carrying out the preprocessing and vectorization of a historical defect report, and building a retrieval database; receiving a defect report of a to-be-generated test case, and obtaining a related historical defect report through similarity retrieval; constructing a prompt template based on a retrieval result, and guiding the large language model to generate a defect reproduction test case with a test method structure; verifying the validity of the code of the generated test case; according to the method, the generation accuracy of the defect reproduction test case can be effectively improved, and the software defect reproduction efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for generating a software defect reproduction test case, and in particular to a method and system for generating a software defect reproduction test case based on retrieval enhancement generation, belonging to the technical field of software testing. Background Art

[0002] Software testing is a crucial process in the software development lifecycle. Within this process, writing test cases is a crucial means of ensuring software quality. Test cases effectively verify system functionality, identify potential defects, and verify their complete resolution after fixes. However, in the actual development process, developers often face numerous challenges, such as missing or incorrect test cases. This makes it difficult to accurately locate defects, hindering the smooth progress of project remediation efforts and ultimately impacting project development cycles and quality assurance.

[0003] To address challenges such as a lack of test cases and difficulty reproducing defects, researchers have conducted extensive research on automated test case generation techniques. Current mainstream approaches typically use program code or specific key classes as input, leveraging static or dynamic analysis to explore the program's control and data flow paths to generate test cases. While these techniques have advanced defect reproduction research to some extent, practical applications still face challenges such as insufficient test accuracy and low reproduction rates, requiring further improvement and exploration.

[0004] In recent years, researchers have discovered that large language models (LLMs) demonstrate significant advantages in natural language processing and code generation tasks. In particular, when these models are combined with retrieval-augmented generation (RAG) technology, researchers can effectively mitigate model hallucinations by fusing external knowledge with language model inputs and improve the generation quality of specific tasks. However, researchers currently face challenges in efficiently integrating RAG technology with defect reproduction test generation tasks and establishing stable and reliable verification mechanisms. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a method and system for generating software defect reproduction test cases based on retrieval-enhanced generation, which can improve the efficiency of defect reproduction.

[0006] Technical solution: The present invention provides a method for generating software defect reproduction test cases based on retrieval-enhanced generation, comprising:

[0007] (1) Preprocess and vectorize historical defect reports to establish a retrieval database;

[0008] (2) Receive defect reports for test cases to be generated and obtain relevant historical defect reports through similarity retrieval;

[0009] (3) Constructing prompt templates based on the search results to guide the large language model to generate defect reproduction test cases with test method structures;

[0010] (4) Verify the validity of the generated test case code.

[0011] Furthermore, the step (1) includes:

[0012] (11) Extract key fields from historical defect reports and divide each report into independent retrieval blocks to remove redundant content; the key fields include title, description, code snippet and stack information;

[0013] (12) Constructing a retrieval database, using a text embedding model, converting the retrieval block text processed in step (11) into a high-dimensional semantic vector representation, and storing it in the retrieval database; the dimension of the high-dimensional semantic vector is 100-500 dimensions.

[0014] Furthermore, step (12) of constructing a search database is specifically as follows:

[0015] An index was built based on Facebook AI Similarity Search, and cosine similarity was used as the subsequent retrieval criterion to establish a historical defect vector retrieval database.

[0016] Furthermore, the step (2) includes:

[0017] (21) preprocessing the received defect report of the test case to be generated, wherein the preprocessing includes reconstruction, extraction of key fields, and removal of redundant content;

[0018] (22) Perform vector representation processing on the pre-processed defect report and build a temporary database to store the generated vectors and corresponding text content;

[0019] (23) Using a similarity algorithm, retrieve several historical reports that are most similar to the current defect report from the retrieval database, and then re-rank them based on the similarity score; wherein the threshold range of the similarity algorithm is 0.3-0.7, and the retrieval quantity range is 3-10.

[0020] Furthermore, the step (3) includes:

[0021] (31) constructing a prompt template prompt, integrating the current defect report and the historical defect report obtained in step (2) into a text format preset by the prompt template; the prompt template includes two parts: a system role definition and a user input, wherein the system role definition is used to specify the professional role of the model, and the user input part includes context information, defect report content, and test method generation instructions;

[0022] (32) Input the prompt template into a general dialogue generation language model for reasoning, and generate a defect reproduction test case with a test method structure;

[0023] (33) Extract and verify whether the generated test cases meet the format requirements of the test method structure. If the verification is passed, proceed to the subsequent steps.

[0024] Furthermore, the parameters of the large language model in step (32) include: a temperature parameter, a sampling parameter, a maximum length parameter, a word frequency penalty parameter, and an existence penalty parameter, and the parameters are used to control the randomness and diversity of the generated results.

[0025] Furthermore, the step (4) includes:

[0026] (41) Using lexical features, the test methods generated by the large language model are matched with the existing test classes in the target software engineering project. The best matching test class is determined as the injection carrier to achieve context binding and dependency completion of the test methods.

[0027] (42) Compiling and executing the test case on the defective version and the repaired version respectively for verification; the verification standard is: the test case compiles and fails to execute on the defective version, and compiles and executes successfully on the repaired version.

[0028] Furthermore, the step (42) includes:

[0029] First, check whether the test case can be successfully compiled on both the defective version and the fixed version;

[0030] If the compilation succeeds, check whether the test code fails to execute on the defective version and succeeds on the fixed version;

[0031] If the above conditions are met, the test case reproduction is determined to be successful; otherwise, the test case reproduction is determined to have failed and needs to be regenerated.

[0032] Based on the same inventive concept, the present invention also provides a system for generating software defect reproduction test cases based on retrieval-enhanced generation, comprising:

[0033] Initialization module, used to pre-process and vectorize historical defect reports and build a retrieval database;

[0034] The retrieval module is used to receive the defect report of the test case to be generated and obtain the relevant historical defect reports through similarity retrieval;

[0035] The generation module is used to build prompt templates based on the search results and guide the large language model to generate defect reproduction test cases with test method structures;

[0036] The verification module is used to verify the validity of the generated test case code.

[0037] Based on the same inventive concept, the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the method for generating software defect reproduction test cases based on retrieval-enhanced generation according to any of the above items.

[0038] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: 1. The combination of retrieval enhancement mechanism (RAG) and large language model can effectively alleviate the hallucination problem of large language model and improve the accuracy of generating defect reproduction test cases; 2. Automatically execute the entire process of defect report processing, retrieval, prompt construction, test generation and verification, greatly improving the efficiency of software defect reproduction and reducing labor costs; 3. Through class matching and dependency injection mechanism, it ensures that the generated test cases have good executable and reproducibility, significantly improving test reliability and engineering applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0040] Figure 2 Schematic diagram of a flow chart of an embodiment of the present invention;

[0041] Figure 3 This is a schematic diagram of defect report preprocessing according to an embodiment of the present invention;

[0042] Figure 4 A schematic diagram of retrieving similarity reports and re-ranking according to an embodiment of the present invention;

[0043] Figure 5 A schematic diagram of generating test code for a large language model according to an embodiment of the present invention;

[0044] Figure 6 The figure is a flow chart of the test case verification and iteration mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0046] As attached Figure 1 As shown, the method for generating software defect reproduction test cases based on retrieval-enhanced generation in this embodiment includes:

[0047] S100: Preprocess and vectorize historical defect reports to establish a retrieval database;

[0048] S200: Receive a defect report for a test case to be generated, and obtain related historical defect reports through similarity retrieval;

[0049] S300: Build a prompt template based on the search results to guide the large language model to generate defect reproduction test cases with a test method structure;

[0050] S400: Verify the validity of the generated test case code

[0051] Specifically, this embodiment is based on the defect instances in the Defects4J 2.0 dataset shown in Table 1. Figure 2 Take the GIANT method shown as an example.

[0052] Table 1 Defect instance table

[0053]

[0054] In step S100, historical defect data is collected, pre-processed and vectorized, and a standardized retrieval database is established to provide basic support for the subsequent generation of defect reproduction test cases. This is achieved through the following sub-steps:

[0055] S110: Historical defect data collection and preprocessing: Extract key fields (including title, description, code snippet and stack information) from historical defect reports, divide each report into independent retrieval blocks, and remove redundant images, links, emoticons and other invalid content, such as Figure 3 shown.

[0056] This example uses Defects4J 2.0 and screens out 750 valid defect instances. Each instance includes a defect revision (PRE_FIX_REVISION), a fixed revision (POST_FIX_REVISION), and its associated test cases. The specific processing methods are:

[0057] (1) Extract text from the defect summary, description, stack trace, and related code snippets;

[0058] (2) Cleaning rules include: removing HTML tags, removing special characters, and encoding in UTF-8 format;

[0059] (3) Use regular expressions to extract function names, class names, and exception information as important fields;

[0060] (4) Each processed record is saved in JSON format, and the structure includes: id, project, bug_id, summary, description, stacktrace, and code_snippet fields.

[0061] S120: Defect Report Vectorization and Retrieval Database Construction: Embedding technology is used to vectorize preprocessed defect text. The model selected is Sentence-BERT (all-MiniLM-L6-v2). The processing method is to concatenate the summary, description, and stack trace into an input, generating a 384-dimensional vector for each. The retrieval method is to build an index based on FAISS (Facebook AI Similarity Search) to support efficient vector retrieval. Finally, cosine similarity is used as the subsequent retrieval criterion to establish a historical defect vector retrieval database. This database supports subsequent similarity report retrieval and prompt construction.

[0062] All experiments were conducted in an Ubuntu 20.04 LTS environment on a machine equipped with 32 GB of memory, an Intel Xeon Gold 5318Y CPU, and an NVIDIA A30 GPU. Data processing scripts were written in Python 3.8, and Java code parsing was performed using the javalang 0.13.0 library.

[0063] S200: Current Defect Processing and Similar Report Retrieval: Standardize the defect report of the test case to be generated and search the retrieval database for the most relevant historical report based on the vectorization results to provide external context support for Prompt construction. This is achieved through the following sub-steps:

[0064] S210: Preprocessing and vectorization of current defect reports: For each defect report of a test case to be generated, the same method as S110 is used to perform data cleaning and unified processing.

[0065] S220: Then concatenate the summary, description, and stacktrace as input and call Sentence-BERT to encode them into a 384-dimensional vector to ensure data structure consistency.

[0066] S230: Similar defect retrieval and re-ranking: Based on the vector space retrieval mechanism, the similarity between the current defect and the historical defects is calculated, and several defect instances with the highest similarity scores are selected and arranged in descending order according to the scores, with the most similar report at the front.

[0067] The retrieval module uses the standard cosine similarity metric to ensure matching accuracy. The number of retrievals is set to n = 5, which can be adjusted according to actual needs; the retrieval results are sorted from high to low by score, and the top k most relevant ones are retained, with k defaulting to 5; instances with scores below 0.5 will be discarded to ensure high relevance. The retrieved similarity reports will serve as an external knowledge source for subsequent Prompt construction. Figure 4 As shown, the schematic process of similarity report retrieval and re-ranking is shown.

[0068] S300: Generation of Reproducible Test Cases: Build prompts based on defect reports and similar reports, guide the large language model to generate defect reproduction test cases with test method structure, and lay the foundation for automated testing. This is achieved through the following sub-steps:

[0069] S310: To build a clear prompt template, the current defect report and the sorted similar defect reports must be integrated into the preset text format of the prompt template and input into the large language model. Specifically, the core content such as the title, description, and code snippet of the current defect report is integrated with the most relevant external knowledge blocks sorted in ascending order of similarity to form a complete context, and Markdown syntax is used to distinguish different modules such as system prompts, defect reports, similar cases, and task instructions.

[0070] In the system prompt, the large language model is assigned the role of "Senior Software Engineer" to enhance professional understanding. At the same time, the template is specified to end with the "public void testXXX" format to clearly guide the model to complete the test method body, and the model is required to self-check the output results to verify the compilability of the code and whether the assertions cover the key features of the defects. This ensures that the model generates defect reproduction test cases that meet the requirements, forming a complete logical chain from information integration, format specification to role positioning, instruction guidance and result verification.

[0071] Specifically, the design of Prompt follows OpenAI's official best practices and has the following format:

[0072] Prompt construction: Based on the current bug description and retrieved historical similarity reports, we designed a system prompt role (senior software engineer) according to OpenAI best practices, annotated key content using Markdown syntax, and generated a test method in the format of public void testXXX() at the end of the prompt. Incorporating historical bug context into the prompt helps improve generation quality.

[0073] The prompt template is as follows:

[0074] Prompt Template

[0075] {

[0076] "Role":"System",

[0077] "Content":"""You are a senior software test engineer.You must provide a JUnit test method that can trigger the bug described in the bugreport.Context section may help you to trigger this bug."""

[0078] }

[0079] {

[0080] "Role":"User",

[0081] "Content":

[0082] """

[0083] #Context:

[0084] [Similarity Chunks]

[0085] #Bug report:

[0086] ##Bug Title:

[0087] [Bug Report Title]

[0088] ##Description:

[0089] [Bug Report Description]

[0090] #Test method of trigger bug:

[0091] Recheck your code to make sure it triggers[Bug Report Title]

[0092] ```

[0093] java

[0094] public void testMethodName(){

[0095] """

[0096] }

[0097] S320: Calling the large language model to generate test code: The prompts constructed using the template are input into the large language model, and a single-turn interaction is used for inference. The API interface is called to implement model inference to generate defect reproduction test cases with test method structures. Due to the randomness of the generation results of the large language model, even the same prompt may output different content. Therefore, to reduce the impact of randomness, a multi-round experimental regeneration strategy is adopted. At the same time, to reduce the cost of API calls, the query and generation are only re-run when the current generated test code is confirmed to be invalid, avoiding resource waste.

[0098] We used the ChatGPT-3.5-turbo API for single-turn dialogue inference, with the following parameters: temperature = 0.7; top_p = 1.0; max_tokens = 512; frequency_penalty = 0.0; and presence_penalty = 0.0. Each interaction started with a new dialogue state to ensure consistency and independence of generated results. The generated results were wrapped with ``` symbols to facilitate subsequent parsing and extraction.

[0099] S330: Extract and initially verify generated test cases: After obtaining the model output, heuristic rules are used to extract the code region enclosed by brackets (```). The extracted code is then verified to be in the complete test method format. Only complete and syntactically correct test code proceeds to the next step of the verification process. Incomplete or malformed code is directly judged as a generation failure, and the model query is re-initiated.

[0100] Specifically, based on the heuristic rules, the code content wrapped between ``` is extracted and preliminarily checked to see if it complies with the standard structure of the Java test method, including: whether it contains the @Test annotation; whether there is a method in the format of public void testXXX(); if not, the generation is determined to be invalid and the regeneration process is entered into S430; if it is satisfied, the verification process is entered into S400. Figure 5 As shown in the figure, the schematic process from prompt input to test code generation is shown.

[0101] S400: Verification and iteration of generated test case code: Dependency analysis and injection are performed on the generated test case code, and reproducibility verification is performed. If verification fails, it is regenerated through iteration until successful reproduction or the maximum number of iterations is reached. This is achieved through the following sub-steps:

[0102] S410: Test Code Dependency Parsing and Injection: Because the large language model cannot understand the entire project structure, the generated test cases often lack necessary dependencies. To this end, this step uses lexical analysis to match the test methods generated by the large language model with existing test classes in the target software engineering project. The best matching test class is identified as the injection carrier to achieve context binding and dependency completion for the test method.

[0103] Specifically, the generated code is first analyzed for dependent objects, parsed using Javalang's AST (Abstract Syntax Tree). If undefined objects or class references are found, common import statements are automatically completed. Then, using a lexical similarity matching strategy, the newly generated test methods are automatically inserted into the most matching existing test class. Once dependency injection is complete, a compilable test case file is generated.

[0104] The details are shown in Algorithm 1:

[0105]

[0106] First, use the findBestMatchingClass function to identify the class c that is most lexically similar to the given test method tm. best , and retrieve the dependencies of the method in lines 1-2. We calculate the matching score for each test class according to formula (1):

[0107]

[0108] Where T t and T ci They are the token sets in the generated test method and the i-th test class respectively.

[0109] Next, it utilizes getUnresolved in line 3 to identify the unresolved dependencies needed_deps. For each unresolved dependency, it checks whether its class definition exists; if not, it uses findMostCommonImport to find the most common import to fill the gap in lines 5-8. Finally, the algorithm uses the injectTest and injectDependencies functions in lines 13-14 to inject the test method and its required dependencies into the updated test suite. middle.

[0110] S420: Compilation and execution verification: After dependency resolution is completed, the test case is compiled and executed on the buggy version (BuggyVersion) and the fixed version (FixedVersion) respectively.

[0111] Specifically, use Maven or Gradle to compile the project in the defective version (PRE_FIX_REVISION) and fixed version (POST_FIX_REVISION) environments respectively. If the compilation fails, record the reason and mark it as a reproduction failure; if the compilation succeeds, execute the test case and check the assertion results; if the test fails on the defective version but succeeds on the fixed version, it is considered a successful reproduction.

[0112] The verification criteria are: the test case compiles and fails on the defective version, and compiles and succeeds on the fixed version. The verification logic is as follows:

[0113]

[0114] S430: Iterative generation and optimization: If the current test case fails to reproduce the defect, readjust the prompt or regenerate a new test method and continue iterative query. Set a maximum of 10 iterations for each defect to balance generation quality and resource consumption. Figure 6 As shown in the figure, the process of test verification and iteration mechanism is demonstrated.

[0115] Based on the same inventive concept, this embodiment also provides a system for generating software defect reproduction test cases based on retrieval-enhanced generation, including:

[0116] Initialization module, used to pre-process and vectorize historical defect reports and build a retrieval database;

[0117] The retrieval module is used to receive the defect report of the test case to be generated and obtain the relevant historical defect reports through similarity retrieval;

[0118] The generation module is used to build prompt templates based on the search results and guide the large language model to generate defect reproduction test cases with test method structures;

[0119] The verification module is used to verify the validity of the generated test case code.

[0120] Based on the same inventive concept, this embodiment also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the method for generating software defect reproduction test cases based on retrieval-enhanced generation according to any of the above items.

[0121] It can be seen from the above embodiments that the present invention can combine defect report retrieval enhancement and large language model reasoning capabilities to automatically and efficiently generate defect reproduction test cases, greatly improving the testing efficiency and accuracy in the defect repair process, and has good engineering application value and promotion prospects.

Claims

1. A method for generating software defect reproduction test cases based on retrieval-enhanced generation, characterized in that: include: (1) Preprocess and vectorize historical defect reports to establish a retrieval database; (2) Receive defect reports for test cases to be generated and obtain relevant historical defect reports through similarity retrieval; (3) Constructing prompt templates based on the search results to guide the large language model to generate defect reproduction test cases with test method structures; (4) Verify the validity of the generated test case code.

2. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 1, characterized in that: The step (1) comprises: (11) Extract key fields from historical defect reports and divide each report into independent retrieval blocks to remove redundant content; the key fields include title, description, code snippet and stack information; (12) Constructing a retrieval database, using a text embedding model, converting the retrieval block text processed in step (11) into a high-dimensional semantic vector representation, and storing it in the retrieval database; the dimension of the high-dimensional semantic vector is 100-500 dimensions.

3. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 2, characterized in that: The step (12) of constructing a search database is specifically as follows: An index was built based on Facebook AI Similarity Search, and cosine similarity was used as the subsequent retrieval criterion to establish a historical defect vector retrieval database.

4. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 1, characterized in that: The step (2) comprises: (21) preprocessing the received defect report of the test case to be generated, wherein the preprocessing includes reconstruction, extraction of key fields, and removal of redundant content; (22) Perform vector representation processing on the pre-processed defect report and build a temporary database to store the generated vectors and corresponding text content; (23) Using a similarity algorithm, retrieve several historical reports that are most similar to the current defect report from the retrieval database, and then re-rank them based on the similarity score; wherein the threshold range of the similarity algorithm is 0.3-0.7, and the retrieval quantity range is 3-10.

5. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 1, characterized in that: The step (3) comprises: (31) constructing a prompt template prompt, integrating the current defect report and the historical defect report obtained in step (2) into a text format preset by the prompt template; the prompt template includes two parts: a system role definition and a user input, wherein the system role definition is used to specify the professional role of the model, and the user input part includes context information, defect report content, and test method generation instructions; (32) Input the prompt template into a general dialogue generation language model for reasoning, and generate a defect reproduction test case with a test method structure; (33) Extract and verify whether the generated test cases meet the format requirements of the test method structure. If the verification is passed, proceed to the subsequent steps.

6. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 4, characterized in that: The parameters of the large language model in step (32) include: a temperature parameter, a sampling parameter, a maximum length parameter, a word frequency penalty parameter, and an existence penalty parameter, and the parameters are used to control the randomness and diversity of the generated results.

7. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 1, characterized in that: The step (4) comprises: (41) Using lexical features, the test methods generated by the large language model are matched with the existing test classes in the target software engineering project. The best matching test class is determined as the injection carrier to achieve context binding and dependency completion of the test methods. (42) Compiling and executing the test case on the defective version and the repaired version respectively for verification; the verification standard is: the test case compiles and fails to execute on the defective version, and compiles and executes successfully on the repaired version.

8. The method for generating software defect reproduction test cases based on retrieval-enhanced generation according to claim 1, characterized in that: The step (42) comprises: First, check whether the test case can be successfully compiled on both the defective version and the fixed version; If the compilation succeeds, check whether the test code fails to execute on the defective version and succeeds on the fixed version; If the above conditions are met, the test case reproduction is determined to be successful; otherwise, the test case reproduction is determined to have failed and needs to be regenerated.

9. A software defect reproduction test case generation system based on retrieval-enhanced generation, characterized in that: include: Initialization module, used to pre-process and vectorize historical defect reports and build a retrieval database; The retrieval module is used to receive the defect report of the test case to be generated and obtain the relevant historical defect reports through similarity retrieval; The generation module is used to build prompt templates based on the search results and guide the large language model to generate defect reproduction test cases with test method structures; The verification module is used to verify the validity of the generated test case code.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method for generating software defect reproduction test cases based on retrieval-enhanced generation according to any one of claims 1 to 8 are implemented.