A code test case generation system and method based on directional fine tuning
Patent Information
- Application Number
- CN202610989802.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0013]为了解决现有基于大语言模型的代码测试用例生成方法中存在的测试逻辑质量评价不足、微调样本构造不充分、偏好样本表面差异干扰较大以及生成测试用例强度与执行稳定性难以兼顾等问题,本发明提供了一种基于定向微调的代码测试用例生成系统及方法
[0045]本发明针对现有大语言模型代码测试用例生成方法中存在的逻辑质量评价不足、样本构造不精确、表面特征干扰较大以及测试强度与执行稳定性难以兼顾等问题,提供了一种可执行、可评估、可训练和可扩展的代码测试用例生成方案。相比于现有技术,本发明具有如下优点:
Smart Images

Figure CN122817080A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of software testing technology, and relates to a code test case generation system and method, specifically a code test case generation system and method based on targeted fine-tuning, test logic quality assessment and preference sample construction. Background Technology
[0002] Software testing is a crucial part of the software development and quality assurance process. Unit tests are used to verify the behavior of functions, methods, or modules in a program, enabling the early detection of functional defects, interface call errors, missing boundary conditions, and inadequate exception handling. High-quality unit tests should not only be executable but also contain valid assertions, cover typical inputs, boundary inputs, and abnormal paths, and have a certain ability to reveal potential program defects.
[0003] With the increasing application of large language models in tasks such as code generation, program repair, defect detection, and automated test generation, generating test cases using large language models has become an important technical direction in the field of software test automation. Current techniques typically input the code of the function under test, its function signature, natural language description, interface documentation, or context code into a large language model, which then directly generates the corresponding unit test code. This approach can reduce the workload of manually writing test cases to some extent and improve the efficiency of initial test case construction.
[0004] In existing methods for generating test cases for large language models, common approaches include direct generation based on prompt words, generation based on few-sample prompts, generation with contextual supplementation based on retrieval enhancement, and domain-adapted generation based on supervised fine-tuning. Some existing solutions also introduce execution checks, syntax checks, or coverage statistics after generation to determine whether the generated tests can run or whether they can cover certain statements and branches in the code under test. These techniques are effective in improving the efficiency of test case generation.
[0005] Among existing publicly available technologies, techniques such as large-model fine-tuning, retrieval-enhanced generation, sample construction, and generation result review have been applied to scenarios such as text generation, question-and-answer generation, dialogue decision-making, and professional domain content generation. For example, CN120578754A discloses a method, system, device, and medium for generating report drafts based on a large model, which achieves professional draft generation through document parsing, central sentence extraction, RAG retrieval, and large-model generation. CN120632050A discloses a dual-engine government affairs question-and-answer method based on large-model fine-tuning and RAG retrieval, which improves the accuracy of government affairs question-and-answer through domain knowledge base, vector retrieval, and model fine-tuning. CN121071100A discloses a method for making control logic decisions in complex dialogues based on large-model fine-tuning and dynamic examples, which assists the model in outputting control logic decisions by constructing dialogue data and recalling similar and negative cases. The above technologies demonstrate that large-model fine-tuning, dynamic example construction, and generation result control have become common technical means in intelligent generation systems.
[0006] However, in code test case generation scenarios, existing technologies still have the following shortcomings:
[0007] (1) Existing methods for generating test cases from large language models often focus on whether the test code is syntactically correct, can be executed successfully, or achieves a certain code coverage, while paying insufficient attention to the logical quality of the test cases. Even if a test case can run normally, it may only call the function under test without effective assertions, or it may only cover the regular path while omitting boundary inputs, abnormal branches, and key behavioral differences. Relying solely on executability or coverage makes it difficult to accurately determine whether a test case truly has behavioral verification capabilities.
[0008] (2) Existing methods lack a multi-dimensional evaluation mechanism for test logic quality. Unit test quality is usually affected by multiple factors such as assertion validity, boundary condition coverage, exception path coverage, branch differentiation ability, defect disclosure ability, and execution stability. Existing generation methods often only screen based on execution results or simple coverage, making it difficult to identify logical-level quality defects such as weak assertions, test prediction errors, boundary omissions, lack of exception handling, call contract errors, and survival of important variants.
[0009] (3) Existing large-scale model fine-tuning methods, when used for test case generation, typically rely on manually labeled samples, reference answer samples, or directly generated samples for training, lacking a sample selection and preference construction mechanism for test logic quality. For test generation tasks, there are often relative quality differences between different candidate tests, but these differences are not easily represented by a single label or a single score. If unselected test samples are used directly for fine-tuning, the model may learn superficial features such as format, length, and number of assertions, but cannot stably learn real test logic capabilities such as boundary checks, abnormal path handling, and defect disclosure.
[0010] (4) Although existing preference learning or pairwise sample training methods can guide the model to generate outputs that better meet the target through superior and inferior sample pairs, there are still difficulties in constructing preference sample pairs in the code test case generation scenario. If the quality of positive samples is insufficient, the model will have difficulty obtaining clear test logic demonstrations; if negative samples only show syntax errors or non-executable errors, the model may only learn low-level format differences; if the differences between positive and negative samples are too large in length, structure, number of assertions or code complexity, the model may learn based on surface differences rather than based on the quality of test logic.
[0011] (5) Existing technologies lack effective handling of the relationship between the strength of generated test cases and execution reliability. Stronger test cases usually contain more assertions, more complex input combinations, or more abnormal path checks, but such tests may also be more prone to problems such as call contract mismatch, incorrect abnormal assumptions, or execution failure. Existing methods usually tend to improve coverage or generation strength in isolation, lacking a mechanism for joint evaluation and risk warning of test logic strength and execution stability, making it difficult to meet the common requirements of test effectiveness and reliability in real-world software testing scenarios.
[0012] Therefore, there is an urgent need for a technical solution for generating code test cases that can perform multi-dimensional logical quality assessment of candidate test cases, select high-quality positive samples and negative samples with diagnostic value based on the test logic quality, construct preferred sample pairs while controlling for differences in surface features, and use these preferred sample pairs to fine-tune the large language model in a targeted manner. This would enable the model to focus more on assertion validity, boundary coverage, abnormal path coverage, and defect disclosure capabilities when generating test cases, while also evaluating the execution stability of the generated test cases, thereby improving the quality and usability of generated code test cases. Summary of the Invention
[0013] To address the shortcomings of existing code test case generation methods based on large language models, such as insufficient evaluation of test logic quality, inadequate fine-tuning sample construction, significant interference from surface differences in preferred samples, and the difficulty in balancing the strength and execution stability of generated test cases, this invention provides a code test case generation system and method based on targeted fine-tuning. This invention improves the quality of generated test cases in terms of assertion validity, boundary coverage, anomaly path coverage, and defect disclosure through techniques such as candidate test case generation, multi-dimensional test logic quality evaluation, positive sample quality screening, negative sample diagnostic construction, controlled pairing based on surface consistency, targeted fine-tuning, and stability evaluation of generated results. It also reduces bias caused by model learning of surface features, making the generated results more suitable for automated software testing and software quality assurance scenarios.
[0014] The objective of this invention is achieved through the following technical solution:
[0015] A code test case generation system based on targeted fine-tuning includes a task data acquisition module, a candidate test generation module, a legality checking module, a test logic quality assessment module, a positive sample screening module, a negative sample construction module, a preferred sample pairing module, a targeted fine-tuning module, a target test generation module, and a result output module, wherein:
[0016] The task data acquisition module is used to acquire the code, signature, description, entry point, and reference test data of the function under test, and organize them into a unified task data format.
[0017] The candidate test generation module is used to construct test generation prompts based on task data and call a large language model to generate multiple candidate test cases.
[0018] The legality check module is used to perform syntax parsing, abstract syntax tree checking, target function call checking, and redefinition checking on candidate test cases or target test cases to filter out test cases that do not meet the requirements.
[0019] The test logic quality assessment module is used to evaluate the execution results, coverage, assertion features, boundary conditions, abnormal paths, mutation tests, and failure types of test cases, and generate multi-dimensional test logic quality records.
[0020] The positive sample screening module is used to calculate the positive sample quality score based on the multidimensional test logic quality record and to screen high-quality positive samples.
[0021] The negative sample construction module is used to filter natural negative samples from naturally generated logical defect samples, or to generate programmed negative samples through controlled perturbation.
[0022] The preferred sample pairing module is used to construct preferred sample pairs from high-quality positive samples and negative samples with diagnostic value, and to perform surface consistency control on the preferred sample pairs.
[0023] The directional fine-tuning module is used to train the large language model based on preference sample pairs, enabling the model to learn generation strategies oriented towards test logic quality.
[0024] The target test generation module is used to call the large language model after targeted fine-tuning and generate target test cases based on the new code under test task.
[0025] The result output module is used to output the target test case code and its quality assessment results, and generate risk warnings based on the strength of the test logic and the stability of execution.
[0026] A method for generating code test cases based on targeted fine-tuning using the above system includes the following steps:
[0027] Step S1: Use the task data acquisition module to obtain the task data of the code under test, construct test cases to generate prompt information, and use it as the basis for candidate test case generation, test logic quality assessment and model-oriented fine-tuning.
[0028] Step S2: The candidate test generation module generates candidate test cases based on the test code task data.
[0029] Test prompts are constructed based on the code, signature, and task description of the function under test. The prompts are then input into a large language model to generate multiple candidate test cases.
[0030] Step S3: Use the validity check module to check the validity of the candidate test cases:
[0031] Syntax parsing, abstract syntax tree checking, function call checking, and target function redefinition checking are performed on candidate test cases. Candidate test cases with syntax errors, those that cannot be parsed, those that redefine the target function, those that do not call the target function, or those that obviously do not meet the test format requirements are filtered out. Candidate test cases that pass the legality check enter the multi-dimensional test logic quality assessment process.
[0032] Step S4: Use the test logic quality assessment module to perform a multi-dimensional test logic quality assessment on the candidate test cases:
[0033] Multidimensional test logic quality assessment includes execution result assessment, coverage assessment, assertion feature assessment, boundary condition assessment, abnormal path assessment, mutation test assessment, and failure type assessment.
[0034] Step S5: The positive sample screening module screens high-quality positive samples based on the multi-dimensional test logic quality assessment results.
[0035] The positive sample quality score is calculated based on the mutation score, coverage, assertion features, boundary condition coverage, abnormal path coverage, number of failure types, and execution stability of the candidate test cases.
[0036] Step S6: The negative sample construction module constructs negative samples with diagnostic value based on the logical defects of the candidate test cases.
[0037] Step S7: Construct surface-consistent controlled preference sample pairs using the preference sample pairing module:
[0038] High-quality positive samples under the same test task are used as preferred test cases, and negative samples with diagnostic value are used as inferior test cases. Preferred sample pairs are constructed, and surface consistency control is performed on preferred and inferior test cases when constructing preferred sample pairs.
[0039] Step S8: The targeted fine-tuning module uses preference samples to perform targeted fine-tuning on the large language model.
[0040] The test code task data is used as input, and the preferred and unpreferred test cases are used as paired training samples to fine-tune the large language model.
[0041] Step S9: The target test generation module generates target test cases using the fine-tuned model.
[0042] The system receives new code under test task data, constructs target test generation prompts based on the code under test, function signature, task description, and entry point, inputs the target test generation prompts into the large language model after targeted fine-tuning, and generates target test cases. After generating target test cases, the system performs the same legality checks and multi-dimensional test logic quality evaluations as the candidate test cases, and obtains the execution status, coverage, assertion features, boundary condition coverage, abnormal path coverage, mutation score, and failure type records of the target test cases.
[0043] Step S10: The result output module outputs the target test cases and their quality assessment results.
[0044] Output the target test case code and the corresponding test logic quality assessment results.
[0045] This invention addresses the problems of insufficient logical quality evaluation, inaccurate sample construction, significant interference from surface features, and the difficulty in balancing test intensity and execution stability in existing methods for generating code test cases for large language models. It provides an executable, evaluable, trainable, and scalable code test case generation scheme. Compared to existing technologies, this invention has the following advantages:
[0046] (1) It can improve the completeness of code test case quality evaluation. Existing code test case generation methods usually use syntax correctness, execution pass rate or code coverage as the main evaluation criteria, which are difficult to reflect whether the test cases have the ability to verify real behavior. This invention evaluates test cases together by indicators such as execution results, coverage, mutation score, assertion features, boundary condition coverage, abnormal path coverage and failure type classification, so that the quality evaluation of test cases is no longer limited to whether it can run, but can further reflect whether the assertions are effective, whether the boundaries are covered, whether abnormal paths are considered and whether the test has the ability to reveal defects.
[0047] (2) It can improve the accuracy of selecting high-quality training samples. This invention selects positive samples based on the multi-dimensional test logic quality evaluation results, and uses test cases with stable execution, effective assertions, sufficient boundary coverage, reasonable handling of abnormal paths, high mutation scores and few failure logs as high-quality positive samples. This processing method can reduce the probability of low-quality test samples entering the fine-tuning training process, so that the model can obtain clearer test logic demonstrations, thereby improving the model's learning effect on test predictions, boundary inputs and anomaly checks.
[0048] (3) It can improve the diagnostic value of negative samples. In existing fine-tuning training, negative samples often manifest as syntax errors, formatting errors, or completely unexecutable samples, and the model tends to learn only low-level formal differences. The negative samples constructed in this invention include naturally generated logical defect samples and procedurally generated negative samples through controlled perturbations. These negative samples can retain the basic structure of the test code while exhibiting identifiable defects in assertions, boundary conditions, abnormal paths, test predictions, or defect disclosure capabilities. In this way, negative samples can more effectively expose weaknesses in the test logic, increasing the density of diagnostic information in preference training.
[0049] (4) It can reduce the bias caused by the model learning surface features. When constructing preferred sample pairs, this invention performs surface consistency control on the code length, number of test functions, number of assertions, assertion density, comment ratio, and structural complexity of preferred and unpredictable test cases. When the surface difference of a sample pair exceeds a preset threshold, the system removes or rematches the sample pair. This processing method can reduce the risk of the model learning preferred relationships based on surface features such as length, format, and number of assertions, and make the model more inclined to learn test logic differences such as assertion validity, boundary coverage, abnormal path coverage, and defect disclosure capabilities.
[0050] (5) Improves the targeting of large language model fine-tuning. Existing large model fine-tuning methods mostly use general samples or common reference answers for training, which are difficult to directly address the logical quality goals in code test case generation. This invention uses high-quality positive samples and diagnostically valuable negative samples to construct paired training data, and through targeted fine-tuning, improves the probability of generating preferred test logic patterns and reduces the probability of generating inferior logical defect patterns. This approach enables the model training objective to match the test logic quality requirements in the code test case generation task.
[0051] (6) It can balance the logical strength and execution stability of test cases. After generating target test cases, this invention jointly evaluates their execution stability, assertion strength, boundary coverage, abnormal path coverage, and mutation score. When the generated test cases have high logical strength but a high risk of execution failure, the system outputs a stability risk warning; when the generated test cases can execute stably but the assertions are weak or the defect disclosure capability is insufficient, the system outputs a logical strength deficiency warning. This processing method can help users identify quality risks in test cases and provide a basis for subsequent screening, correction, or regeneration.
[0052] (7) It can improve the automation level of the code test case generation process. This invention integrates task data acquisition, candidate test case generation, legality checking, multi-dimensional quality assessment, positive sample screening, negative sample construction, preferred sample pairing, model-oriented fine-tuning, target test generation, and result output into a continuous processing flow. This process can reduce the workload of manually writing test cases, manually screening training samples, and manually judging test quality, and improve the processing efficiency of code test case generation, training sample construction, and quality assessment.
[0053] (8) It can generate traceable, statistical, and verifiable quality evaluation results. In one implementation, the system can construct a candidate test case pool based on function-level code tasks, evaluate the execution, coverage, mutation testing, assertion features, and failure types of the candidate test cases, and construct preferred sample pairs from the available candidate samples. In experimental applications, the system can generate candidate test cases for batch function-level tasks and select samples that can be used for quality analysis and preference training, forming a traceable, statistical, and verifiable data foundation. In this way, the present invention can provide stable data support for model training, test case selection, and generation effect evaluation.
[0054] (9) Improves the applicability and scalability of the system. The method described in this invention does not rely on a fixed type of function under test or a fixed testing framework. In actual implementation, it can be adapted according to the software project language, testing framework, and quality indicators. For Python function-level tasks, the pytest testing format can be used; for other programming languages or testing environments, the legality check, execution evaluation, coverage statistics, and mutation testing modules can be replaced with implementations under the corresponding language or toolchain. This structure facilitates deployment and expansion in different code testing scenarios. Attached Figure Description
[0055] Figure 1 This is a flowchart illustrating the overall process of generating code test cases based on targeted fine-tuning.
[0056] Figure 2 A schematic diagram of the module structure of the system for generating code test cases based on targeted fine-tuning;
[0057] Figure 3 This is a schematic diagram of the multidimensional test logic quality assessment process;
[0058] Figure 4 A schematic diagram illustrating the process of positive sample screening, negative sample construction, and preference sample pairing.
[0059] Figure 5 This is a schematic diagram of a surface-consistency controlled pairing process;
[0060] Figure 6 A schematic diagram of the process for generating target test cases and assessing stability risks. Detailed Implementation
[0061] The technical solution of the present invention will be further described below with reference to the accompanying drawings, but it is not limited thereto. Any modifications or equivalent substitutions to the technical solution of the present invention that do not depart from the spirit and scope of the technical solution of the present invention should be covered within the protection scope of the present invention.
[0062] This invention provides a code test case generation system based on targeted fine-tuning. The system includes a task data acquisition module, a candidate test generation module, a legality checking module, a test logic quality assessment module, a positive sample screening module, a negative sample construction module, a preferred sample pairing module, a targeted fine-tuning module, a target test generation module, and a result output module, wherein:
[0063] The task data acquisition module is used to acquire the code, signature, description, entry point, and reference test data of the function under test, and organize them into a unified task data format.
[0064] The candidate test generation module is used to construct test generation prompts based on task data and call a large language model to generate multiple candidate test cases.
[0065] The legality check module is used to perform syntax parsing, abstract syntax tree checking, target function call checking, and redefinition checking on candidate test cases or target test cases to filter out test cases that do not meet the requirements.
[0066] The test logic quality assessment module is used to evaluate the execution results, coverage, assertion features, boundary conditions, abnormal paths, mutation tests, and failure types of test cases, and generate multi-dimensional test logic quality records.
[0067] The positive sample screening module is used to calculate the positive sample quality score based on the multidimensional test logic quality record and to screen high-quality positive samples.
[0068] The negative sample construction module is used to filter natural negative samples from naturally generated logical defect samples, or to generate programmed negative samples through controlled perturbation.
[0069] The preferred sample pairing module is used to construct preferred sample pairs from high-quality positive samples and negative samples with diagnostic value, and to perform surface consistency control on the preferred sample pairs.
[0070] The directional fine-tuning module is used to train the large language model based on preference sample pairs, enabling the model to learn generation strategies oriented towards test logic quality.
[0071] The target test generation module is used to call the large language model after targeted fine-tuning and generate target test cases based on the new code under test task.
[0072] The result output module is used to output the target test case code and its quality assessment results, and generate risk warnings based on the strength of the test logic and the stability of execution.
[0073] The present invention also provides a method for generating code test cases based on targeted fine-tuning implemented by the above system, the method comprising the following steps:
[0074] Step S1: Use the task data acquisition module to obtain the task data of the code under test, construct test cases to generate prompt information, and use it as the basis for candidate test case generation, test logic quality assessment and model-oriented fine-tuning.
[0075] The test code task data includes one or more of the following: the test function code, function signature, function name, entry point, natural language description, reference test case, reference implementation, input / output constraints, and task source information.
[0076] In one implementation, the code under test task data is Python function-level task data, and each task includes at least a task identifier, function signature, description of the function under test, and entry function name. The system performs format validation on the task data, deletes tasks that cannot be parsed or lack necessary fields, and writes the retained tasks into a unified task pool.
[0077] Step S2: The candidate test generation module generates candidate test cases based on the test code task data.
[0078] Test prompts are constructed based on the code, signature, and task description of the function under test. These prompts are then input into a large language model to generate multiple candidate test cases. The candidate test cases are executable test code, preferably in pytest format.
[0079] When generating candidate test cases, constraints are imposed on the prompts, requiring the test code to call the function under test and include checks on actual behavior; redefining, overriding, wrapping, or replacing the function under test in the test code is prohibited; copying reference implementations is prohibited; and generating call methods inconsistent with the interface of the function under test is prohibited. After candidate test cases are generated, the system records the task identifier, generation model, generation round, decoding method, and candidate number for each candidate test case.
[0080] Step S3: Use the legality check module to perform legality checks on the candidate test cases.
[0081] Candidate test cases undergo syntax parsing, abstract syntax tree checking, function call checking, and target function redefinition checking to eliminate those with syntax errors, those that cannot be parsed, those that redefine the target function, those that do not call the target function, or those that clearly do not meet the test format requirements. Candidate test cases that pass the validity check proceed to the multi-dimensional test logic quality assessment process.
[0082] Step S4: Use the test logic quality assessment module to perform multi-dimensional test logic quality assessment on candidate test cases.
[0083] In this step, the multidimensional test logic quality assessment includes execution result assessment, coverage assessment, assertion feature assessment, boundary condition assessment, anomaly path assessment, mutation test assessment, and failure type assessment, among which:
[0084] The execution result evaluation is used to determine whether the candidate test cases can be run in the preset test environment, and records the status such as execution success, execution failure, timeout, runtime exception and call contract error;
[0085] The coverage evaluation is used to calculate the line coverage and branch coverage of the candidate test cases to the code under test;
[0086] The assertion feature evaluation is used to statistically analyze the number of assertions, assertion density, number of test functions, average number of assertions per test function, number of exception checks, and number of pytest.raises usages in candidate test cases.
[0087] The boundary condition evaluation is used to determine whether candidate test cases cover zero values, null values, extreme values, duplicate values, single elements, multiple elements, special strings, abnormal inputs, or other boundary scenarios related to the semantics of the function under test.
[0088] The abnormal path evaluation is used to determine whether the candidate test cases cover the abnormal input, abnormal return, error branch or abnormal handling path that the function under test may trigger;
[0089] The mutation test evaluation is used to generate one or more mutants in the code under test, and to determine whether the candidate test cases can kill the mutants by the execution results of the candidate test cases, thereby obtaining the defect disclosure capability index of the candidate test cases.
[0090] The failure type assessment is used to record specific quality defects in candidate test cases. These quality defects include one or more of the following: weak or missing assertions, call contract errors, abnormal path gaps, boundary condition gaps, branch differentiation gaps, survival of important variants, execution timeouts, and redefinition of the objective function.
[0091] In one implementation, candidate test cases Variation score It can be represented as:
[0092] in, Indicates candidate test cases The number of mutants that can be killed. This indicates the total number of variants generated for the tested code. If If the mutation score is 0, the system can skip the mutation score calculation or mark the task as a non-evaluable mutation test task. A higher mutation score indicates a stronger ability of the candidate test cases to identify changes in the behavior of the code under test.
[0093] In one implementation, the system calculates a comprehensive test logic quality score for candidate test cases based on execution results, coverage, mutation score, assertion features, boundary condition coverage, abnormal path coverage, and failure type penalty items. For candidate test cases... Its overall test logic quality score It can be represented as:
[0094]
[0095] in, This indicates the score for the execution result. Indicates the coverage score. Indicates the mutation score. Indicates the assertion feature score, Indicates the boundary condition coverage score. Indicates the score for abnormal path coverage. Indicates the penalty for failure type. to This represents the weighting coefficient of the corresponding indicator. The system can preset or adjust these weighting coefficients based on the test task type, the language of the code being tested, the test framework, or the application scenario. Each indicator can be normalized before calculation to ensure that quality indicators of different dimensions are within the same or similar numerical range.
[0096] Step S5: The positive sample screening module screens high-quality positive samples based on the multi-dimensional test logic quality assessment results.
[0097] Positive sample quality scores are calculated based on the candidate test cases' mutation score, coverage, assertion features, boundary condition coverage, abnormal path coverage, number of failure types, and execution stability. A higher positive sample quality score indicates that the candidate test case is more suitable as a demonstration sample for the model's learning test logic.
[0098] In one implementation, the positive sample quality score The overall score can be based on the test logic quality. and execution stability score The calculation yielded:
[0099] in, Indicates candidate test cases Positive sample quality score, This represents the overall score for the quality of the test logic. This indicates the performance stability score. This represents a balance coefficient between quality strength and execution stability. This method allows the system to avoid selecting positive samples based solely on a single coverage rate or a single execution result.
[0100] In one implementation, candidate test cases that pass execution, contain valid assertions, cover boundary scenarios, cover abnormal paths, have mutation scores that meet a preset mutation score threshold, and have a number of failure types lower than a preset failure number threshold are identified as high-quality positive samples. The mutation score is determined based on the ratio of the number of mutants killed by the candidate test case to the total number of mutants; the number of failure types is determined based on the number of defect records for the candidate test case regarding execution failure, runtime anomalies, call contract errors, weak assertions, boundary condition gaps, abnormal path gaps, and the survival of important mutants.
[0101] When multiple candidate test cases exist for the same task under test, the system sorts the candidate test cases according to the positive sample quality score and prioritizes those with positive sample quality scores within a preset proportion range or above a preset quality score threshold. The positive sample quality score can be calculated by weighting the overall test logic quality score and the execution stability score. The preset mutation score threshold, preset failure number threshold, preset proportion range, and preset quality score threshold can be set or adjusted according to the code language being tested, the test framework, the task difficulty, or the size of the training data.
[0102] Step S6: The negative sample construction module constructs negative samples with diagnostic value based on the logical defects of the candidate test cases.
[0103] In this step, negative samples include natural negative samples and programmed negative samples, wherein:
[0104] The natural negative samples are low-quality logical samples that naturally appear during the candidate test case generation process. These low-quality logical samples include test cases with missing assertions, invalid assertions, missing boundary inputs, incorrect exception path handling, test oracle errors, call contract errors, or significant variant survival.
[0105] The programmatic negative samples are generated from high-quality or intermediate-quality test cases through controlled perturbations. These controlled perturbations include one or more of the following: weakening assertions, removing critical boundary inputs, removing anomaly checks, replacing expected outputs, reducing critical branch tests, and retaining the test structure but reducing the validity of test predictions. The negative samples generated through these controlled perturbations still superficially possess a complete test code structure, but have identifiable defects in their test logic, serving to provide diagnostic training signals to the model.
[0106] Step S7: Construct surface-consistent controlled preference sample pairs using the preference sample pairing module.
[0107] High-quality positive samples under the same test task are selected as preferred test cases, and negative samples with diagnostic value are selected as inferior test cases, thus constructing preferred sample pairs. When constructing preferred sample pairs, surface consistency control is applied to the preferred and inferior test cases. This surface consistency control includes code length control, test function number control, assertion number control, assertion density control, comment ratio control, code structure complexity control, and input format control. If the difference in surface characteristics between the preferred and inferior test cases exceeds a preset threshold, the preferred sample pair is discarded or re-matched, ensuring that the retained preferred sample pairs primarily reflect differences in test logic quality, rather than differences in code length, format, or structural complexity.
[0108] In one implementation, preferred test cases are... With inferior test cases Surface uniformity distance between It can be represented as:
[0109] in, Indicates the preferred test cases. Indicates the inferior test case; Indicates the first A surface feature, the surface feature including one or more of the following: code length, number of test functions, number of assertions, assertion density, comment ratio, structural complexity, and input format; Indicates the first The weights corresponding to each surface feature; This indicates the number of features involved in surface consistency control.
[0110] When the following conditions are met:
[0111] If the preferred test case and the unpreferred test case are not met, the preferred sample pair is retained; if the above conditions are not met, the sample pair is discarded or returned to the same task sample matching stage for re-pairing. This represents the surface consistency distance threshold.
[0112] Step S8: The targeted fine-tuning module uses preference samples to perform targeted fine-tuning on the large language model.
[0113] The test code task data is used as input, and the preferred and unpreferred test cases are used as paired training samples to perform targeted fine-tuning on the large language model. During targeted fine-tuning, the model is made to increase the probability of generating test logic patterns similar to the preferred test cases and decrease the probability of generating logical defect patterns similar to the unpreferred test cases.
[0114] In one implementation, the targeted fine-tuning employs a direct preference optimization method. Training samples include task prompts, preferred test cases, and disliked test cases. The model learns assertion design, boundary checks, outlier handling, and defect disclosure capabilities through pairwise preference relationships, rather than simply learning test code format or length characteristics.
[0115] In one implementation, for the code under test task input ( ), Optimal test cases ( ) and inferior test cases ( ), preference training loss function It can be represented as:
[0116]
[0117] in, Indicates the model to be trained. Represents the reference model. This represents the parameter controlling the intensity of preference. This represents the Sigmoid function. Through this training objective, the model, given the same test code task input, increases the probability of generating optimal test logic patterns and decreases the probability of generating suboptimal logic patterns.
[0118] Step S9: The target test generation module generates target test cases using the model after targeted fine-tuning.
[0119] Receive new code under test task data, construct target test generation prompts based on the code under test, function signature, task description and entry point, input the target test generation prompts into the large language model after targeted fine-tuning, and generate target test cases.
[0120] After generating target test cases, the system performs the same legality checks and multi-dimensional test logic quality assessments on the target test cases as on the candidate test cases, and obtains the execution status, coverage, assertion features, boundary condition coverage, abnormal path coverage, mutation score, and failure type records of the target test cases.
[0121] Step S10: The result output module outputs the target test cases and their quality assessment results.
[0122] The system outputs the target test case code and the corresponding test logic quality assessment results, including whether the test case passed execution, the number of assertions, assertion density, coverage, boundary condition coverage, abnormal path coverage, mutation score, failure type, and execution stability risk warning.
[0123] In one implementation, the system can generate risk warnings based on the logic strength score and execution stability score of the target test case. If the target test case has a high logic strength score but low execution stability, the system generates a stability risk warning; if the target test case is stable in execution but has weak assertions, insufficient boundary coverage, insufficient abnormal path coverage, or a low mutation score, the system generates a logic strength deficiency warning. This allows users to filter, modify, or regenerate test cases based on the quality assessment results.
[0124] Through the above technical solution, this invention combines test case generation, test logic quality assessment, sample quality screening, negative sample diagnostic construction, surface consistency controlled pairing, and large language model targeted fine-tuning into a complete technical process. This enables the model-generated test cases to no longer solely focus on executability or coverage, but simultaneously consider assertion validity, boundary coverage, abnormal path coverage, defect disclosure capability, and execution stability. Furthermore, it solves the following technical problems:
[0125] 1. Existing code test case generation methods based on large language models usually focus on the syntactic correctness, executability, or code coverage of the test code, lacking a targeted generation mechanism for the quality of test logic. As a result, although the generated test cases may be able to run, they may have problems such as insufficient assertions, missing boundary inputs, insufficient coverage of abnormal paths, weak test predictions, and insufficient defect disclosure capabilities, making it difficult to reliably meet the requirements for test effectiveness in software quality assurance scenarios.
[0126] 2. The lack of a multi-dimensional quality evaluation mechanism in existing test case generation methods. Existing methods typically screen generated tests based on execution results, coverage, or human experience. It is difficult to uniformly measure assertion validity, boundary condition coverage, abnormal path coverage, branch differentiation ability, mutant killing ability, and execution stability. It is also difficult to accurately identify logical defects in test cases, such as call contract errors, abnormal path gaps, boundary condition gaps, weak assertions, and survival of important mutants.
[0127] 3. The existing large-scale model fine-tuning sample construction methods are mismatched with the code test case generation task. Existing fine-tuning methods mostly use ordinary reference samples, manually labeled samples, or generated samples without fine-grained screening. They lack a mechanism to screen high-quality positive samples and negative samples with diagnostic value based on the quality of test logic. This makes it easy for the model to learn superficial features such as the format, length, and number of assertions of the test code, but it cannot stably learn test logic capabilities such as assertion design, boundary checks, exception path handling, and defect disclosure.
[0128] 4. The existing preference samples suffer from significant interference from surface differences during the construction process. For code test case generation tasks, if the preferred and unpredictable test cases differ greatly in code length, number of test functions, number of assertions, comment ratio, or structural complexity, the model may judge the quality of the samples based on surface features rather than learning the generation strategy based on the quality of the test logic. Therefore, a preference sample construction method is needed that can control the surface consistency of positive and negative samples while highlighting the differences in test logic.
[0129] 5. The challenge of balancing logical strength and execution stability in test case generation. Stronger test cases may include more assertions, boundary conditions, and exception checks, but they may also introduce a higher risk of execution failure. Existing methods lack a mechanism for jointly evaluating the strength of test logic and execution reliability, making it difficult to identify stability risks in generated test cases in a timely manner, and also hindering the provision of a basis for subsequent test modification, selection, and application.
[0130] Example:
[0131] This embodiment provides a method for generating code test cases based on targeted fine-tuning, such as... Figure 1 As shown, the method includes three stages: task and candidate test preparation, test logic quality modeling and preference sample construction, and targeted fine-tuning and target test generation.
[0132] The system receives the test code task data. This test code task data includes the test function code, function signature, function name, entry point, task description, reference test data, reference implementation, and execution environment information. For Python function-level code tasks, the task data can be stored in JSON, JSONL, CSV, or database record format. Each task record must include at least a task identifier, function signature, entry function name, and natural language description.
[0133] After acquiring the task data, the task data acquisition module 100 performs a completeness check on the task fields. If a task lacks an entry function name, function signature, or executable code, it is marked as an invalid task and will not enter the candidate test generation process. If the task fields are complete, it is written to the unified task pool. The unified task pool is used to support candidate test case generation, test logic quality assessment, sample screening, and subsequent target test generation.
[0134] The candidate test generation module 200 constructs test generation prompts based on the task data of the code under test. The test generation prompts include the code of the function under test, the function signature, the task description, the entry point, the test framework requirements, and generation constraints. The generation constraints include: the test code should call the function under test; the test code should include real-world behavior checks; the test code must not redefine, rewrite, wrap, or replace the function under test; the test code must not copy the reference implementation; and the test code must not use a calling method inconsistent with the function signature.
[0135] The candidate test generation module 200 inputs test generation prompts into the large language model to generate multiple candidate test cases. In one implementation, the candidate test cases adopt the pytest format, and each candidate test case includes one or more test functions. The system records the task identifier, generation model, generation round, decoding method, candidate number, and generation time for each candidate test case, forming a candidate test case pool.
[0136] The legality checking module 300 performs legality checks on candidate test cases. These checks include syntax parsing, abstract syntax tree (AST) checking, target function call checking, and target function redefinition checking. Syntax parsing determines whether the test code can be parsed by the interpreter; AST checking identifies the test code structure; target function call checking determines whether the test case calls the function under test; and target function redefinition checking determines whether the test code redefines, overrides, or replaces the function under test. If a candidate test case fails the legality check, it is eliminated; if it passes, it proceeds to the test logic quality assessment process.
[0137] Reference Figure 3The test logic quality assessment module 400 performs multi-dimensional test logic quality assessment on candidate test cases. This module includes an execution result assessment unit 401, a coverage assessment unit 402, an assertion feature assessment unit 403, a boundary condition assessment unit 404, an anomaly path assessment unit 405, a mutation test assessment unit 406, a failure type assessment unit 407, and a test logic quality record generation unit 408, wherein:
[0138] The execution result evaluation unit 401 runs candidate test cases in a preset execution environment and records execution statuses such as execution success, execution failure, execution timeout, runtime exception, and call contract error.
[0139] The coverage evaluation unit 402 calculates the line coverage and branch coverage of the candidate test cases on the code under test.
[0140] The assertion feature evaluation unit 403 statistically analyzes the number of assertions, assertion density, number of test functions, average number of assertions per test function, number of anomaly checks, and number of times pytest.raises is used in candidate test cases.
[0141] The boundary condition evaluation unit 404 identifies boundary scenarios based on the task description of the function under test, the input type, and the reference test data, and determines whether the candidate test cases cover zero values, null values, extreme values, single elements, multiple elements, duplicate values, special strings, illegal inputs, or other boundary inputs related to the semantics of the function under test.
[0142] The abnormal path evaluation unit 405 determines whether the candidate test cases cover abnormal input, abnormal return, error branch, or abnormal handling path.
[0143] The mutation testing and evaluation unit 406 generates at least one mutant based on the code under test and executes the mutant using candidate test cases. If the candidate test cases can produce a detectable difference between the mutant's execution result and the original program behavior, the mutant is determined to be killed; if the candidate test cases cannot distinguish between the mutant and the original program, the mutant is determined to be alive. A mutation score is calculated based on the number of killed mutants and the total number of mutants, which characterizes the defect disclosure capability of the candidate test cases.
[0144] The failure type evaluation unit 407 classifies the quality defects in the candidate test cases. The failure types include one or more of the following: weak or missing assertions, call contract errors, abnormal path gaps, boundary condition gaps, branch differentiation gaps, survival of important variants, execution timeouts, syntax errors, and redefinition of objective functions.
[0145] The test logic quality record generation unit 408 summarizes the above evaluation results to form a test logic quality record. The test logic quality record includes execution status, coverage, assertion features, boundary condition coverage, abnormal path coverage, mutation score, failure type, and risk marker. The test logic quality record generation unit 408 calculates the comprehensive test logic quality score of candidate test cases based on execution results, coverage, mutation score, assertion features, boundary condition coverage, abnormal path coverage, and failure type penalty items.
[0146] Reference Figure 4 The positive sample screening module 500 filters high-quality positive samples based on the test logic quality records. The positive sample quality score calculation unit 501 calculates the positive sample quality score based on execution stability, mutation score, assertion features, boundary condition coverage, abnormal path coverage, and the number of failure types. The positive sample quality score can be calculated using a weighted method, where mutation score, boundary coverage, abnormal path coverage, assertion density, and execution pass status are positive indicators, and the number of failure types, execution timeout, call contract error, and weak assertions are negative indicators. Candidate test cases enter the high-quality positive sample pool 502 when they meet the following conditions: the candidate test case can be executed successfully; the candidate test case contains valid assertions; the candidate test case covers at least one normal input scenario; the candidate test case covers at least one of boundary input, abnormal path, or critical branch; the mutation score of the candidate test case is higher than a preset threshold; and the number of failure types of the candidate test case is lower than a preset threshold. If there are multiple candidate test cases that meet the conditions for the same test task, the system sorts them according to the positive sample quality score from high to low and prioritizes the samples with higher scores to enter the high-quality positive sample pool 502.
[0147] The negative sample construction module 600 is used to construct negative samples with diagnostic value. The natural negative sample screening unit 601 filters naturally generated logical defect samples from the candidate test case pool. These natural negative samples include weak assertion samples, boundary omission samples, abnormal path gap samples, call contract error samples, test prediction error samples, and important variant survival samples. Natural negative samples retain typical logical defects that may occur during actual test generation of the large language model, reflecting the model's true weaknesses in the test case generation task. The programmatic negative sample construction unit 602 constructs programmatic negative samples based on existing test cases. Construction methods include weakening assertions, deleting key boundary inputs, deleting exception checks, replacing expected outputs, reducing key branch tests, and retaining the test structure but reducing the validity of test predictions—one or more of these methods. For example, for test cases that originally asserted that the objective function returned True or False, some strong assertions can be replaced with weak assertions such as "the result is not empty"; for test cases containing abnormal path checks, exception check statements can be deleted; for test cases covering boundary inputs, boundary input-related test functions can be deleted. The procedural negative samples obtained after the above processing still superficially possess a complete test code structure, but their test logic quality is reduced. A diagnostic negative sample pool 603 is used to aggregate natural negative samples and procedural negative samples. Each negative sample corresponds to at least one defect record, which describes the type of logical problem present in the negative sample, such as weak assertions, anomalous path gaps, boundary condition gaps, or variant survival. This defect record is used for subsequent sample pairing and preference training.
[0148] Reference Figure 4 and Figure 5The preferred sample pairing module 700 constructs preferred sample pairs from high-quality positive samples and diagnostic negative samples. The same-task sample matching unit 701 matches positive and negative samples under the same tested task, ensuring that the positive and negative samples correspond to the same tested function, entry point, and task description. This matching method avoids mistaking differences in task difficulty or function semantics as differences in sample quality. The feature extraction process calculates the surface features of the positive and negative samples to be paired. These surface features include code length, number of test functions, number of assertions, assertion density, comment ratio, structural complexity, and input format. Code length can be represented by the number of tokens or characters; the number of test functions can be statistically analyzed using an abstract syntax tree; the number of assertions can be obtained by identifying assert statements, exception assertions, or test framework assertion methods; assertion density can be calculated from the number of assertions and code length or number of test functions; structural complexity can be calculated from the number of branches, number of loops, nesting levels, or test function structure. The surface consistency control unit 702 filters the samples to be paired item by item. Code length filtering unit 711 compares the code length differences between positive and negative samples. If the difference exceeds a preset threshold, the sample pair is removed or re-matched. Test function number filtering unit 712 compares the number of test functions. If the difference in the number of test functions between positive and negative samples is too large, the sample pair is not retained. Assertion number filtering unit 713 compares the number of assertions, assertion density filtering unit 714 compares the assertion density, comment ratio filtering unit 715 compares the comment ratio, and structural complexity filtering unit 716 compares the complexity of branching, nesting, or calling structures. After the above filtering, sample pairs with similar surface features and whose main differences lie in the test logic quality level are retained. The pairing determination process makes a judgment based on the surface consistency filtering results. If the positive and negative samples meet the surface consistency threshold, the controlled preference sample pair output unit 717 outputs a controlled preference sample pair in the form of chosen / rejected, where chosen is a high-quality positive sample and rejected is a diagnostic negative sample. If the sample pair does not meet the surface consistency threshold, the system can remove it or return to the same task sample matching stage to reselect negative or positive samples for pairing. The preference sample pair generation unit 703 combines positive and negative samples controlled by surface consistency into preference sample pairs. The balancing sample pair assembly unit 704 performs balanced assembly according to task source, task difficulty, positive sample type, negative sample type, and quality level to avoid preference sample pairs being concentrated in a single task source or a single defect type. The preference sample pair dataset 705 is used for targeted fine-tuning and preference training.
[0149] The targeted fine-tuning module 800 trains a large language model based on a preference sample pair dataset. The training data includes task prompts, chosen test cases, and rejected test cases. The training objective is to increase the probability of the model generating chosen test logic patterns and decrease the probability of generating rejected logic defect patterns. In one implementation, targeted fine-tuning employs a direct preference optimization method. During direct preference optimization training, the reference model remains frozen, and the model to be trained adjusts its parameters based on the chosen / rejected sample pairs, making the model more inclined to generate test cases with valid assertions, sufficient boundary coverage, reasonable anomaly path handling, and strong defect disclosure capabilities under the same task prompts. The targeted fine-tuning module 800 can also employ a combination of supervised fine-tuning and preference training. The system first uses high-quality positive samples to perform supervised fine-tuning on the large language model, enabling the model to acquire basic test generation capabilities; then, it uses controlled preference sample pairs for preference training, allowing the model to further learn the differences in test logic quality between positive and negative samples. The training method can be adjusted according to computing resources, model size, and data size.
[0150] Reference Figure 6 The target test generation module 900 generates target test cases for new code under test tasks. After receiving the new code under test task input, the target test generation unit 901 constructs target test prompts based on the code of the function under test, function signature, task description, and entry point, and calls the fine-tuned large language model to generate target test cases. The target test validity check unit 902 performs syntax parsing, abstract syntax tree checking, target function call checking, and redefinition checking on the target test cases. If a target test case fails the validity check, the system outputs a validity failure flag and can regenerate the target test case. If the target test case passes the validity check, it proceeds to the target test logic quality evaluation unit 903. The target test logic quality evaluation unit 903 performs the same multi-dimensional quality evaluation on the target test cases as on the candidate test cases, generating a target test logic quality record. Supporting evaluation information includes reference test data, quality thresholds, and execution environment. The quality evaluation result output module outputs the execution status, coverage, assertion features, boundary coverage, abnormal path coverage, mutation score, quality score, and risk flag of the target test cases. The logic strength assessment unit 904 determines the logic strength of the target test case based on assertion strength, boundary coverage, exception path coverage, and mutation score. The execution stability assessment unit 905 determines the execution stability of the target test case based on execution pass rate, runtime exceptions, execution timeouts, and call contract errors. The risk warning generation unit 906 jointly assesses the logic strength and execution reliability, and generates corresponding risk warnings.
[0151] Target test cases Logical strength score It can be calculated based on assertion strength, boundary coverage, outlier path coverage, and mutation score; target test cases. Execution stability score It can be calculated based on the execution success status, runtime exceptions, execution timeouts, and contract call errors. When the following conditions are met:
[0152] The system generates a stability risk warning. Among them, Indicates the logic strength threshold. This represents the execution stability threshold.
[0153] When the following conditions are met:
[0154] The system generates a logic strength insufficient warning. If the target test case satisfies:
[0155] The system will then retain the target test case and output a quality pass mark.
[0156] If the target test case has high logical strength but low execution stability, the system generates a stability risk warning, indicating that the test case may have a call contract error, overly strong exception assumptions, or execution failure risk. If the target test case has high execution stability but low logical strength, the system generates a logical strength deficiency warning, indicating that the test case may have weak assertions, insufficient boundary coverage, or insufficient defect disclosure capability. If both the logical strength and execution stability of the target test case meet the preset thresholds, the system retains the test case and outputs a quality pass mark.
[0157] The target test case output unit 907 outputs the target test code, quality assessment results, and risk warnings. The output includes the target test cases, test logic quality records, quality score, failure type, stability risk warnings, and logic strength warnings. Users can directly use the target test cases based on the output results, or choose to regenerate, manually modify, or filter qualified test cases.
[0158] In a specific application scenario, this invention can be used for generating Python function-level unit tests. The system constructs a candidate test case pool using function-level task data and evaluates the candidate test cases based on execution, coverage, assertion features, boundary conditions, outlier paths, mutation testing, and failure types. The system filters high-quality positive samples from the candidate test cases and constructs natural and procedural negative samples. After surface consistency control, the system forms a controlled preference sample pair dataset and uses this dataset to perform targeted fine-tuning of a large language model. For new Python function tasks, the system calls the fine-tuned model to generate pytest format test cases and outputs test logic quality evaluation results and stability risk warnings.
[0159] This invention is not limited to the Python language. For Java, JavaScript, C++, or other programming languages, the task data acquisition module 100, candidate test generation module 200, legality check module 300, and test logic quality assessment module 400 can be replaced with parsers, test frameworks, coverage tools, and mutation testing tools for the corresponding language. For example, Java can use the JUnit test format, and JavaScript can use the Jest or Mocha test format. As long as the system still generates code test cases through multi-dimensional test logic quality assessment, positive and negative sample construction, surface consistency controlled pairing, and targeted fine-tuning, it falls within the scope of this invention.
[0160] In this embodiment, the large language model can be an open-source large model, a general-purpose large language model, a model pre-trained on a software engineering corpus, or a dedicated model adapted to an enterprise codebase. The targeted fine-tuning method can employ direct preference optimization, a combination of supervised fine-tuning and preference training, reward model-assisted preference training, or other training methods capable of relative preference learning using chosen / rejected sample pairs. Changing the model type and training method does not affect the technical essence of this invention in improving the quality of test case generation through test logic quality assessment and controlled preference samples.
[0161] Through the above implementation methods, the present invention can combine candidate generation, quality assessment, sample screening, negative sample construction, preference matching, model fine-tuning, and risk assessment of generation results into a complete closed loop in the code test case generation process. This enables the large language model to not only focus on the form and executability of test code when generating test cases, but also to learn test logic features such as assertion validity, boundary coverage, abnormal path coverage, and defect disclosure capabilities, and to provide prompts on the execution stability of the generated test cases.
[0162] This invention can also be implemented using electronic devices, servers, cloud computing platforms, or local computing devices. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable by the processor. When executed by the processor, the computer program implements the above-described code test case generation method. The method can also be implemented in the form of a computer-readable storage medium.
Claims
1. A code test case generation system based on targeted fine-tuning, characterized in that... The system includes a task data acquisition module, a candidate test generation module, a legality check module, a test logic quality assessment module, a positive sample screening module, a negative sample construction module, a preference sample pairing module, a targeted fine-tuning module, a target test generation module, and a result output module, wherein: The task data acquisition module is used to acquire the code, signature, description, entry point, and reference test data of the function under test, and organize them into a unified task data format. The candidate test generation module is used to construct test generation prompts based on task data and call a large language model to generate multiple candidate test cases. The legality check module is used to perform syntax parsing, abstract syntax tree checking, target function call checking, and redefinition checking on candidate test cases or target test cases to filter out test cases that do not meet the requirements. The test logic quality assessment module is used to evaluate the execution results, coverage, assertion features, boundary conditions, abnormal paths, mutation tests, and failure types of test cases, and generate multi-dimensional test logic quality records. The positive sample screening module is used to calculate the positive sample quality score based on the multidimensional test logic quality record and to screen high-quality positive samples. The negative sample construction module is used to filter natural negative samples from naturally generated logical defect samples, or to generate programmed negative samples through controlled perturbation. The preferred sample pairing module is used to construct preferred sample pairs from high-quality positive samples and negative samples with diagnostic value, and to perform surface consistency control on the preferred sample pairs. The directional fine-tuning module is used to train the large language model based on preference sample pairs, enabling the model to learn generation strategies oriented towards test logic quality. The target test generation module is used to call the large language model after targeted fine-tuning and generate target test cases based on the new code under test task. The result output module is used to output the target test case code and its quality assessment results, and generate risk warnings based on the strength of the test logic and the stability of execution.
2. The code test case generation system based on targeted fine-tuning according to claim 1, characterized in that... The test logic quality assessment module includes an execution result assessment unit, a coverage assessment unit, an assertion feature assessment unit, a boundary condition assessment unit, an anomaly path assessment unit, a mutation test assessment unit, a failure type assessment unit, and a test logic quality record generation unit, wherein: The execution result evaluation unit is responsible for running candidate test cases in a preset execution environment and recording the execution status of successful execution, failed execution, execution timeout, runtime exception, and call contract error. The coverage evaluation unit is responsible for calculating the line coverage and branch coverage of the candidate test cases on the code under test; The assertion feature evaluation unit is responsible for calculating the number of assertions, assertion density, number of test functions, average number of assertions per test function, number of exception checks, and number of times pytest.raises is used in candidate test cases. The boundary condition evaluation unit is responsible for identifying boundary scenarios based on the task description of the function under test, the input type, and the reference test data, and determining whether the candidate test cases cover zero values, null values, extreme values, single elements, multiple elements, duplicate values, special strings, illegal inputs, or other boundary inputs related to the semantics of the function under test. The abnormal path evaluation unit is responsible for determining whether candidate test cases cover abnormal input, abnormal return, error branch, or abnormal handling path. The mutation test evaluation unit is responsible for generating at least one mutant based on the code under test, and executing the mutant using candidate test cases. If the candidate test cases can make the execution result of the mutant produce a detectable difference from the behavior of the original program, the mutant is determined to be killed; if the candidate test cases cannot distinguish the mutant from the original program, the mutant is determined to be alive; the mutation score is calculated based on the number of killed mutants and the total number of mutants, which characterizes the defect disclosure capability of the candidate test cases. The failure type evaluation unit is responsible for classifying quality defects in candidate test cases. Failure types include one or more of the following: weak or missing assertions, call contract errors, abnormal path gaps, boundary condition gaps, branch differentiation gaps, survival of important variants, execution timeouts, syntax errors, and redefinition of objective functions. The test logic quality record generation unit is responsible for summarizing the above evaluation results, forming test logic quality records, and calculating the comprehensive test logic quality score of candidate test cases based on execution results, coverage, mutation score, assertion features, boundary condition coverage, abnormal path coverage, and failure type penalty items.
3. The code test case generation system based on targeted fine-tuning according to claim 1, characterized in that... The positive sample screening module includes a positive sample quality score calculation unit and a high-quality positive sample pool. The positive sample quality score calculation unit is responsible for calculating the positive sample quality score based on execution stability, mutation score, assertion features, boundary condition coverage, abnormal path coverage, and the number of failure types. Candidate test cases enter the high-quality positive sample pool when they meet the following conditions: the candidate test case can be executed successfully; the candidate test case contains valid assertions; the candidate test case covers at least one common input scenario; the candidate test case covers at least one of boundary inputs, abnormal paths, or critical branches; the mutation score of the candidate test case is higher than a preset threshold; and the number of failure types of the candidate test case is lower than a preset threshold.
4. The code test case generation system based on targeted fine-tuning according to claim 1, characterized in that... The negative sample construction module includes a natural negative sample screening unit, a programmed negative sample construction unit, and a diagnostic negative sample pool, wherein: The natural negative sample filtering unit is responsible for filtering naturally generated logical defect samples from the candidate test case pool; The programmatic negative sample construction unit is responsible for constructing programmatic negative samples based on existing test cases; The diagnostic negative sample pool is responsible for collecting both natural negative samples and programmed negative samples.
5. The code test case generation system based on targeted fine-tuning according to claim 1, characterized in that... The preference sample pairing module includes a task-matching unit, a surface consistency control unit, a code length filtering unit, a test function number filtering unit, an assertion number filtering unit, an assertion density filtering unit, an annotation ratio filtering unit, a structural complexity filtering unit, a controlled preference sample pair output unit, a preference sample pair generation unit, a balanced sample pair assembly unit, and a preference sample pair dataset, wherein: The same task sample matching unit is responsible for matching positive and negative samples under the same test task, so that the positive and negative samples correspond to the same test function, entry point and task description. The surface consistency control unit is responsible for filtering the samples to be paired item by item. The code length filtering unit is responsible for comparing the code length difference between positive and negative samples. If the difference exceeds a preset threshold, the sample pair is removed or rematched. The test function number filtering unit is responsible for comparing the number of test functions. If the difference between the number of positive and negative sample test functions is too large, the sample pair will not be retained. The assertion count filtering unit is responsible for comparing the number of assertions; The assertion density filtering unit is responsible for comparing assertion densities; The annotation ratio filtering unit is responsible for comparing the annotation ratios. The structural complexity filtering unit is responsible for comparing the structural complexity of branches, nesting, or invocation. The controlled preference sample pair output unit is responsible for outputting controlled preference sample pairs in the form of chosen / rejected, where chosen are high-quality positive samples and rejected are diagnostic negative samples; The preference sample pair generation unit is responsible for combining positive and negative samples controlled by surface consistency into preference sample pairs; The balanced sample assembly unit is responsible for balancing and assembling samples according to task source, task difficulty, positive sample type, negative sample type, and quality level. The preference samples are responsible for targeted fine-tuning and preference training of the dataset.
6. The code test case generation system based on targeted fine-tuning according to claim 1, characterized in that... The target test generation module includes a target test generation unit, a target test legality checking unit, a target test logic quality assessment unit, a logic strength judgment unit, an execution stability judgment unit, a risk warning generation unit, and a target test case output unit, wherein: The target test generation unit is responsible for constructing target test prompts based on the code of the function under test, the function signature, the task description, and the entry point, and for calling the large language model after targeted fine-tuning to generate target test cases. The target test legality checking unit is responsible for performing syntax parsing, abstract syntax tree checking, target function call checking, and redefinition checking on the target test cases. If the target test case fails the legality check, the system outputs a legality failure flag and can regenerate the target test case. If the target test case passes the legality check, it enters the target test logic quality evaluation unit. The target test logic quality assessment unit is responsible for performing the same multi-dimensional quality assessment on the target test cases as on the candidate test cases, and generating target test logic quality records. The logic strength judgment unit is responsible for judging the logic strength of the target test case based on assertion strength, boundary coverage, abnormal path coverage, and mutation score. The execution stability judgment unit is responsible for judging the execution stability of the target test cases based on the execution pass rate, runtime exceptions, execution timeouts, and calling contract errors. The risk warning generation unit is responsible for jointly judging the logic strength and execution reliability, and generating corresponding risk warnings; The target test case output unit is responsible for outputting the target test code, quality assessment results, and risk warnings.
7. A method for generating code test cases based on targeted fine-tuning using the system described in any one of claims 1-6, characterized in that... The method includes the following steps: Step S1: Use the task data acquisition module to obtain the task data of the code under test, construct test cases to generate prompt information, and use it as the basis for candidate test case generation, test logic quality assessment and model-oriented fine-tuning. Step S2: The candidate test generation module generates candidate test cases based on the test code task data. Test prompts are constructed based on the code, signature, and task description of the function under test. The prompts are then input into a large language model to generate multiple candidate test cases. Step S3: Use the validity check module to check the validity of the candidate test cases: Syntax parsing, abstract syntax tree checking, function call checking, and target function redefinition checking are performed on candidate test cases. Candidate test cases with syntax errors, those that cannot be parsed, those that redefine the target function, those that do not call the target function, or those that obviously do not meet the test format requirements are filtered out. Candidate test cases that pass the legality check enter the multi-dimensional test logic quality assessment process. Step S4: Use the test logic quality assessment module to perform a multi-dimensional test logic quality assessment on the candidate test cases: Multidimensional test logic quality assessment includes execution result assessment, coverage assessment, assertion feature assessment, boundary condition assessment, abnormal path assessment, mutation test assessment, and failure type assessment. Step S5: The positive sample screening module screens high-quality positive samples based on the multi-dimensional test logic quality assessment results. The positive sample quality score is calculated based on the mutation score, coverage, assertion features, boundary condition coverage, abnormal path coverage, number of failure types, and execution stability of the candidate test cases. Step S6: The negative sample construction module constructs negative samples with diagnostic value based on the logical defects of the candidate test cases. Step S7: Construct surface-consistent controlled preference sample pairs using the preference sample pairing module: High-quality positive samples under the same test task are used as preferred test cases, and negative samples with diagnostic value are used as inferior test cases. Preferred sample pairs are constructed, and surface consistency control is performed on preferred and inferior test cases when constructing preferred sample pairs. Step S8: The targeted fine-tuning module uses preference samples to perform targeted fine-tuning on the large language model. The test code task data is used as input, and the preferred and unpreferred test cases are used as paired training samples to fine-tune the large language model. Step S9: The target test generation module generates target test cases using the fine-tuned model. Receive new code under test task data, construct target test generation prompts based on the code under test, function signature, task description and entry point, input the target test generation prompts into the large language model after targeted fine-tuning, and generate target test cases; After generating target test cases, perform the same legality checks and multi-dimensional test logic quality assessments on the target test cases as on the candidate test cases, and obtain the execution status, coverage, assertion features, boundary condition coverage, abnormal path coverage, mutation score, and failure type records of the target test cases; Step S10: The result output module outputs the target test cases and their quality assessment results. Output the target test case code and the corresponding test logic quality assessment results.
8. The code test case generation method based on targeted fine-tuning according to claim 7, characterized in that... In step S5, candidate test cases Variation score Represented as: in, Indicates candidate test cases The number of mutants that can be killed. This indicates the total number of variants generated for the tested code; Positive sample quality score Based on the overall score of test logic quality and execution stability score The calculation yielded: in, Indicates candidate test cases Positive sample quality score, This represents the overall score for the quality of the test logic. This indicates the performance stability score. This represents a balance coefficient between quality strength and performance stability; Overall Test Logic Quality Score Represented as: in, This indicates the score for the execution result. Indicates the coverage score. Indicates the variation score. Indicates the assertion feature score, Indicates the boundary condition coverage score. Indicates the score for abnormal path coverage. Indicates the penalty for failure type. to This represents the weighting coefficient of the corresponding indicator.
9. The code test case generation method based on targeted fine-tuning according to claim 7, characterized in that... In step S7, the test cases are selected. With inferior test cases Surface uniformity distance between Represented as: in, Indicates the preferred test cases. Indicates the inferior test case; Indicates the first One surface feature; Indicates the first The weights corresponding to each surface feature; Indicates the number of features involved in surface consistency control; When the following conditions are met: The preferred sample pair consisting of the preferred test case and the unpreferred test case is retained; if the above conditions are not met, the sample pair is removed or the process returns to the same task sample matching stage for re-pairing. This represents the surface consistency distance threshold.
10. The code test case generation method based on targeted fine-tuning according to claim 7, characterized in that... In step S8, the input of the code under test is... Optimize test cases and inferior selection test cases Preferred training loss function Represented as: in, Indicates the model to be trained. Represents the reference model. This represents the parameter controlling the intensity of preference. This represents the Sigmoid function.
Citation Information
Patent Citations
Method, system and equipment for generating report lecture based on large model and medium
CN120578754A
Double-engine government affair question and answer method based on large model fine tuning and RAG retrieval
CN120632050A
Method for performing control logic decision in complex dialogue based on large model fine tuning and dynamic sample
CN121071100A