Test case evaluation model training and application method, system, equipment and medium
By using domain-specific inference models for pre-training and supervising training in test case generation, evaluation basis is generated, and the problem of traditional test case generation relying on manual experience is solved, and efficient and automated test case evaluation is achieved.
Patent Information
- Application Number
- CN202510658183.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-19
AI Technical Summary
The traditional way of generating test cases depends on manual experience, lack of unified standards and systematic feedback, which makes it difficult to maintain the quality of test cases and low manual review efficiency.
By determining the initial inference model based on the application field, pre-training with knowledge data and labeling error categories and modification suggestions, generating evaluation basis, supervising training of the inference model, and obtaining a qualified test case evaluation model.
It reduces the cost of manual labeling, improves the efficiency and quality of test case evaluation, and realizes automated test case evaluation.
Smart Images

Figure CN120508505A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of deep learning model training, and specifically to a training method, system, device and medium for a test case evaluation model. Background Art
[0002] With the rapid development of automotive technology, users' demands for vehicle quality and reliability are increasing. As a core component of the automotive industry, test case generation has become increasingly important. However, due to the varying levels of test engineers' writing skills, traditional test case generation methods suffer from numerous issues, significantly hindering the efficiency and quality of the testing process. For example, test case design relies heavily on the experience and personal style of the tester, lacking unified standards and specifications. Test case evaluation systems are inadequate, lacking a systematic feedback strategy, resulting in test cases that lack high quality. In various fields (such as automotive testing projects), a large number of test cases are often required, and their quality assessment is typically conducted manually, which is labor-intensive, time-consuming, and inefficient. Summary of the Invention
[0003] In view of this, the present application provides a training method, system, device and medium for a test case evaluation model, aiming to solve or partially solve the problems existing in the background technology.
[0004] A first aspect of the present application provides a method for training a test case evaluation model, the method comprising:
[0005] Determine an initial reasoning model applicable to the application domain to which the test case belongs;
[0006] Pre-training the initial reasoning model using knowledge data from the application domain to obtain a reasoning model that can understand and process natural language text from the application domain;
[0007] determining error categories and modification suggestions for test cases used to train the inference model;
[0008] Inputting the test case and the determined error category and modification suggestion into a non-inference model for processing to determine an evaluation basis for the test case, and labeling the test case with the evaluation basis, the error category, and the modification suggestion, wherein the evaluation basis is the correct thinking process of the inference model on the test case;
[0009] The reasoning model is supervised and trained by completing the labeled test cases to obtain a qualified test case evaluation model.
[0010] Optionally, the test case and the determined error category and modification suggestion are input into a non-inference model for processing to determine an evaluation basis for the test case, including:
[0011] Filling the test case into a first placeholder in a pre-built standard prompt word template, and filling the error category and modification suggestion into a second placeholder in the standard prompt word template to obtain a target prompt word, wherein the standard prompt word template includes standard definitions of various error categories, a pre-instruction for guiding the model into a state of receiving input, and a post-instruction for clarifying the task goal of the model;
[0012] The target prompt word is input into the non-inference model, and the non-inference model is guided by the pre-instruction to obtain the test case, the error category and the modification suggestion, and the non-inference model is guided by the post-instruction to generate the evaluation basis of the test case based on the standard definition of the error category.
[0013] Optionally, determine a standard prompt word template, including:
[0014] Pre-building an initial prompt word template, wherein the initial prompt word template includes basic definitions of various error categories;
[0015] Filling the test case into the first placeholder in the initial prompt word template, and filling the error category and modification suggestion of the test case into the second placeholder in the initial prompt word template to obtain an initial target prompt word;
[0016] Inputting the initial target prompt word into a non-inference model, guiding the non-inference model to generate an evaluation basis corresponding to the test case based on a basic definition of an error category;
[0017] Conducting a quality assessment on the evaluation basis obtained to determine whether the corresponding quality requirements are met;
[0018] If the corresponding quality requirements are not met, the basic definitions of various error categories in the initial prompt word template are adjusted to obtain a new initial prompt word template for a new round of evaluation basis quality assessment;
[0019] In the case where the corresponding quality requirements are met, the initial prompt word template is determined to be a standard prompt word template.
[0020] Optionally, the marking of error categories and modification suggestions for test cases includes at least one of three marking angles, the three marking angles including: the marking angle of test case preconditions, the marking angle of test case test steps, and the marking angle of test case expected results; each marking angle has corresponding evaluation basis.
[0021] Optionally, supervised training of the inference model is performed using the completed labeled test cases to train a qualified test case evaluation model, including:
[0022] Based on the output results of the inference model training, the error categories, modification suggestions, and evaluation basis of the annotated test cases are aligned to match the format of the output results;
[0023] The reasoning model is supervised and trained using aligned test cases to obtain a qualified test case evaluation model.
[0024] Optionally, pre-training the initial reasoning model using knowledge data of the application domain to obtain a reasoning model that can understand and process natural language texts in the application domain includes:
[0025] Acquire a large amount of unlabeled data, wherein the unlabeled data includes various corpora in the application field;
[0026] Preprocessing the unlabeled data to obtain a data format that can be processed by the initial inference model, wherein the preprocessing includes at least data deduplication, data normalization and standardization, data smoothing and transformation, missing data correction, image recognition, and OCR processing;
[0027] The preprocessed unlabeled data is input into the initial reasoning model for pre-training to obtain a reasoning model that can understand and process natural language text in the application field.
[0028] Optionally, supervised training of the inference model is performed using the completed labeled test cases to train a qualified test case evaluation model, including:
[0029] Pre-setting fine-tuning parameters for model training, wherein the fine-tuning parameters include at least learning rate, batch size, and optimizer type;
[0030] Based on the preset fine-tuning parameters, the reasoning model is supervised and trained by completing the annotated test cases to obtain a qualified test case evaluation model.
[0031] A second aspect of the present application provides a method for evaluating a test case evaluation model, the method comprising:
[0032] Inputting the test case to be evaluated into a test case evaluation model for evaluation processing to obtain a corresponding evaluation result, wherein the evaluation result includes at least an error category and modification suggestions for the test case to be evaluated, and the test case evaluation model is the test case evaluation model in the test case evaluation model training method described in the first aspect of the present application;
[0033] Based on the evaluation results, the test case is optimized.
[0034] A third aspect of the present application provides a training system for a test case evaluation model, the system comprising:
[0035] A reasoning model determination module is used to determine an initial reasoning model applicable to the application field according to the application field to which the test case belongs;
[0036] A pre-training module, configured to pre-train the initial reasoning model using knowledge data from the application domain to obtain a reasoning model capable of understanding and processing natural language texts from the application domain;
[0037] an error and suggestion determination module, configured to determine error categories and modification suggestions for test cases used to train the inference model;
[0038] an evaluation basis determination module, configured to input the test case and the determined error category and modification suggestion into a non-inference model for processing to determine an evaluation basis for the test case, and annotate the test case with the evaluation basis, the error category, and the modification suggestion, wherein the evaluation basis is the correct thinking process of the inference model on the test case;
[0039] The supervised training module is used to perform supervised training on the inference model by completing the labeled test cases to train and obtain a qualified test case evaluation model.
[0040] The fourth aspect of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored on the memory and running on the processor, wherein when the computer program is executed by the processor, the steps in the training method of a test case evaluation model as described in the first aspect of the present application are implemented.
[0041] The fifth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the training method of a test case evaluation model as described in the first aspect of the present application.
[0042] The training method of a test case evaluation model provided in this application has the following advantages:
[0043] The present application provides a training method for a test case evaluation model. First, according to the application field to which the test case belongs, an initial reasoning model applicable to the application field is determined; the initial reasoning model is pre-trained by the knowledge data of the application field to obtain a reasoning model that can understand and process the natural language text of the application field; the error category and modification suggestion of the test case for training the reasoning model are determined; the test case and the determined error category and modification suggestion are input into the non-reasoning model for processing to determine the evaluation basis of the test case, and the test case is labeled with the evaluation basis, error category and modification suggestion. The evaluation basis is the correct thinking process of the reasoning model for the test case; the reasoning model is supervised by the completed labeled test case to train a qualified test case evaluation model. The present application finds that for the non-reasoning model, it directly executes user instructions and cannot observe its internal understanding process of the labeled data. This non-reasoning model directly gives error categories and modification suggestions when processing test cases, but cannot provide tuning direction when it cannot correctly identify or classify problems in the test case. The reasoning model is different. It can generate a thinking process before executing user instructions. This evaluation basis (that is, the thinking process) can reveal how the reasoning model understands the labeled data. By analyzing the thinking process, the deficiencies of the labeled data can be diagnosed. For example, the reasoning model may misunderstand the error category due to the deviation of the labeled data. This information is reflected in the thinking process, guiding the user to further correct the labeled data and optimize the performance of the reasoning model. Therefore, in view of the advantage of the reasoning model being easy to optimize, this application selects the reasoning model as the base model for the final test case evaluation. Unlike the training of non-reasoning models, which only requires engineers to mark the error categories and provide corresponding modification suggestions for the test cases used to supervise the training of non-reasoning models, the training of reasoning models also requires the evaluation basis to be marked for the test cases to supervise the thinking process of the reasoning model. It can be understood that the evaluation basis for the test case annotation belongs to the correct thinking process of the reasoning model (such as the correct logical chain, factual evidence, rules or reasoning steps based on which the reasoning model processes the test case). The evaluation basis is used to supervise and evaluate whether the thinking process of the test case during the training of the reasoning model conforms to the correct thinking process corresponding to the evaluation basis. This evaluation basis usually involves a lot of thinking content, and manual labeling is very time-consuming. Therefore, the present application provides a method for determining the evaluation basis of a test case. This method uses a large non-inferential model to determine the evaluation basis of the test case and use it to annotate the test case, thereby achieving the purpose of saving manpower. At the same time, after the inference model is trained and qualified to obtain an applicable test case evaluation model, the test case evaluation model can be directly used to evaluate test cases in a specific application field without the need for manual review, which can effectively improve the evaluation efficiency of test cases. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 A flowchart of a method for training a test case evaluation model according to an embodiment of the present application is shown;
[0046] Figure 2 A schematic diagram of generating evaluation criteria in a training method for a test case evaluation model according to one embodiment of the present application;
[0047] Figure 3 A schematic diagram of a training system for a test case evaluation model is shown as an embodiment of the present application. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0049] refer to Figure 1 , Figure 1 The flowchart of a method for training a test case evaluation model is shown as an embodiment of the present application. Figure 1 As shown, the method includes:
[0050] Step S1: Determine an initial reasoning model applicable to the application field to which the test case belongs.
[0051] In this embodiment, first, based on the application field to which the test case to be evaluated belongs, an AI open source large model (i.e., an initial inference model) suitable for the application field is selected for subsequent model pre-training and training to obtain a qualified test case evaluation model that can be used to evaluate the test case in the application field. Among them, the application field to which the test case to be evaluated belongs can be any application field, such as the automotive field, the medical equipment field, the financial field, etc. For the convenience of subsequent description, the application field to which the test case to be evaluated belongs is the automotive field.
[0052] In this embodiment, when the test case belongs to the automotive application domain, an open-source AI model suitable for the automotive field is selected based on the following dimensions: model performance, domain adaptability, model size and complexity, license and compliance, cross-platform compatibility, security and robustness. Furthermore, to provide a more consistent evaluation basis for error categories and modification suggestions, this application selects an AI open-source reasoning model that is ultimately suitable for the automotive field from the open-source reasoning AI model, i.e., the initial reasoning model.
[0053] Step S2: pre-training the initial reasoning model using the knowledge data of the application domain to obtain a reasoning model that can understand and process natural language texts in the application domain.
[0054] In this embodiment, after the initial reasoning model applicable to the automotive field is selected and obtained through step S1, the selected initial reasoning model is pre-trained with knowledge data in the automotive field. For example, automobile-related professional knowledge background, including specific technologies, regulations, etc., is used to pre-train the initial reasoning model, so that after completing the pre-training of the initial reasoning model, a reasoning model can be obtained that has automotive field knowledge and can understand and process natural language texts in the automotive field.
[0055] In the present application, step S2 may include: obtaining a large amount of unlabeled data, the unlabeled data including various types of corpus in the application field; preprocessing the unlabeled data to obtain a data format that can be processed by the initial reasoning model, the preprocessing at least including data deduplication, data normalization and standardization, data smoothing and transformation, missing gap filling, image recognition, and OCR processing; inputting the preprocessed unlabeled data into the initial reasoning model for pre-training to obtain a reasoning model that can understand and process the natural language text in the application field.
[0056] In this example, a large amount of unlabeled data from the automotive field is first collected. This data includes at least various corpora such as automotive design specifications, basic automotive knowledge, testing methodologies, systems, and regulations. The collected unlabeled data is preprocessed through a combination of manual processing and tool software (such as Pandas and Apache Spark), applying technologies such as data deduplication, data normalization and standardization, data smoothing and transformation, missing information correction, image recognition, and optical character recognition (OCR) to remove noise and irrelevant data. The cleaned and processed unlabeled data is then converted into a data format that can be processed by the inference model.
[0057] In this embodiment, after completing the preprocessing of the unlabeled data, the unlabeled data (e.g., input in the form of tokens, etc.) converted into a data format that can be processed by the inference model is input into the selected initial inference model for pretraining. After completing the pretraining of the initial inference model, a reasoning model with automotive domain knowledge that can understand and process natural language text in the automotive domain can be obtained. During the pretraining process, the reasoning model will learn the inherent laws and representation methods of the unlabeled data, and at the same time, it will learn a lot of useful knowledge, such as the grammar, semantics, and contextual relationships of the automotive domain language. This knowledge can help the reasoning model better understand and process natural language text in the automotive domain, thereby improving the natural language processing performance in the automotive domain. In addition, pretraining can also improve the generalization ability of the reasoning model, so that the reasoning model can also perform better when processing new, unseen data. After the pretraining is completed, the pretrained reasoning model is saved. This reasoning model will serve as the basic model for the subsequent fine-tuning supervised training phase to solve specific test case evaluation tasks.
[0058] Step S3: Determine error categories and modification suggestions for test cases used to train the inference model.
[0059] In the present embodiment, the inference model obtained by the pre-training of step S2 will be further supervised training in the future, and supervised training requires a large amount of sample data. Since the application field to which the test case belongs is the automotive field, these large amounts of sample data are also a large number of test cases in the automotive field. For each of the large number of test cases, the error category involved and the modification suggestions for each error category must be determined, that is, a test case may involve multiple error categories, and each error category has a corresponding modification suggestion. The error category and modification suggestions of the test case will be used to mark the test case, and will also be used to determine the evaluation basis of the test case in the future. Since the time cost of determining the error category and modification suggestions of the test case is relatively low, it is preferably manually determined directly by professional engineering personnel. It should be understood that other automatic determination methods (such as building a model to determine the error category and modification suggestions of the test case) can also be adopted.
[0060] Step S4: Input the test case and the determined error category and modification suggestion into the non-inference model for processing to determine the evaluation basis of the test case, and mark the test case with the evaluation basis, the error category and modification suggestion. The evaluation basis is the correct thinking process of the inference model on the test case.
[0061] In this embodiment, after determining the error category and modification suggestion corresponding to each test case through step S3, the test case and the error category and modification suggestion of the test case are input into a pre-trained non-inference model for processing, so that the non-inference model determines the evaluation basis of the test case based on the error category and modification suggestion of the test case. After obtaining the evaluation basis of the test case, the test case is labeled with the evaluation basis and the error category and modification suggestion of the test case. The labeled test case will be used to supervise the training of the reasoning model determined in step S2 to obtain a qualified test case evaluation model that can be used to evaluate test cases in the automotive field. Through the same implementation method, it is possible to label a large number of test cases with error categories, modification suggestions and evaluation basis. Among them, the evaluation basis of the test case determined by the non-inference model in this application is actually a label of the test case. The evaluation basis records the correct thinking process of the reasoning model on the test case, which is used to supervise the thinking process of the reasoning model in processing the test case when the reasoning model is trained later.
[0062] Step S5: supervise the reasoning model through the labeled test cases to obtain a qualified test case evaluation model.
[0063] In this embodiment, after obtaining a large number of labeled test cases through step S4, the obtained large number of labeled test cases are used to supervise the inference model determined through step S2. When the key indicators of the inference model (such as accuracy, precision, and recall) after training meet their corresponding set conditions (such as when the accuracy, precision, and recall all reach their corresponding certain values), a qualified test case evaluation model that can be used to evaluate test cases in the automotive field is determined to be obtained. It should be understood that the test cases used in the process of supervised training of the inference model also include some correct test cases without errors. For such correct test cases, there is no label with error category and modification suggestion, but a label indicating that they are correct test cases.
[0064] In this embodiment, the sample data used in this application for supervised training of the inference model is a test case in the automotive field. At the same time, the test case includes labels of error categories involved in the test case and modification suggestions for the error categories involved, which are manually marked by professional engineers. It also includes an evaluation basis label for the test case determined by a non-inference model based on the test case and the error category label and modification suggestion label of the test case. During supervised training of the inference model, the test case and the error category label, modification suggestion label and evaluation basis label of the test case are input into the inference model for supervised training. At the same time, during the training process, the output results of the inference model include the predicted error category, modification suggestion and evaluation basis of the test case, and the error category label, modification suggestion label and evaluation basis of the test case are used to compare with the error category, modification suggestion and evaluation basis of the test case predicted by the inference model to determine the training effect of the inference model. When the key indicators of the inference model do not meet the corresponding set conditions after training, the inference model is continued to be trained until a qualified test case evaluation model is obtained. The test case evaluation model will be used to evaluate the writing quality of various test cases in the automotive field. By inputting the test cases in the automotive field into the test case evaluation model for quality evaluation, the test case evaluation model will output the error category involved in the test case and give modification suggestions for the error category involved in the test case.
[0065] The present application provides a training method for a test case evaluation model. First, according to the application field to which the test case belongs, an initial reasoning model applicable to the application field is determined; the initial reasoning model is pre-trained by the knowledge data of the application field to obtain a reasoning model that can understand and process the natural language text of the application field; the error category and modification suggestion of the test case for training the reasoning model are determined; the test case and the determined error category and modification suggestion are input into the non-reasoning model for processing to determine the evaluation basis of the test case, and the test case is labeled with the evaluation basis, error category and modification suggestion. The evaluation basis is the correct thinking process of the reasoning model for the test case; the reasoning model is supervised by the completed labeled test case to train a qualified test case evaluation model. The present application finds that for the non-reasoning model, it directly executes user instructions and cannot observe its internal understanding process of the labeled data. This non-reasoning model directly gives error categories and modification suggestions when processing test cases, but cannot provide tuning direction when it cannot correctly identify or classify problems in the test case. The reasoning model is different. It can generate a thinking process before executing user instructions. This evaluation basis (that is, the thinking process) can reveal how the reasoning model understands the labeled data. By analyzing the thinking process, the deficiencies of the labeled data can be diagnosed. For example, the reasoning model may misunderstand the error category due to the deviation of the labeled data. This information is reflected in the thinking process, guiding the user to further correct the labeled data and optimize the performance of the reasoning model. Therefore, in view of the advantage of the reasoning model being easy to optimize, this application selects the reasoning model as the base model for the final test case evaluation. Unlike the training of non-reasoning models, which only requires engineers to mark the error categories and provide corresponding modification suggestions for the test cases used to supervise the training of non-reasoning models, the training of reasoning models also requires the evaluation basis to be marked for the test cases to supervise the thinking process of the reasoning model. It can be understood that the evaluation basis for the test case annotation belongs to the correct thinking process of the reasoning model (such as the correct logical chain, factual evidence, rules or reasoning steps based on which the reasoning model processes the test case). The evaluation basis is used to supervise and evaluate whether the thinking process of the test case during the training of the reasoning model conforms to the correct thinking process corresponding to the evaluation basis. This evaluation basis usually involves a lot of thinking content, and manual labeling is very time-consuming. Therefore, the present application provides a method for determining the evaluation basis of a test case. This method uses a large non-inferential model to determine the evaluation basis of the test case and use it to annotate the test case, thereby achieving the purpose of saving manpower. At the same time, after the inference model is trained and qualified to obtain an applicable test case evaluation model, the test case evaluation model can be directly used to evaluate test cases in a specific application field without the need for manual review, which can effectively improve the evaluation efficiency of test cases.
[0066] In combination with the above embodiments, in one embodiment, the present application also provides a method for training a test case evaluation model. In the test case evaluation model training method, step S4 may include steps S41 to S42:
[0067] Step S41: Fill the test case into the first placeholder in a pre-built standard prompt word template, and fill the error category and modification suggestion into the second placeholder in the standard prompt word template to obtain the target prompt word, wherein the standard prompt word template includes standard definitions of various error categories, and includes pre-instructions to guide the model into the input receiving state, and includes post-instructions to clarify the task objectives of the model.
[0068] In this embodiment, regarding the method for determining the evaluation basis of a test case provided in the previous embodiment of this application, this application further discovered that directly given a test case and the error category and modification suggestions for the test case, and allowing the non-inference large model to directly generate the evaluation basis for the test case, the quality of the evaluation basis obtained in this way is not very good, and there is a large deviation between the thinking process (that is, the generated evaluation basis) and the test expert's thinking. Therefore, for the method for determining the evaluation basis of a test case in this application, another embodiment is provided, which is to introduce a prompt word template when determining the evaluation basis of the test case through the non-inference large model, and the prompt word template includes at least the standard definitions of all error categories involved in test cases in the automotive field. When using the non-inference large model to determine the evaluation basis of the test case, the non-inference model is guided to closely focus on the standard definitions of the error categories in the prompt word template to generate the evaluation basis. Through this method, this application found that the quality of the evaluation basis obtained is good, and the thinking process (that is, the generated evaluation basis) can well match the thinking of the test expert.
[0069] Specifically, if Figure 2 As shown, the test case is filled into the first placeholder in the pre-built standard prompt word template. The first placeholder is the position reserved in the standard prompt word template for receiving the test case data, such as Figure 2 The content after the "###Test Case" part and before the "###Groundtruth Evaluation" is the test case that fills the first placeholder. At the same time, the error category and modification suggestions of the test case are filled into the second placeholder in the pre-built standard prompt word template. The second placeholder is the position reserved for receiving the error category and modification suggestions of the test case in the standard prompt word template, such as Figure 2The last two paragraphs of the "###Groundtruth Evaluation" section in the second placeholder are the error category and modification suggestions of the test case. After filling the test case and its error category and modification suggestions into the standard prompt word template, the corresponding target prompt word is obtained, such as Figure 2 All the contents in the rightmost rectangle are the target prompt words obtained based on a test case. It should be understood that Figure 2 It shows that the error category based on which the generated target prompt word is based belongs to the precondition error category of the test case. Among them, the standard prompt word template pre-built in this application also records all error categories involved in the test cases in the automotive field and the standard definitions of all error categories, such as Figure 2 The content of the part "Error categories that may occur in known preconditions are" is all the error categories involved in the test case's preconditions and the standard definitions of all error categories. At the same time, the standard prompt word template also includes a pre-instruction to guide the non-inference model to enter the input receiving state, and the pre-instruction is used to guide the non-inference model to receive the test case filled with the first placeholder and the second placeholder, and the error category and modification suggestions of the test case. At the same time, the standard prompt word template also includes a post-instruction to clarify the task objectives of the model, and the post-instruction is used to guide the non-inference model to generate the corresponding evaluation basis for the test case based on the received test case, the error category and modification suggestions of the test case, and the standard definitions of all error categories.
[0070] Step S42: Input the target prompt word into the non-inference model, guide the non-inference model to obtain the test case, the error category and the modification suggestion through the pre-instruction, and guide the non-inference model to generate the evaluation basis of the test case based on the standard definition of the error category through the post-instruction.
[0071] In this embodiment, the target prompt word obtained in step S41 is input into a pre-trained non-inference model for processing. This model is guided to closely follow the standard definition of all error categories in the target prompt word to generate corresponding evaluation criteria for the test case containing the target prompt word's own error category and modification suggestions. After obtaining the evaluation criteria for the test case, the test case is annotated with the evaluation criteria, the error category, and modification suggestions for the test case. The annotated test case is then used to supervise the inference model determined in step S2, thereby obtaining a qualified test case evaluation model that can be used to evaluate test cases in the automotive field.
[0072] In conjunction with the above embodiments, in one embodiment, the present application also provides a method for training a test case evaluation model. In the test case evaluation model training method, determining a standard prompt word template includes:
[0073] Step S01: pre-constructing an initial prompt word template, wherein the initial prompt word template includes basic definitions of various error categories.
[0074] Step S02: Fill the test case into the first placeholder in the initial prompt word template, and fill the error category and modification suggestion of the test case into the second placeholder in the initial prompt word template to obtain the initial target prompt word.
[0075] Step S03: inputting the initial target prompt word into the non-inference model, guiding the non-inference model to generate evaluation criteria corresponding to the test case based on the basic definition of the error category.
[0076] Step S04: Performing a quality assessment on the obtained evaluation basis to determine whether it meets the corresponding quality requirements.
[0077] Step S05: If the corresponding quality requirements are not met, the basic definitions of various error categories in the initial prompt word template are adjusted to obtain a new initial prompt word template for a new round of evaluation basis quality assessment;
[0078] Step S06: When the corresponding quality requirements are met, the initial prompt word template is determined to be a standard prompt word template.
[0079] In this embodiment, in order to ensure that the non-inference model generates better quality evaluation basis based on the prompt word template, the present application pre-constructs an initial prompt word template. The structure of the initial prompt word template is consistent with that of the final standard prompt word template. The difference is that the definition of all error categories involved in the test case in the automotive field is a basic definition determined by the test expert. Whether the definition of the error category is accurate directly affects the quality of the final generated evaluation basis, and thus affects the training effect of the model. Therefore, before the actual evaluation basis of the test case is generated, the present application first determines the quality of the evaluation basis generated based on the given prompt word template. If the quality is not good, the definition of each error category is further adjusted, and then the quality of the evaluation basis generated based on the given prompt word template is re-determined until a good quality evaluation basis is obtained. At this time, the current prompt word template is determined as the standard prompt word template, and the definition of each error category in the standard prompt word template is the corresponding standard definition.
[0080] Specifically, an initial prompt word template is pre-constructed. This initial prompt word template has the same structure as the final standard prompt word template, except that the definitions of all error categories involved in automotive test cases are based on basic definitions determined by test experts. The test case is entered into the first placeholder of the initial prompt word template, while the error category and modification suggestions for the test case are entered into the second placeholder of the pre-constructed standard prompt word template. After the test case, its error category, and modification suggestions are entered into the standard prompt word template, a corresponding initial target prompt word is obtained. The obtained initial target prompt word is input into a pre-trained non-inference model for processing, guiding the non-inference model to closely follow the basic definitions of all error categories in the initial target prompt word to generate corresponding evaluation criteria for the test case. After obtaining the evaluation criteria for the test case, the test experts evaluate whether the quality of the generated evaluation criteria meets the corresponding quality requirements. An optional implementation method is to first have the test experts compile the standard evaluation criteria corresponding to the test case. After obtaining the evaluation criteria generated by the non-inference model, the test experts compare the evaluation criteria with the standard evaluation criteria. If the two are highly similar, the evaluation criteria generated by the non-inference model is determined to meet the corresponding quality requirements. If it is determined that the corresponding quality requirements are met, the current initial prompt word template is determined to be the standard prompt word template, and the definitions of each error category recorded therein are also standard definitions for each error category. If it is determined that the corresponding quality requirements are not met, the test experts will adjust the basic definitions of each error category in the initial prompt word template and conduct a new round of evaluation based on the adjusted initial prompt word template. In other words, steps S02 to S06 are repeated using the adjusted initial prompt word template.
[0081] In conjunction with the above embodiments, in one implementation, the present application also provides a method for training a test case evaluation model. In this method, the error categories and modification suggestions for test cases are annotated using at least one of three annotation angles: the angle for annotating test case preconditions, the angle for annotating test case test steps, and the angle for annotating test case expected results; each annotation angle has a corresponding evaluation basis.
[0082] In this embodiment, for the sake of structural considerations of test cases in the automotive field, in order to enable the qualified test case evaluation model obtained by supervised training to evaluate the test cases in the automotive field in terms of error categories, modification suggestions and evaluation basis, this application includes three annotation angles for the annotation of error categories and modification suggestions for test cases. It should be understood that only at least one of them can be annotated. The three annotation angles are the annotation angle of the test case preconditions, the annotation angle of the test case test steps, and the annotation angle of the test case expected results. At the same time, there is a corresponding evaluation basis for each annotation angle, that is, the test case is annotated with several angles (such as the two angles of preconditions and test steps), and correspondingly, several evaluation bases will be generated (correspondingly including the evaluation basis of the precondition angle and the evaluation basis of the test steps) to annotate the test case. Among them, the expected result of the test case is the correct output or behavior that the test case should produce under specific input or operation given by the test case developer.
[0083] In combination with the above embodiments, in one embodiment, the embodiments of the present application also provide a method for training a test case evaluation model. In the test case evaluation model training method, the reasoning model is supervised and trained by the marked test cases to obtain a qualified test case evaluation model, including: based on the output results of the reasoning model training, the error categories, modification suggestions and evaluation basis of the marked test cases are aligned to match the format of the output results; the reasoning model is supervised and trained by the aligned test cases to obtain a qualified test case evaluation model.
[0084] In this embodiment, in order to facilitate the training of the supervised reasoning model by annotating the error categories, modification suggestions and evaluation basis of the test cases. Based on the output results of the reasoning model training, this application aligns the error categories, modification suggestions and evaluation basis of the test cases that have been annotated for supervising the reasoning model training. That is, the order of the error categories, modification suggestions and evaluation basis involved in the output results output during the reasoning model training is modification suggestion, error category, evaluation basis, and the corresponding order of the error categories, modification suggestions and evaluation basis annotated for the test cases is also set to modification suggestion, error category, evaluation basis. The reasoning model is then supervised and trained with the aligned test cases to obtain a qualified test case evaluation model.
[0085] In combination with the above embodiments, in one embodiment, the embodiments of the present application further provide a method for training a test case evaluation model. In the test case evaluation model training method, the reasoning model is supervised trained by completing the labeled test cases to train and obtain a qualified test case evaluation model, including: pre-setting fine-tuning parameters for model training, the fine-tuning parameters including at least learning rate, batch size, and optimizer type; based on the pre-set fine-tuning parameters, the reasoning model is supervised trained by completing the labeled test cases to train and obtain a qualified test case evaluation model.
[0086] In this embodiment, before supervised training of the inference model determined in step S2, fine-tuning parameters for model training are pre-set based on the requirements of the current task. The fine-tuning parameters need to be adjusted according to the specific situation to achieve the best training effect. The fine-tuning parameters include at least the learning rate, batch size, and optimizer type. Based on these pre-set fine-tuning parameters, supervised training of the inference model is then performed using labeled test cases to obtain a qualified test case evaluation model. When the test case labels are aligned, supervised training of the inference model is performed using labeled and aligned test cases to obtain a qualified test case evaluation model. Specifically, the parameters of the pre-trained inference model are loaded into the current model. These parameters include the weights and biases of the inference model. Loading the parameters of the pre-trained inference model can help the model quickly adapt to the current task and improve training efficiency. Using the prepared dataset consisting of labeled and aligned test cases and the set fine-tuning parameters, the model is fine-tuned. Fine-tuning training uses a mini-batch gradient descent algorithm to iteratively update the model parameters. At each iteration, the model's loss function is calculated, and the model's parameters are updated using the backpropagation algorithm. After fine-tuning training is complete, the model is evaluated to determine its performance on the current task. Evaluation metrics can include precision, recall, and F1 score. If the model's performance meets the requirements, a corresponding test case evaluation model is generated, which can be used to evaluate the test case being evaluated. If it does not meet the requirements, further fine-tuning training is required until it meets the requirements.
[0087] Based on the same inventive concept, the present application provides an evaluation method for a test case evaluation model, the method comprising: inputting the test case to be evaluated into the test case evaluation model for evaluation processing to obtain a corresponding evaluation result, the evaluation result at least including the error category and modification suggestions of the test case to be evaluated, the test case evaluation model being the test case evaluation model in the training method of a test case evaluation model described in the first aspect of the present application; and optimizing the test case based on the evaluation result.
[0088] In this embodiment, a test case to be evaluated is input into a test case evaluation model that has been trained using the test case evaluation model training method provided by the first aspect of this application for evaluation. The test case evaluation model outputs a corresponding evaluation result, which includes at least the error category involved in the test case to be evaluated and the corresponding modification suggestions. The test case evaluation model also outputs the evaluation basis for the test case to be evaluated. Based on the obtained evaluation result of the test case to be evaluated, the test case to be evaluated is optimized.
[0089] In this embodiment, the present application improves the professionalism and consistency of the evaluation criteria by introducing a reasoning model and proposing a generative paradigm for the thought process (i.e., the aforementioned standard prompt word template). Ultimately, the evaluation criteria determination method provided by this application achieved an accuracy rate of over 90% across all evaluation dimensions, significantly reducing annotation costs while significantly improving generation efficiency. First, in actual applications, it was found that after counting common error categories related to test efficiency, although traditional non-reasoning small models could be trained by annotating the error categories and modification suggestions of test cases, this method could not consistently output error categories, modification suggestions, and evaluation criteria, making it difficult to meet production needs. To address this problem, we introduced a reasoning model. By enhancing the model's reasoning capabilities, it was able to provide more consistent evaluation criteria, thereby providing more reasonable error categories and modification suggestions. However, despite the improvement in consistency of the reasoning model, its judgment of specific professional errors still did not fully conform to professional intuition. To this end, the reasoning model needs to be further trained to make it more in line with professional needs. Unlike simple error categories and modification suggestion annotations, training a reasoning model requires further annotation of the evaluation criteria, that is, detailed annotation of the thought process that determines the error category. However, the workload of labeling this thinking process is much higher than that of labeling ordinary evaluation basis. The existing automatic generation and labeling methods mainly rely on multiple sampling of the reasoning model and manual verification to screen out the error categories and modification suggestions. The correct thinking process is used as the evaluation basis. However, this method is relatively random and time-consuming, making it difficult to meet the needs efficiently. To solve this problem, the present application further proposes a thinking process generation paradigm (i.e., a standard prompt word template) to guide a larger non-inference model to assist in generating the thinking process (i.e., the evaluation basis). This method not only greatly saves labeling time, but also makes the generated evaluation basis more in line with the thinking mode of the test experts, thereby significantly improving the professionalism and practicality of the model. Based on these high-quality evaluation bases, error categories and modification suggestions, the present application finally trained a large model for use case review and optimization that is more in line with production requirements, greatly improving the efficiency of the testing process.
[0090] Based on the same inventive concept, an embodiment of the present application provides a training system for a test case evaluation model, such as Figure 3 As shown, the system 300 includes:
[0091] The reasoning model determination module 301 is used to determine an initial reasoning model applicable to the application field according to the application field to which the test case belongs;
[0092] A pre-training module 302 is configured to pre-train the initial reasoning model using knowledge data from the application domain to obtain a reasoning model that can understand and process natural language texts from the application domain;
[0093] an error and suggestion determination module 303 for determining error categories and modification suggestions for test cases used to train the inference model;
[0094] An evaluation basis determination module 304 is configured to input the test case, the determined error category, and the modification suggestion into a non-inference model for processing to determine an evaluation basis for the test case, and annotate the test case with the evaluation basis, the error category, and the modification suggestion. The evaluation basis is the correct thinking process of the inference model on the test case.
[0095] The supervised training module 305 is used to perform supervised training on the reasoning model through the completed labeled test cases, so as to train and obtain a qualified test case evaluation model.
[0096] Optionally, the evaluation basis determination module 304 includes:
[0097] a target prompt word determination module, configured to fill the test case into a first placeholder in a pre-built standard prompt word template, and fill the error category and modification suggestion into a second placeholder in the standard prompt word template to obtain a target prompt word, wherein the standard prompt word template includes standard definitions of various error categories, a pre-instruction for guiding the model into an input receiving state, and a post-instruction for clarifying the task goal of the model;
[0098] An evaluation basis determination submodule is used to input the target prompt word into the non-inference model, guide the non-inference model to obtain the test case, the error category and the modification suggestion through the pre-instruction, and guide the non-inference model to generate the evaluation basis of the test case based on the standard definition of the error category through the post-instruction.
[0099] Optionally, the system 300 further includes a standard prompt word template determination module, configured to determine a standard prompt word template; the standard prompt word template determination module includes:
[0100] An initial prompt word template construction module, used to pre-construct an initial prompt word template, wherein the initial prompt word template includes basic definitions of various error categories;
[0101] an initial target prompt word determination module, configured to fill a test case into a first placeholder in the initial prompt word template, and fill an error category and modification suggestion of the test case into a second placeholder in the initial prompt word template to obtain an initial target prompt word;
[0102] An evaluation basis determination module, configured to input the initial target prompt word into a non-inference model, and guide the non-inference model to generate an evaluation basis corresponding to the test case based on a basic definition of an error category;
[0103] A quality assessment module, configured to perform a quality assessment on the obtained evaluation basis to determine whether the corresponding quality requirements are met;
[0104] A first evaluation result determination module is configured to adjust the basic definitions of various error categories in the initial prompt word template if the corresponding quality requirements are not met, and obtain a new initial prompt word template for a new round of evaluation basis quality assessment;
[0105] The second evaluation result determination module is configured to determine that the initial prompt word template is a standard prompt word template if the corresponding quality requirement is met.
[0106] Optionally, the marking of error categories and modification suggestions for test cases in the system 300 includes at least one of three marking angles, the three marking angles including: the marking angle of test case preconditions, the marking angle of test case test steps, and the marking angle of test case expected results; each marking angle has corresponding evaluation basis.
[0107] Optionally, the supervised training module 305 includes:
[0108] An alignment processing module is used to align the error categories, modification suggestions and evaluation basis of the annotated test cases based on the output results of the inference model training to match the format of the output results;
[0109] The supervised training submodule is used to perform supervised training on the inference model through the aligned test cases to obtain a qualified test case evaluation model.
[0110] Optionally, the pre-training module 302 includes:
[0111] A data acquisition module, configured to acquire a large amount of unlabeled data, wherein the unlabeled data includes various corpora in the application field;
[0112] A preprocessing module is used to preprocess the unlabeled data to obtain a data format that can be processed by the initial inference model. The preprocessing includes at least data deduplication, data normalization and standardization, data smoothing and transformation, missing error correction, image recognition, and OCR processing;
[0113] The pre-training submodule is used to input the pre-processed unlabeled data into the initial inference model for pre-training to obtain an inference model that can understand and process the natural language text in the application field.
[0114] Optionally, the supervised training module 305 includes:
[0115] A fine-tuning parameter setting module is used to pre-set fine-tuning parameters for model training, wherein the fine-tuning parameters include at least learning rate, batch size, and optimizer type;
[0116] The supervised training submodule is used to perform supervised training on the inference model based on the pre-set fine-tuning parameters by completing the labeled test cases to train and obtain a qualified test case evaluation model.
[0117] Based on the same inventive concept, an embodiment of the present application provides an evaluation system for a test case evaluation model, the system comprising:
[0118] An evaluation module, configured to input a test case to be evaluated into a test case evaluation model for evaluation processing, and obtain a corresponding evaluation result, wherein the evaluation result includes at least an error category and modification suggestions for the test case to be evaluated, wherein the test case evaluation model is the test case evaluation model in the test case evaluation model training method described in the first aspect of the present application;
[0119] An optimization module is used to optimize the test case based on the evaluation result.
[0120] Based on the same inventive concept, an embodiment of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored on the memory and running on the processor. When the computer program is executed by the processor, it implements the steps in the training method of a test case evaluation model as described in the first aspect of the present application.
[0121] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the training method of a test case evaluation model as described in the first aspect of the present application.
[0122] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0123] It should be noted that for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0124] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0125] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0126] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0127] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0129] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0130] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0131] The above is a detailed introduction to the training method, system, equipment and medium of a test case evaluation model provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A training method for a test case evaluation model, characterized in that: The method comprises: Determine an initial reasoning model applicable to the application domain to which the test case belongs; Pre-training the initial reasoning model using knowledge data from the application domain to obtain a reasoning model that can understand and process natural language text from the application domain; determining error categories and modification suggestions for test cases used to train the inference model; Inputting the test case and the determined error category and modification suggestion into a non-inference model for processing to determine an evaluation basis for the test case, and annotating the test case with the evaluation basis, the error category, and the modification suggestion, wherein the evaluation basis is the correct thinking process of the inference model on the test case; The reasoning model is supervised and trained by completing the labeled test cases to obtain a qualified test case evaluation model.
2. A test case evaluation model training method according to claim 1, characterized in that: Inputting the test case and the determined error category and modification suggestion into a non-inference model for processing to determine an evaluation basis for the test case, including: Filling the test case into a first placeholder in a pre-built standard prompt word template, and filling the error category and modification suggestion into a second placeholder in the standard prompt word template to obtain a target prompt word, wherein the standard prompt word template includes standard definitions of various error categories, a pre-instruction for guiding the model into a state of receiving input, and a post-instruction for clarifying the task goal of the model; The target prompt word is input into the non-inference model, and the non-inference model is guided by the pre-instruction to obtain the test case, the error category and the modification suggestion, and the non-inference model is guided by the post-instruction to generate the evaluation basis of the test case based on the standard definition of the error category.
3. The training method of a test case evaluation model according to claim 2, characterized in that: Determine the standard prompt word template, including: Pre-building an initial prompt word template, wherein the initial prompt word template includes basic definitions of various error categories; Filling the test case into the first placeholder in the initial prompt word template, and filling the error category and modification suggestion of the test case into the second placeholder in the initial prompt word template to obtain an initial target prompt word; Inputting the initial target prompt word into a non-inference model, guiding the non-inference model to generate an evaluation basis corresponding to the test case based on a basic definition of an error category; Conducting a quality assessment on the evaluation basis obtained to determine whether the corresponding quality requirements are met; If the corresponding quality requirements are not met, the basic definitions of various error categories in the initial prompt word template are adjusted to obtain a new initial prompt word template for a new round of evaluation basis quality assessment; In the case where the corresponding quality requirements are met, the initial prompt word template is determined to be a standard prompt word template.
4. The training method of a test case evaluation model according to claim 1, characterized in that: The marking of error categories and modification suggestions for test cases includes at least one of three marking angles, the three marking angles including: the marking angle of test case preconditions, the marking angle of test case test steps, and the marking angle of test case expected results; each marking angle has corresponding evaluation basis.
5. The training method of a test case evaluation model according to claim 1, characterized in that: The reasoning model is supervised and trained by completing the annotated test cases to obtain a qualified test case evaluation model, including: Based on the output results of the inference model training, the error categories, modification suggestions, and evaluation basis of the annotated test cases are aligned to match the format of the output results; The reasoning model is supervised and trained using aligned test cases to obtain a qualified test case evaluation model.
6. The training method of a test case evaluation model according to claim 1, characterized in that: Pre-training the initial reasoning model using the knowledge data of the application domain to obtain a reasoning model that can understand and process natural language text in the application domain includes: Acquire a large amount of unlabeled data, wherein the unlabeled data includes various corpora in the application field; Preprocessing the unlabeled data to obtain a data format that can be processed by the initial inference model, wherein the preprocessing includes at least data deduplication, data normalization and standardization, data smoothing and transformation, missing data correction, image recognition, and OCR processing; The preprocessed unlabeled data is input into the initial reasoning model for pre-training to obtain a reasoning model that can understand and process natural language text in the application field.
7. The training method of a test case evaluation model according to claim 1, characterized in that: The reasoning model is supervised and trained by completing the annotated test cases to obtain a qualified test case evaluation model, including: Pre-setting fine-tuning parameters for model training, wherein the fine-tuning parameters include at least learning rate, batch size, and optimizer type; Based on the preset fine-tuning parameters, the reasoning model is supervised and trained by completing the annotated test cases to obtain a qualified test case evaluation model.
8. A method for evaluating a test case evaluation model, characterized in that: The method comprises: Inputting the test case to be evaluated into a test case evaluation model for evaluation processing to obtain a corresponding evaluation result, wherein the evaluation result at least includes an error category and modification suggestions for the test case to be evaluated, wherein the test case evaluation model is the test case evaluation model in the test case evaluation model training method according to any one of claims 1 to 7; Based on the evaluation results, the test case is optimized.
9. A training system for a test case evaluation model, characterized in that: The system comprises: A reasoning model determination module is used to determine an initial reasoning model applicable to the application field according to the application field to which the test case belongs; A pre-training module, configured to pre-train the initial reasoning model using knowledge data from the application domain to obtain a reasoning model capable of understanding and processing natural language texts from the application domain; an error and suggestion determination module, configured to determine error categories and modification suggestions for test cases used to train the inference model; an evaluation basis determination module, configured to input the test case and the determined error category and modification suggestion into a non-inference model for processing to determine an evaluation basis for the test case, and annotate the test case with the evaluation basis, the error category, and the modification suggestion, wherein the evaluation basis is the correct thinking process of the inference model on the test case; The supervised training module is used to perform supervised training on the inference model by completing the labeled test cases to train and obtain a qualified test case evaluation model.
10. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and running on the processor, wherein when the computer program is executed by the processor, the steps in the training method of a test case evaluation model as described in any one of claims 1 to 7 are implemented.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the training method of a test case evaluation model as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Evaluation model training method, data processing method and device
CN121094052A
Evaluation model training method, data processing method and device
CN121094052B