Test case set generation method and device, equipment and storage medium

By generating test case sets, using proxy models and white box models to generate test cases, the problems of insufficient and incomplete testing of aviation deep learning models are solved, and the performance and testing effect of the model in complex environments is improved.

CN120596367APending Publication Date: 2025-09-05CHINA AVIATON SYST ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510587156.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the test verification, deep learning models in the aviation field face insufficient data volume and insufficient diversity, and difficult to cover complex environments. The traditional test case construction methods cannot be directly transplanted, resulting in insufficient generalization capabilities of the model and great impact on adversarial samples, and insufficient and incomplete testing.

Method used

By obtaining the test data set, the model to be tested, the pre-stored model and the user expectation semantics, generating the proxy model or directly using the white box model, combining the test data set and the pre-stored model to generate test cases, improve data diversity and targeting, and evaluate the robustness and reliability of the model.

Benefits of technology

Ensure the adequacy and comprehensiveness of the test, improve the performance of the model in complex environments, improve the diversity and pertinence of the test data, and enhance the robustness and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596367A_ABST
    Figure CN120596367A_ABST
Patent Text Reader

Abstract

The invention discloses a test case set generation method and device, equipment and a storage medium. The method comprises the steps of obtaining a to-be-tested data set, a to-be-tested model, test task text description, a pre-storage model and user expectation semantics; under the condition that the type of the to-be-tested model comprises a black box model, generating an agent model based on the test data set and a pre-storage module; under the condition that the type of the to-be-tested model comprises a white box model, determining the to-be-tested model as an agent model; and generating a test case according to the test data set, the test task text description, the pre-stored model, the proxy model and the user expected semantics. According to the scheme, the original test case is expanded by generating the test case, so that the diversity and pertinence of the test data can be improved, and the to-be-tested model can show excellent performance in various complex situations, thereby ensuring the sufficiency and comprehensiveness of the test.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of aviation intelligent testing technology, and in particular to a method, apparatus, device and storage medium for generating a test case set. Background Art

[0002] The aviation sector, characterized by its complex natural environment, high flight risks, and high costs, is facing an increasing demand for testing and validation. Furthermore, AI models in aviation require thorough validation before they can be put into actual flight. The performance evaluation and validation of deep learning models relies on high-quality, high-coverage test data. However, current test cases for aviation systems suffer from the following deficiencies: First, insufficient data volume and diversity make it difficult to cover the diverse scenarios and unexpected variations found in real-world environments, making it difficult for models to cope with extreme conditions in real-world environments, potentially leading to devastating consequences. Second, deep learning models lack generalization capabilities and are susceptible to adversarial examples. Even slight perturbations in the input data can easily lead to incorrect predictions. Furthermore, traditional software test case construction methods cannot be directly applied to deep neural network testing. Current deep learning test cases primarily utilize random assignment strategies to generate test cases, requiring manual labeling of test data. This approach is susceptible to subjective influences, lacks systematicity and comprehensiveness, and requires significant labor costs. Furthermore, the lack of direct correlation between test cases and the network itself makes it difficult to fully test the network and ensure sufficient and comprehensive testing. Summary of the Invention

[0003] The embodiments of the present application provide a method, apparatus, device, and storage medium for generating a test case set. By evaluating the model to be tested, the method then selects a corresponding method to generate test cases to expand the original test cases, and then evaluates the model to be tested based on the expanded test cases, ultimately generating a corresponding test report. This not only improves the diversity and pertinence of the test data, but also effectively evaluates the robustness and reliability of the neural network in the model, ensuring that the model to be tested can demonstrate excellent performance in a variety of complex situations, thereby ensuring the adequacy and comprehensiveness of the test.

[0004] In a first aspect, an embodiment of the present application further provides a method for generating a test case set, the method comprising:

[0005] Obtain the test dataset, the model to be tested, the test task text description, the pre-stored model, and the user's expected semantics;

[0006] When the type of model to be tested includes a black box model, generating a proxy model based on the test dataset and the pre-stored module;

[0007] In a case where the type of the model to be tested includes a white box model, determining the model to be tested as a proxy model;

[0008] Generate test cases based on test datasets, test task text descriptions, pre-stored models, proxy models, and user expected semantics.

[0009] Optionally, before generating a test case based on the test dataset, the test task text description, the pre-stored model, the proxy model, and the user's expected semantics, the method further includes:

[0010] Determine whether the data volume of the test data set meets the preset requirements;

[0011] If the preset requirements are not met, a request message is sent to the database so that the database can filter available test data according to the preset requirements to supplement the test data set.

[0012] Optionally, the above-mentioned generation of a proxy model based on a test dataset and a pre-stored model includes:

[0013] determining an alternative model among pre-stored models;

[0014] Test the alternative model on the test dataset and generate predicted labels;

[0015] Train the surrogate model based on the predicted labels and the test dataset;

[0016] The trained surrogate model is determined as the surrogate model.

[0017] Optionally, the above training of the alternative model based on the predicted labels and the test dataset includes:

[0018] Training the surrogate model based on the predicted label and the test dataset, and determining the output of the surrogate model as the predicted output;

[0019] Under the condition of consistency constraints on the predicted output and the predicted label, the parameters of the surrogate model are modified until the surrogate model converges.

[0020] Optionally, the above-mentioned test case generation based on the test dataset, the test task text description, the pre-stored model, the proxy model and the user expected semantics includes:

[0021] Build a use case expected output-model neuron coverage library based on the test label semantics of the test dataset and the agent model inference;

[0022] Build a semantic model based on the test task text description and the natural language processing model in the pre-stored model;

[0023] Generate test cases based on the expected output of the use case - model neuron coverage library, semantic model and user expected semantics.

[0024] Optionally, the above-mentioned test case generation based on the use case expected output-model neuron coverage library, semantic model and user expected semantics includes:

[0025] Generate a neuron coverage vector library of user expected categories based on the semantic model and user expected semantics;

[0026] The difference between the neuron coverage vector library of the user's expected category and the use case expected output-model neuron coverage library is retrieved to generate test cases in a gradient increasing manner.

[0027] Optionally, the above method further includes:

[0028] Obtain performance metrics and pre-stored test cases;

[0029] Determine the effectiveness of test cases based on performance indicators;

[0030] If the test case is valid, the model to be tested is evaluated based on the test case and pre-stored test cases, and a test report is generated.

[0031] In a second aspect, an embodiment of the present application further provides a device for generating a test case set, the device comprising:

[0032] The acquisition module is used to obtain the test data set, the model to be tested, the test task text description, the pre-stored model and the user's expected semantics;

[0033] A generation module, configured to generate a proxy model based on a test dataset and a pre-stored model when the type of the model to be tested includes a black box model;

[0034] The above-mentioned generation module is also used to determine the model to be tested as a proxy model when the type of the model to be tested includes a white box model; the above-mentioned generation module is also used to generate test cases based on the test data set, the test task text description, the pre-stored model, the proxy model and the user's expected semantics.

[0035] In a third aspect, an embodiment of the present application further provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a method for generating a test case set as provided in any embodiment of the present application is implemented.

[0036] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements a method for generating a test case set as provided in any embodiment of the present application.

[0037] The embodiment of the present application provides a method, apparatus, device and storage medium for generating a test case set, the method comprising: obtaining a data set to be tested, a model to be tested, a text description of a test task, a pre-stored model, and user expected semantics; when the type of the model to be tested includes a black box model, generating a proxy model based on the test data set and the pre-stored module; when the type of the model to be tested includes a white box model, determining the model to be tested as a proxy model; generating test cases based on the test data set, the text description of the test task, the pre-stored model, the proxy model and the user expected semantics. The above scheme can improve the diversity and pertinence of the test data by evaluating the model to be tested and then selecting the corresponding method to generate test cases to expand the original test cases, thereby ensuring that the model to be tested can show excellent performance in various complex situations, thereby ensuring the adequacy and comprehensiveness of the test. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of a method for generating a test case set provided by this application;

[0039] Figure 2 This is a flow chart of another method for generating a test case set provided by this application;

[0040] Figure 3 This is a flow chart of a method for generating a proxy model based on a test data set and a pre-stored model provided in an embodiment of the present application;

[0041] Figure 4 This is a flow chart of a method for generating test cases based on a test data set, a test task text description, a pre-stored model, a proxy model, and user expected semantics, provided by an embodiment of the present application;

[0042] Figure 4a This is a flow chart of a method for generating a test report provided in an embodiment of the present application;

[0043] Figure 5 This is a schematic diagram of the structure of a device for generating a test case set provided in an embodiment of the present application;

[0044] Figure 6 This is a schematic diagram of the structure of another device for generating a test case set provided in an embodiment of the present application;

[0045] Figure 7 It is a structural diagram of the computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] The present disclosure will be further described below with reference to the embodiments shown in the accompanying drawings. It should be understood that the specific embodiments described herein are intended only to illustrate the present disclosure and are not intended to limit the present disclosure. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions of the present disclosure, not all of the structures.

[0047] In addition, in the embodiments of this application, words such as "optionally" or "exemplarily" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "optionally" or "exemplarily" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "optionally" or "exemplarily" is intended to present the relevant concepts in a concrete manner.

[0048] The embodiment of the present application provides a method for generating a test case set, which can be applied to the field of aviation intelligent testing technology. The type of the model to be tested and the data set to be tested are combined to evaluate the model to be tested, and then the corresponding method is selected to generate test cases to expand the original test cases, and the model to be tested is evaluated based on the expanded test cases, and finally a corresponding test report is obtained. This can not only improve the diversity and pertinence of the test data, but also effectively evaluate the robustness and reliability of the neural network in the model, and ensure that the model to be tested can show excellent performance in various complex situations, thereby ensuring the adequacy and comprehensiveness of the test. The method can be executed by the test case set generation device provided in the embodiment of the present application, and the device can be implemented in software and / or hardware. In a specific embodiment, the device can be integrated in a computer device, and the computer device can be, for example, a server. The following embodiments will be described by taking the device integrated in a computer device as an example, such as Figure 1 As shown, the method may include but is not limited to the following steps:

[0049] S101. Obtain a test dataset, a model to be tested, a text description of a test task, a pre-stored model, and user expected semantics.

[0050] In an embodiment of the present application, the test data set is used to generate a proxy model and a test case as test data. The model to be tested is used to determine its type, as well as to determine the proxy model in the corresponding situation. Accordingly, the type of the model to be tested is used to determine whether the model to be tested is a black box or a white box, and then determine the implementation method of generating test cases in subsequent steps. Similarly, the text description of the test task can be understood as the user's intention and / or the sample label or text description of the user's input data. The pre-stored model is a pre-stored network model of various types, such as a replacement model required by the user, a natural language processing model, and the like. The user's expected semantics can be understood as the semantic information directly input by the user, which is used to generate a neuron coverage vector library of the user's expected category in the test case generation stage.

[0051] S102: When the type of the model to be tested includes a black box model, a proxy model is generated based on the test data set and the pre-stored model.

[0052] In an embodiment of the present application, if the model to be tested is a black box model, indicating that the structure of the model to be tested is not provided, a proxy model can be generated based on the test dataset and the pre-stored model to serve as a new model to replace the model to be tested. For example, a surrogate model can be determined from the pre-stored model and trained based on the test dataset.

[0053] S103: When the type of the model to be tested includes a white-box model, determine the model to be tested as a proxy model.

[0054] On the contrary, if the type of the model to be tested is a white box model, which means that the model structure is known, the model can be directly used as a proxy model to generate test cases.

[0055] S104: Generate test cases based on the test data set, the test task text description, the pre-stored model, the proxy model, and the user's expected semantics.

[0056] Exemplarily, the implementation of this step may include: constructing a use case expected output-model neuron coverage library based on the test label semantics of the test data set and the agent model inference; constructing a semantic model based on the test task text description and the natural language processing model in the pre-stored model; generating test cases based on the use case expected output-model neuron coverage library, the semantic model and the user expected semantics.

[0057] The embodiment of the present application provides a method for generating a test case set, the method comprising: obtaining a data set to be tested, a model to be tested, a text description of a test task, a pre-stored model, and user expected semantics; when the type of the model to be tested includes a black box model, generating a proxy model based on the test data set and the pre-stored module; when the type of the model to be tested includes a white box model, determining the model to be tested as a proxy model; generating test cases based on the test data set, the text description of the test task, the pre-stored model, the proxy model, and user expected semantics. The above scheme expands the original test cases by evaluating the model to be tested and then selecting a corresponding method to generate test cases, and can improve the diversity and pertinence of the test data, ensuring that the model to be tested can show excellent performance in various complex situations, thereby ensuring the adequacy and comprehensiveness of the test.

[0058] like Figure 2 As shown, in one example, before executing the above step S104, the embodiment of the present application further provides an implementation method including:

[0059] S201: Determine whether the data volume of the test data set meets the preset requirements.

[0060] For example, the data can be evaluated based on the complexity of the task to see whether it meets the processing requirements. When the amount of test data provided by the user reaches the system requirements, it is considered to meet the requirements and is divided into a test set and a validation set for subsequent generation of proxy models and test cases.

[0061] S202: If the preset requirements are not met, a request message is sent to the database, so that the database can filter available test data according to the preset requirements to supplement the test data set.

[0062] On the contrary, when the data volume of the test data set provided by the user cannot meet the system requirements, a request message is sent to the database. The request message is used by the database to filter available test data according to preset requirements to supplement the test data set.

[0063] like Figure 3 As shown, in one example, the implementation of generating the proxy model based on the test data set and the pre-stored model in step S102 may include but is not limited to the following steps:

[0064] S301: Determine a replacement model in pre-stored models.

[0065] The replacement model in the above steps may be an available model selected by the user from the pre-stored models based on actual application scenario requirements.

[0066] S302: Test the alternative model based on the test data set to generate a prediction label.

[0067] In this step, after generating a predicted label by testing the alternative model, the predicted label can be further softened.

[0068] S303: Train the replacement model based on the predicted labels and the test dataset.

[0069] Exemplarily, the implementation of this step may include training the alternative model based on the predicted label and the test data set, and determining the output of the alternative model as the predicted output, and then modifying the parameters of the alternative model under the condition of consistency constraints on the predicted output and the predicted label until the alternative model converges.

[0070] S304: Determine the trained replacement model as the proxy model.

[0071] That is, after the above steps, if it is determined that the surrogate model after parameter correction converges, the trained surrogate model is determined as the proxy model and the proxy model is output.

[0072] like Figure 4 As shown, in one example, in the above step S102, the implementation method of generating a test case based on the test data set, the test task text description, the pre-stored model, the proxy model and the user expected semantics may include but is not limited to the following steps:

[0073] S401. Construct a use case expected output-model neuron coverage library based on the test label semantics of the test dataset and the proxy model inference.

[0074] Optionally, the proxy model can infer model output (i.e., the predicted output of the surrogate model) and neural network coverage output, where selectable neural network coverage output metrics may include neuron coverage, strong neuron coverage, top-k neuron coverage, etc. Furthermore, the neural network coverage can be stored in the form of a vector in the above-mentioned use case expected output - model neuron coverage library, using the predicted output as an index, for subsequent determination of test cases.

[0075] S402: Construct a semantic model based on the test task text description and the natural language processing model in the pre-stored model.

[0076] S403: Generate test cases based on the expected output of the use case-model neuron coverage library, the semantic model, and the user's expected semantics.

[0077] Exemplarily, the implementation method of this step may include generating a neuron coverage vector library of the user expected category based on the semantic model and the user expected semantics; then retrieving the difference between the neuron coverage vector library of the user expected category and the use case expected output-model neuron coverage library, and generating test cases in a gradient increasing manner.

[0078] Furthermore, in an embodiment of the present application, the following methods can be used to generate the above-mentioned test cases: first, directly use the difference between the user's expected semantics and the semantics output by the test data set to return the generated use case samples; second, calculate the test cases through the coverage difference between the expected neuron coverage vector library and the actual network output; third, change the semantic category through the user's expected semantics, search for coverage in the constructed neuron coverage vector library of the user's expected category, and calculate the difference with the current neural network coverage to generate test cases.

[0079] It should be noted that when the category of the above-mentioned model to be tested includes white box, the method of generating test cases based on the test data set and proxy model is the same as the implementation method in the black box mode, that is, the above three methods can also be used to generate test cases.

[0080] like Figure 4a As shown, in an example, the embodiment of the present application further provides an implementation method, including:

[0081] S401a: Obtain performance indicators and pre-stored test cases.

[0082] S402a. Determine the effectiveness of the test case based on the performance indicators.

[0083] S403a: If the test case is valid, the model to be tested is evaluated based on the test case and pre-stored test cases, and a test report is generated.

[0084] For example, the test cases obtained in the above steps can be quality-assessed. When the data meets various indicator requirements (e.g., test adequacy), the test cases are collected into a sample library, and the model is supplemented with information such as task type and usage intent. The pre-stored test cases are then mixed with the generated test cases to test and evaluate the model under test, and a test report is output.

[0085] The following takes the test case of generating image classification tasks as an example to further introduce the above implementation method in detail.

[0086] Step 1: Data quantity assessment. Receive the model to be tested and the test dataset provided by the user, and perform a quantitative assessment of the data provided by the user, and calculate the total amount and distribution of the statistics.

[0087] Step 2: Data Processing Requirements Assessment. Based on the complexity of the test task, the system evaluates whether the user-provided test dataset meets the processing requirements. For example, for image classification tasks, the user-provided test dataset is on the order of 1,000 images, while for detection tasks, the dataset is on the order of 10,000 images. If the user-provided test dataset meets the system requirements, the data is divided into a test set and a validation set. If the user-provided test dataset is insufficient, the system requests the database to filter out available test data that meets the requirements.

[0088] Step 3: Add the filtered data to the test data set and re-evaluate and divide the data (repeat Step 2). If the data set provided by the database still cannot meet the test requirements, generate a test case and execute Step 4.

[0089] Step 4: Determine the type of model to be tested. If it is a black-box model, proceed to the black-box proxy model generation module. Use the training data to test the model to be tested, obtain predicted labels, and soften the labels. Consistency constraints are imposed on the proxy model output and the output of the model to be tested using methods such as the KL distribution, and the proxy model parameters are adjusted. When the proxy model converges, output the proxy model and proceed to the next step.

[0090] Step 5: Combine the model to be tested or the proxy model converted by Step 4 with the training data to enter the white-box test case generation phase. Using the data provided by the user, specifically the test label semantics of the above-mentioned test data set can be adopted, combined with the proxy model inference to build the expected output of the use case and the model neuron coverage library. Use the NLP model (such as BERT, Word2Vec, etc.) in the pre-stored model to build a semantic model of user intention and expected data, so as to directly query test cases of multiple expected added categories through user semantic input. By retrieving the difference between the neuron coverage vector library of the user's expected category and the neuron coverage vector library of the input category, adversarial samples are generated by gradient increasing to improve the diversity and pertinence of the test cases, ensuring that the model also has excellent performance in various complex situations.

[0091] Step 6: Evaluate the quality of the acquired test cases to ensure that the data meets test adequacy and other requirements. Eligible test cases are added to the sample library and labeled with information such as task type and user intent. The user-provided test data is combined with the system-generated test case data to test the model under test and generate a test report.

[0092] Figure 5 A schematic diagram of the structure of a device for generating a test case set provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the apparatus may include: an acquisition module 501, a generation module 502;

[0093] The acquisition module is used to obtain the test data set, the model to be tested, the test task text description, the pre-stored model, and the user's expected semantics;

[0094] A generation module, configured to generate a proxy model based on a test dataset and a pre-stored model when the type of the model to be tested includes a black box model;

[0095] The generation module is further configured to determine the model to be tested as a proxy model when the type of the model to be tested includes a white box model;

[0096] The generation module is also used to generate test cases based on the test dataset, test task text description, pre-stored models, proxy models and user expected semantics.

[0097] like Figure 6 As shown, the above device may further include a judgment module 503;

[0098] The judgment module is used to judge whether the data volume of the test data set meets the preset requirements; and, if it does not meet the preset requirements, send a request message to the database so that the database can filter the available test data according to the preset requirements to supplement the test data set.

[0099] In one example, the above-mentioned generation module is specifically used to determine a substitute model in a pre-stored model; test the substitute model based on a test data set to generate a prediction label; train the substitute model based on the prediction label and the test data set; and determine the trained substitute model as a proxy model.

[0100] Furthermore, the generation module is also used to train the alternative model based on the predicted label and the test data set, and determine the output of the alternative model as the predicted output; based on the consistency constraint on the predicted output and the predicted label, the parameters of the alternative model are corrected until the alternative model converges.

[0101] In one example, a generation module is specifically configured to construct a use case expected output-model neuron coverage library based on the test label semantics of a test dataset and inferred by a proxy model; construct a semantic model based on the test task text description and a natural language processing model in a pre-stored model; and generate test cases based on the use case expected output-model neuron coverage library, the semantic model, and user expected semantics. For example, a neuron coverage vector library for a user expected category is generated based on the semantic model and user expected semantics; and the difference between the neuron coverage vector library for the user expected category and the use case expected output-model neuron coverage library is retrieved to generate test cases using a gradient-increasing method.

[0102] In one example, the acquisition module is further used to acquire performance indicators and pre-stored test cases;

[0103] The judgment module is also used to judge the effectiveness of the test case based on the performance indicators;

[0104] The generation module is also used to evaluate the model to be tested based on the test case and pre-stored test cases and generate a test report when the test case is valid.

[0105] The generator of the above test case set can execute Figures 1-4 The provided method for generating a test case set has corresponding devices and beneficial effects in the method.

[0106] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the computer device includes a controller 701, a memory 702, an input device 703, and an output device 704; the number of controllers 701 in the computer device can be one or more. Figure 7 In the figure, a controller 701 is taken as an example; the controller 701, memory 702, input device 703 and output device 704 in the computer device can be connected by a bus or other means. Figure 7 The bus connection is taken as an example.

[0107] The memory 702 is a computer-readable storage medium that can be used to store software programs, computer executable programs, and modules, such as Figure 1 The program instructions / modules corresponding to the test case set generation method in the embodiment (for example, the acquisition module 501, the generation module 502, etc. in the test case set generation device). The controller 701 executes the software programs, instructions, and modules stored in the memory 702 to perform various functions of the computer device and data processing, that is, to implement the above-mentioned test case set generation method.

[0108] The memory 702 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the computer. Furthermore, the memory 702 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 702 may further include memory remotely located relative to the controller 701, and these remote memories may be connected to a terminal / server via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0109] The input device 703 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the computer device. The output device 704 may include a display device such as a display screen.

[0110] The embodiment of the present application also provides a storage medium containing computer executable instructions, which, when executed by a computer controller, is used to execute a method for generating a test case set, the method comprising: Figure 1 Steps shown.

[0111] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present application can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer's floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0112] It is worth noting that the modules included in the above-mentioned test case set generation device are only divided according to functional logic, but are not limited to the above-mentioned division method. As long as the corresponding functions can be achieved, it is not used to limit the scope of protection of this application.

[0113] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the appended claims.

Claims

1. A method for generating a test case, characterized in that: The method comprises: Obtain the test dataset, the model to be tested, the test task text description, the pre-stored model, and the user's expected semantics; In a case where the type of the model to be tested includes a black box model, generating a proxy model based on the test data set and the pre-stored model; In a case where the type of the model to be tested includes a white box model, determining the model to be tested as a proxy model; A test case is generated according to the test data set, the test task text description, the pre-stored model, the proxy model and the user expected semantics.

2. The method according to claim 1, characterized in that Before generating a test case according to the test data set, the test task text description, the pre-stored model, the proxy model and the user expected semantics, the method further includes: Determine whether the data volume of the test data set meets the preset requirements; If the preset requirements are not met, a request message is sent to the database, so that the database can filter available test data according to the preset requirements to supplement the test data set.

3. The method according to claim 1, characterized in that The generating of the proxy model based on the test data set and the pre-stored model includes: determining an alternative model among the pre-stored models; Testing the alternative model based on the test data set to generate a predicted label; training the surrogate model based on the predicted labels and the test dataset; The trained surrogate model is determined as the surrogate model.

4. The method according to claim 3, characterized in that The training of the substitution model based on the predicted label and the test data set includes: Training the surrogate model based on the predicted label and the test data set, and determining the output of the surrogate model as the predicted output; The parameters of the surrogate model are modified under the condition that the predicted output and the predicted label are consistent with each other until the surrogate model converges.

5. The method according to any one of claims 1 to 4, characterized in that Generating a test case according to the test data set, the test task text description, the pre-stored model, the proxy model, and the user expected semantics includes: Constructing a use case expected output-model neuron coverage library based on the test label semantics of the test dataset and the agent model inference; Constructing a semantic model based on the test task text description and the natural language processing model in the pre-stored model; Test cases are generated according to the use case expected output-model neuron coverage library, the semantic model and the user expected semantics.

6. The method according to claim 5, characterized in that Generating a test case according to the use case expected output-model neuron coverage library, the semantic model and the user expected semantics includes: Generate a neuron coverage vector library of a user expected category according to the semantic model and the user expected semantics; The difference between the neuron coverage vector library of the user expected category and the neuron coverage library of the use case expected output-model is retrieved, and a test case is generated in a gradient increasing manner.

7. The method according to claim 1, characterized in that The method further comprises: Obtain performance metrics and pre-stored test cases; Determining the effectiveness of the test case based on the performance indicator; If the test case is valid, the model to be tested is evaluated based on the test case and the pre-stored test case, and a test report is generated.

8. A device for generating a test case set, characterized in that: The device comprises: The acquisition module is used to obtain the test data set, the model to be tested, the test task text description, the pre-stored model and the user's expected semantics; a generating module, configured to generate a proxy model based on the test data set and the pre-stored model when the type of the model to be tested includes a black box model; The generating module is further configured to, when the type of the model to be tested includes a white box model, determine the model to be tested as a proxy model; The generation module is further configured to generate a test case based on the test data set, the test task text description, the pre-stored model, the proxy model, and the user expected semantics.

9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for generating a test case set according to any one of claims 1 to 7 is implemented.

10. A device-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for generating a test case set according to any one of claims 1 to 7 is implemented.