Model testing method and device, storage medium and electronic equipment
By receiving and parsing model evaluation requests, determining the target defect category and test set, and generating and executing test tasks, the problem of inefficient testing in the existing technology is solved, and end-to-end model testing and defect detection in custom formats is realized.
Patent Information
- Application Number
- CN202510496158.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
Existing model testing methods cannot implement end-to-end testing, resulting in low testing efficiency.
By receiving the model evaluation request from the user side, analyzing the model name, defect category and test set name, extracting model attribute information, determining the target defect category and test set, generating and executing the target test task, and feedback the results to the user side.
End-to-end model testing is implemented, testing efficiency is improved, and defect detection is supported in custom formats.
Smart Images

Figure CN120336928A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to a method for testing a model, a device for testing a model, a computer-readable storage medium, and an electronic device. Background Art
[0002] In existing model testing methods, end-to-end model testing cannot be achieved, thereby resulting in low model testing efficiency.
[0003] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a method for testing a model, a device for testing a model, a computer-readable storage medium, and an electronic device, so as to at least overcome to some extent the problem of low model testing efficiency caused by the limitations and defects of related technologies.
[0005] According to one aspect of the present disclosure, there is provided a method for testing a model, including:
[0006] Receiving a model evaluation request sent by a user terminal, and parsing the model evaluation request to obtain a name of a model to be detected, a category of a defect to be detected, and a name of a test set;
[0007] Extracting model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determining a target defect category according to the category of the defect to be detected and the model attribute information;
[0008] Determining a target test set according to the name of the test set and the target defect category, and generating a target test task according to the model attribute information, the target defect category, and the target test set;
[0009] Executing the target test task to obtain a target task execution result, and feeding back the target task execution result to the user terminal.
[0010] In an exemplary embodiment of the present disclosure, the name of the model to be detected includes name information of the model itself and model version number information;
[0011] The category of the defect to be detected includes the category of product defects of industrial products that the model to be detected can detect, and / or the category of dataset defects of the target test set required for detecting the model to be detected;
[0012] The model attribute information includes at least one of the following: the original model training dataset and the original model validation dataset corresponding to the model to be detected, the original product defect categories that the model to be detected can detect, the original data volume of the original model training dataset, and the original model accuracy of the model after training the model to be detected based on the original model training dataset.
[0013] In an exemplary embodiment of the present disclosure, determining the target defect category according to the defect category to be detected and the model attribute information includes:
[0014] Classify the defect category to be detected to obtain the product defect category to be detected and the dataset defect category to be detected, and obtain the original product defect category corresponding to the model to be detected in the model attribute information;
[0015] Match the product defect category to be detected and the original product defect category to obtain a product defect category matching result, and determine the target defect category according to the product defect category matching result and the dataset defect category to be detected.
[0016] In an exemplary embodiment of the present disclosure, determining the target defect category according to the product defect category matching result and the dataset defect category to be detected includes:
[0017] In response to the product defect category matching result that all the product defect categories to be detected are included in the original product defect category and the first category quantity of the defect category to be detected is the same as the second category quantity of the original product defect category, use the defect category to be detected as the target defect category;
[0018] In response to the product defect category matching result that the product defect categories to be detected are not completely included in the original product defect category, or the first category quantity is not the same as the second category quantity, adjust the product defect categories to be detected, and determine the target defect category according to the adjusted product defect category and the dataset defect category to be detected.
[0019] In an exemplary embodiment of the present disclosure, adjusting the product defect category to be detected includes:
[0020] According to the product defect category matching result, determine a first product defect list that is included in the product defect category to be detected but not included in the original product defect category, and / or a second product defect list that is included in the original product defect category but not included in the product defect category to be detected;
[0021] Generate a first user prompt message based on the first product defect list and the second product defect list, and send the first user prompt message to the client, so that the client adjusts the product defect categories to be detected according to the first user prompt message.
[0022] In an exemplary embodiment of the present disclosure, determining a target test set according to the test set name and the target defect category includes:
[0023] Establish a category mapping relationship between the test set name and the product defect categories to be detected in the target defect category;
[0024] Based on the category mapping relationship, extract a first test set name having a mapping relationship with the product defect categories to be detected and a second test set name not having a mapping relationship from the test set name;
[0025] Based on the category mapping relationship, extract a first product defect category having a mapping relationship with the test set name and a second product defect category not having a mapping relationship from the product defect categories to be detected;
[0026] Filter out the second test set name in the test set name, and add a test set name corresponding to the second product defect category to the filtered test set name to obtain the target test set.
[0027] In an exemplary embodiment of the present disclosure, executing the target test task to obtain a task execution result includes:
[0028] Determine the target computing resources required to execute the target test task, and match a target execution node from the task execution nodes based on the target computing resources;
[0029] Execute the target test task based on the target execution node to obtain a target task execution result.
[0030] In an exemplary embodiment of the present disclosure, executing the target test task to obtain a target task execution result includes:
[0031] Execute a test task of the model to be detected in the defect dimension of the data set to obtain a first subtask execution result; and / or
[0032] Execute a test task of the model accuracy trend of the model to be detected on multiple different model versions to obtain a second subtask execution result; and / or
[0033] Execute a test task of the model accuracy of the model to be detected on the newly added product defect types to obtain a third subtask execution result; and / or
[0034] Execute the test task for the model differences of the model to be detected in different test environments to obtain the execution result of the fourth subtask; and / or
[0035] Execute the test task for the model inference speed trend of the model to be detected on multiple different model versions to obtain the execution result of the fifth subtask; and
[0036] Generate the execution result of the target task according to the execution result of the first subtask and / or the execution result of the second subtask and / or the execution result of the third subtask and / or the execution result of the fourth subtask and / or the execution result of the fifth subtask.
[0037] In an exemplary embodiment of the present disclosure, the test task on the defect dimension of the data set is used to detect whether the data set of the model to be detected exists, and / or whether the data volume in the data set meets the preset requirements, and / or whether there is redundant data in the data set;
[0038] The test task for the model accuracy trend on multiple different model versions is used to detect the accuracy of the model prediction results of the model to be detected with the same name but different model versions under the same test set;
[0039] The test task for the model accuracy on the newly added product defect type is used to detect the accuracy of the model prediction results of the model to be detected on the newly added product defect type;
[0040] The test task for the model differences in different test environments is used to detect the differences in the accuracy of the model prediction results of the model to be detected in the production environment and the accuracy of the model prediction results of the model to be detected in the test environment;
[0041] The test task for the model inference speed trend on multiple different model versions is used to detect the task execution time required for the model to be detected with the same name but different model versions to execute the prediction task under the same test set.
[0042] In an exemplary embodiment of the present disclosure, generating the execution result of the target task according to the execution result of the first subtask and / or the execution result of the second subtask and / or the execution result of the third subtask and / or the execution result of the fourth subtask and / or the execution result of the fifth subtask includes:
[0043] Call the first content template corresponding to the execution result of the first subtask, and generate the first content display result according to the execution result of the first subtask and the first content template; and / or
[0044] Call the second content template corresponding to the execution result of the second subtask, and generate a second content display result according to the execution result of the second subtask and the second content template; and / or
[0045] Call the third content template corresponding to the execution result of the third subtask, and generate a third content display result according to the execution result of the third subtask and the third content template; and / or
[0046] Call the fourth content template corresponding to the execution result of the fourth subtask, and generate a fourth content display result according to the execution result of the fourth subtask and the fourth content template; and / or
[0047] Call the fifth content template corresponding to the execution result of the fifth subtask, and generate a fifth content display result according to the execution result of the fifth subtask and the fifth content template; and
[0048] Generate a target task execution result according to the first content display result and / or the second content display result and / or the third content display result and / or the fourth content display result and / or the fifth content display result.
[0049] In an exemplary embodiment of the present disclosure, generating a first content display result according to the execution result of the first subtask and the first content template includes:
[0050] Extract the current dataset defect category included in the execution result of the first subtask, and determine the position of the current dataset defect category in the first content template;
[0051] According to the position of the current dataset defect category in the first content template, fill the defect value corresponding to the current dataset defect category into the first content template to obtain a first content filling result;
[0052] Generate dataset defect prompt information and dataset defect optimization information according to the current dataset defect category and the defect value, and fill the dataset defect prompt information and the dataset defect optimization information into the first content filling result to obtain a first content display result.
[0053] In an exemplary embodiment of the present disclosure, the model evaluation request is generated in the following manner:
[0054] In response to a selection operation on the model name list in the model test interface, determine the model name to be detected;
[0055] In response to a selection operation on the list of product defects supported by the model to be detected corresponding to the model name to be detected in the model test interface, determine the defect category to be detected;
[0056] In response to a selection operation on the dataset list associated with the model to be detected corresponding to the name of the model to be detected in the model test interface, determine the test set name;
[0057] Generate a model evaluation request associated with the model to be detected according to the name of the model to be detected, the defect category to be detected, and the test set name.
[0058] According to one aspect of the present disclosure, there is provided a test device for a model, including:
[0059] A request parsing module, configured to receive a model evaluation request sent by a user terminal and parse the model evaluation request to obtain the name of the model to be detected, the defect category to be detected, and the test set name;
[0060] A target defect category determination module, configured to extract model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determine a target defect category according to the defect category to be detected and the model attribute information;
[0061] A target test task generation module, configured to determine a target test set according to the test set name and the target defect category, and generate a target test task according to the model attribute information, the target defect category, and the target test set;
[0062] A target test task execution module, configured to execute the target test task, obtain a target task execution result, and feedback the target task execution result to the user terminal.
[0063] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the test method for a model described in any one of the above is implemented.
[0064] According to one aspect of the present disclosure, there is provided an electronic device, including:
[0065] A processor; and
[0066] A memory, configured to store executable instructions of the processor;
[0067] Wherein, the processor is configured to execute the test method for a model described in any one of the above by executing the executable instructions.
[0068] A method for testing a model provided by an embodiment of the present disclosure, on the one hand, by receiving a model evaluation request sent by a user terminal, parsing the model evaluation request to obtain the name of the model to be detected, the category of the defect to be detected, and the name of the test set; then extracting the model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determining the target defect category according to the category of the defect to be detected and the model attribute information; then determining the target test set according to the name of the test set and the target defect category, and generating a target test task according to the model attribute information, the target defect category, and the target test set; finally, executing the target test task to obtain the target task execution result, and feedbacking the target task execution result to the user terminal, realizing end-to-end model testing, and solving the problem that the model testing efficiency is relatively low in the prior art due to the inability to realize end-to-end model testing; on the other hand, since the model can be tested according to the name of the model to be detected, the category of the defect to be detected, and the name of the test set in the model evaluation request, defect detection in a custom format is realized.
[0069] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0070] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0071] Figure 1 A flowchart schematically showing a method for testing a model according to an exemplary embodiment of the present disclosure.
[0072] Figure 2 A schematic structural diagram of a backend in a model testing system according to an exemplary embodiment of the present disclosure.
[0073] Figure 3 A schematic diagram showing an interaction scenario between modules in a backend according to an exemplary embodiment of the present disclosure.
[0074] Figure 4 A schematic diagram showing an example of a model testing interface according to an exemplary embodiment of the present disclosure.
[0075] Figure 5 A schematic diagram showing an example of a relationship mapping scenario between a test data set - algorithm - product defect category - training data set according to an exemplary embodiment of the present disclosure.
[0076] Figure 6Schematic diagram showing an example graph of the name, number of training sets, and training accuracy information of an extracted defect according to an example embodiment of the present disclosure.
[0077] Figure 7 Schematic diagram showing a statistical table obtained by evaluating the training accuracy of a batch of algorithm models on multiple versions according to an example embodiment of the present disclosure.
[0078] Figure 8 Schematic diagram showing a statistical table obtained by evaluating the processing effect of a specified algorithm on newly added defect types according to an example embodiment of the present disclosure.
[0079] Figure 9 Schematic diagram showing a statistical table obtained by verifying the difference between the algorithm model in the production environment and the test environment according to an example embodiment of the present disclosure.
[0080] Figure 10 Schematic diagram showing a statistical table obtained by verifying the impact of defect inference code update on the inference speed according to an example embodiment of the present disclosure.
[0081] Figure 11 Schematic diagram showing an example scenario graph presenting the accuracy change trend in the form of a line chart according to an example embodiment of the present disclosure.
[0082] FIG. 12(a) and FIG. 12(b) schematically show an example scenario graph presenting accuracy data in tabular form according to an example embodiment of the present disclosure.
[0083] Figure 13 Schematic diagram showing an example structure diagram of a test device for a model according to an example embodiment of the present disclosure.
[0084] Figure 14 Schematic diagram showing an electronic device for implementing a test method of a model according to an example embodiment of the present disclosure. Detailed implementation
[0085] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0086] In addition, the drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in the form of software, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0087] In the evaluation process of some related defect detection models, it can be achieved in the following way: First, obtain the key fields of the annotation information file, and by reading the product serial number of the product image to be detected, query the information file (before detection) annotated by AOI (Automated Optical Inspection) and the information file (after detection) annotated by the defect detection model using the product serial number. Then, read these two files line by line; calculate the intersection over union (IoU) of the numerical values of the annotation key fields; compare and analyze the calculation results of the IoU (compare the results before and after detection) to obtain the evaluation index value; use the evaluation index value to count the data and obtain the final evaluation result of the detection data. At the same time, based on this method to evaluate the product defect detection model, although it is possible to clearly view indicators such as the product name, the number of defects in the annotation file corresponding to the product name, the number of defects in the prediction file, the IoU inference result, the confidence level, the number of over-detected defects, and the number of missed defects, and it is also possible to present algorithm evaluation indicators such as over-detection and missed detection for all defect types in the form of a table, so that the product side can evaluate the detection ability of this version of the deep learning algorithm for such products according to the final data. However, this method has the following defects: On the one hand, it is impossible to achieve end-to-end model testing, thus reducing the model testing efficiency; on the other hand, it is impossible to perform defect detection in a custom format; that is, testers cannot determine the dimensions to be tested according to actual needs; furthermore, it is impossible to give corresponding optimization suggestions based on the test results, thus reducing the iterative optimization efficiency of the model.
[0088] Based on this, in the exemplary embodiment, a model testing method is first provided. This method can run on terminal devices, servers, server clusters, cloud servers, etc.; of course, those skilled in the art can also run the method of the present disclosure on other platforms according to requirements, and no special limitation is made in this exemplary embodiment. Specifically, referring to Figure 1 as shown, the model testing method may include the following steps:
[0089] Step S110. Receive a model evaluation request sent by the user terminal, and parse the model evaluation request to obtain the name of the model to be detected, the category of the defect to be detected, and the name of the test set;
[0090] Step S120. Extract the model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determine the target defect category according to the category of the defect to be detected and the model attribute information;
[0091] Step S130. Determine the target test set according to the name of the test set and the target defect category, and generate a target test task according to the model attribute information, the target defect category, and the target test set;
[0092] Step 140. Execute the target test task, obtain the target task execution result, and feedback the target task execution result to the client.
[0093] In the model testing method described above, on the one hand, by receiving a model evaluation request sent by the client and parsing the model evaluation request, the name of the model to be detected, the category of defects to be detected, and the name of the test set are obtained; then, based on the name of the model to be detected, the model attribute information corresponding to the model to be detected is extracted, and the target defect category is determined according to the category of defects to be detected and the model attribute information; then, the target test set is determined according to the name of the test set and the target defect category, and based on the model attribute information, the target defect category, and the target test set, a target test task is generated; finally, the target test task is executed, the target task execution result is obtained, and the target task execution result is feedback to the client, realizing end-to-end model testing and solving the problem in the prior art that the model testing efficiency is relatively low due to the inability to achieve end-to-end model testing; on the other hand, since model testing can be performed according to the name of the model to be detected, the category of defects to be detected, and the name of the test set in the model evaluation request, defect detection in a custom format is realized.
[0094] Hereinafter, the model testing method described in the exemplary embodiments of the present disclosure will be further explained and illustrated with reference to the accompanying drawings.
[0095] First, the technical implementation principle of the exemplary embodiments of the present disclosure will be explained and illustrated. Specifically, in the actual process of product defect detection of industrial products, due to the excessive variety of product defects, the variety of corresponding industrial defect detection models is also particularly large (for example, different sites, different product backgrounds, and different defect forms lead to a variety of different product defect categories); moreover, new defect categories are constantly emerging on the production line of industrial products; therefore, how to quickly obtain the iterative effect of the model version of the product defect detection model (that is, the classification situation and accuracy situation of product defects, etc.) so as to achieve the purpose of improving the defect analysis speed of the production line and quickly improving the defect interception efficiency has become an urgent problem to be solved. For this reason, the exemplary embodiments of the present disclosure provide a model testing method, which quickly completes the end-to-end algorithm effect evaluation analysis through a pipeline process and gives optimization suggestions, providing strong support for the rapid iteration of the algorithm and helping to improve the yield and quality of the production environment.
[0096] Secondly, the test system of the model involved in the exemplary embodiments of the present disclosure will be explained and described. Specifically, the test system of the model may include a client and a backend, and the client is communicatively connected to the backend. At the same time, the client described herein is the user interface display side, which is used to display the model test interface, so that the user can perform human-computer interaction based on the model test interface and then generate a corresponding model evaluation request. The backend described herein refers to the computing side that executes the model test task, and the backend can be implemented based on a local computer, or based on a server or a cloud server. This example does not make special restrictions on this.
[0097] It should also be supplemented and explained here that for the test system of the model described in the exemplary embodiments of the present disclosure, on the one hand, for algorithm developers, they can not only select multiple versions of the model for trend evaluation, but also evaluate the change trend of training accuracy with the model version. At the same time, if the model supports multiple types of defects, the change trend of the accuracy of each type of defect with the version can also be obtained. Moreover, multiple versions of the algorithm can be selected to evaluate the effect of increasing or decreasing training data, and the accuracy change of the algorithm with the increase or decrease of training data can be verified, so as to achieve the purpose of improving the algorithm development efficiency. On the other hand, for algorithm testers, algorithm accuracy comparison tests can be implemented: after batch-selecting the defects and data sets supported by the algorithm, select the intersection, union, and difference set of the defects and the corresponding defects in the data set for defect index evaluation. At the same time, algorithm processing time comparison tests can also be implemented: perform algorithm processing time comparison tests for the same algorithm under different environmental resources, so as to achieve the purpose of improving the test efficiency. On the other hand, for managers / algorithm users, they can select the same algorithm and environmental resources as the production environment to verify the processing effect of the algorithm on the specified data set, so as to estimate the effect of the algorithm in the production environment.
[0098] Furthermore, referring to Figure 2 As shown, the backend may include a test set management module 210, a defect management module 220, an attribution layer management module 230, an algorithm management module 240, a training set management module 250, a template management module 260, and an algorithm evaluation module 270. In the actual application process, a pipeline-style test system can be formed based on the test set management module 210, the defect management module 220, the attribution layer management module 230, the algorithm management module 240, the training set management module 250, the template management module 260, and the algorithm evaluation module 270, which can not only evaluate the algorithm accuracy, processing speed, and defect interception trend in multiple dimensions and from multiple angles, but also give targeted reminders of the next optimization points to the users of the test system. Next, the roles played by each management module in the actual model test process will be further explained and described. Figure 3 For each management module in the actual model test process will be further explained and described.
[0099] (1) Test set management module 210, which is responsible for managing the data set. The main contents of this module are as follows: ① Object: the name of the data set; ② Attributes: the specific content of the data set, the quantity, the data set creation time, the data set update time, the mapping relationship between the data set and the defect (one-to-many horizontally and many-to-many vertically), etc.; ③ Actions: create a data set, update a data set, delete a data set, and query a data set, etc.
[0100] (2) Defect management module 220, which is responsible for managing defects. The main contents of this module are as follows: ① Object: the name of the product defect; ② Attributes: the description of the product defect category (1 to N), the mapping relationship between the product defect category and the corresponding data set of the category; ③ Actions: input product defect information, edit product defect information, delete product defect information, and search for product defect information, etc.
[0101] (3) Defect attribution layer management module 230; specifically, in the actual application process, for iterative algorithm services, the defect categories to be detected are usually fixed; therefore, by setting up the defect attribution layer management module, it is convenient to use the batch management function to select the target test defects. Further, the main contents of this module are as follows: ① Objects: the name of the attribution layer, associated defects, description information; ② Actions: create a defect attribution layer, delete a defect attribution layer, modify a defect attribution layer, search for a defect attribution layer; and when creating the defect attribution layer management module, N types of defects can be selected and added in the layer, and the same defect can also be added in different layers.
[0102] (4) Algorithm management module 240, which is responsible for algorithm management, including algorithm models (i.e., defect detection models) and traditional algorithms. The main contents of this module are as follows: ① Object: the name of the algorithm; ② Attributes: model file / algorithm package, the product defect types supported by the model / algorithm, the version information of the model / algorithm, the dimensions of the algorithm evaluation indicators and the expected results; and when the algorithm supports multiple different product defect categories, indicators such as the overall average accuracy of the algorithm and the average accuracy of each type of defect can be obtained; ③ Actions: create a model entry, establish the association relationship between the model and the training set.
[0103] (5) Training set management module 250, which is responsible for managing the training set. The main contents of this module are as follows: ① Object: the name of the training set; ② Attributes: the content of the training set, the quantity of the training set, the name of the corresponding product defect; ③ Actions: create a training set, update a training set, delete a training set, and query a training set. And the same product defect name may correspond to one or more training sets.
[0104] (6) Template Management Module 260, which can, based on the model, data, and defects selected by the user, combined with the algorithm evaluation results, use the built-in templates of the system to complete the output of the algorithm's business capabilities. Among them, the built-in template strategy can be as follows: ① When both the test set and the training set are selected, automatically complete the comparison of the categories of the test set and the training set, obtain the defects of "defects missing the training set and missing the test set", and at the end of the algorithm evaluation metrics, automatically add an optimization prompt: In this round of evaluation, it is found that there are defects with a training set but no test set: X1, X2, X3, and defects with a test set but no training set: Y1, Y2, Y3. In the next iteration of the algorithm, it is recommended to supplement the above defects, with a recommended quantity of more than XX sheets. ② When both the test set and the training set are selected, automatically complete the comparison of the quantity of the test set, the quantity of the training set, and the built-in quantity threshold, obtain the defects of "defects with insufficient test set quantity, insufficient training set quantity", and at the end of the algorithm evaluation metrics, automatically add an optimization prompt: In this round of evaluation, it is found that there are defects with insufficient training set quantity of N: X1, X2, X3, and defects with insufficient test set quantity of N: Y1, Y2, Y3. In the next iteration of the algorithm, it is recommended to supplement the test set quantity for defect X (at least N1 more sheets need to be supplemented), and the training set quantity for defect Y (at least N2 more sheets need to be supplemented). ③ When testing different model versions of the same defect, automatically form an index trend line chart; among them, the abscissa is the version, the ordinate is the index dimension, and the expected index is displayed in the form of a target line; and, according to the change status of the trend value, automatically add an optimization prompt: Comparing with the historical iteration versions of the statistical algorithm, the average XX rate of XX defect has increased by XX / decreased by XX, and defect accuracy optimization is required.
[0105] (7) Algorithm Evaluation Module 270, which mainly completes the following tasks: ① Select the model file, select the defect to be tested, select the test set to be used, select the hardware resources, and create a test task; ② Start the test task and complete the model test; at the same time, in the specific test process, the training set will be automatically associated and the evaluation template will be automatically loaded. Further, there are many selection strategies supported by the algorithm evaluation, such as: verifying the processing ability of the model on the specified data set, verifying the generalization ability trend of different versions of the model on the fixed data set, and verifying the accuracy of the defects that are both in the list of defect types supported by the model and in the test set list, etc.
[0106] Hereinafter, the specific generation process of the model evaluation request will be explained and described. Specifically, the specific generation process of the model evaluation request can be implemented in the following manner: in response to a selection operation on the model name list in the model test interface, determine the model name to be detected; in response to a selection operation on the list of product defects supported by the model to be detected corresponding to the model name to be detected in the model test interface, determine the defect category to be detected; in response to a selection operation on the list of data sets associated with the model to be detected corresponding to the model name to be detected in the model test interface, determine the test set name; generate a model evaluation request associated with the model to be detected according to the model name to be detected, the defect category to be detected, and the test set name. Among them, the model test interface described here can be referred to Figure 4 as shown; in the actual application process, the user can, according to actual needs, determine the model to be tested, the defect category to be tested, the test set name required, etc.; of course, the environmental information required to execute the test task can also be determined, etc., and this example does not make special restrictions on this; when the user side detects an interaction event acting on the interaction control for starting the evaluation task, a model evaluation request associated with the model to be detected can be generated according to the selected model name to be detected, the defect category to be detected, and the test set name, and the model evaluation request is sent to the backend. Based on this, the problem that custom-format defect detection cannot be performed in the existing solution can be solved, and testers can, according to actual needs, determine the model to be tested and the dimension to be tested, so as to improve the accuracy of the task execution result obtained on the basis of improving the test efficiency.
[0107] Hereinafter, it will be combined with Figures 2 - 4 to Figure 1 the test method of the model shown in
[0108] In step S110, receive the model evaluation request sent by the user side, and parse the model evaluation request to obtain the model name to be detected, the defect category to be detected, and the test set name.
[0109] Specifically, the model name to be detected described here may include the name information of the model itself and the model version number information; for example, the scratch detection model V2.0; the defect category to be detected described here includes the product defect category of industrial products that the model to be detected can detect, and / or the data set defect category of the target test set required for detecting the model to be detected; for example, the product defect category of industrial products may include, but is not limited to, scratch defect category, breakage defect category, chipping defect category, etc., and the data set defect category of the target test set may include, but is not limited to, lack of training set, lack of test set, too little data volume included in the training set, or too little data volume included in the test set, etc.
[0110] It should be noted here that during the actual process of testing the network model, the test data set can be divided into a public test set and a dedicated test set according to the corresponding relationship between the data set category, defect category, and algorithm. Among them, the public test set refers to a data set that can be used for at least 2 to N types of defects; the dedicated test set refers to a data set that only corresponds to one product defect category. Similarly, the training data set is also classified into a public training set and a dedicated training set. Further, when performing a test task, different types and versions of algorithm models and / or algorithm combinations can be selected according to different business requirements, and different algorithms can support 1 to N types of defect detections. Among them, the relationship mapping between the test data set - algorithm - product defect category - training data set can be referred to Figure 5 as shown
[0111] In step S120, model attribute information corresponding to the model to be detected is extracted based on the name of the model to be detected, and the target defect category is determined according to the defect category to be detected and the model attribute information
[0112] In this exemplary embodiment, first, model attribute information corresponding to the model to be detected is extracted. Among them, the model attribute information recorded here may include, but is not limited to, the original model training dataset and the original model validation dataset corresponding to the model to be detected, the original product defect categories that the model to be detected can detect, the original data volume of the original model training dataset, and the original model accuracy of the model after training the model to be detected based on the original model training dataset. That is to say, the model attribute information can be used to represent the training set, validation set, model function, specific data volume, and model accuracy associated with the model, etc. On this premise, when extracting the model attribute information, the file / software package corresponding to the model to be tested can be selected based on the mapping relationship between the test dataset - algorithm - product defect category - training dataset. At the same time, if the dataset is fixed, batch selection of multiple algorithm versions to be evaluated is supported. It should be added here that the mapping relationship between the model to be detected and the dataset can be a one-to-one mapping or a one-to-many, many-to-one, or many-to-many mapping. In the actual application process, the mapping relationships in different scenarios are different. For example, when testing the accuracy improvement of a certain model in multiple versions, a fixed dataset M (a test set containing defects 1, 2, N) and multiple model versions will be selected. Another example is when testing the processing effects of a certain type of product image in three algorithm models, a fixed test set M1 (this type of product image, such as black or white images, etc.) and 3 algorithm models will be selected. In actual application, the system will automatically extract the name of each defect participating in the training, the training set quantity, and the training accuracy information based on the name of the model to be detected, the training set information corresponding to the model, and the training accuracy information. Among them, the specific extraction results can be referred to Figure 6 as shown.
[0113] Secondly, the target defect category is determined according to the defect category to be detected and the model attribute information. Specifically, it can be achieved in the following way: classify the defect category to be detected to obtain the product defect category to be detected and the dataset defect category to be detected, and obtain the original product defect category corresponding to the model to be detected in the model attribute information; match the product defect category to be detected and the original product defect category to obtain the product defect category matching result, and determine the target defect category according to the product defect category matching result and the dataset defect category to be detected. That is, in the actual application process, it is first necessary to determine whether the dataset defect category to be detected is included in the target defect category. If so, it is necessary to separate the product defect category to be detected and the dataset defect category to be detected, and then further determine whether the product defect category to be detected needs to be adjusted, so as to achieve the purpose of improving the accuracy of the obtained target defect category.
[0114] In an exemplary embodiment, the target defect category is determined according to the product defect category matching result and the defect category of the dataset to be detected, which can be implemented in the following manner: in response to the product defect category matching result that all the defect categories of the product to be detected are included in the original product defect categories, and the number of the first categories of the defect categories to be detected is the same as the number of the second categories of the original product defect categories, the defect categories to be detected are used as the target defect category; in response to the product defect category matching result that the defect categories of the product to be detected are not completely included in the original product defect categories, or the number of the first categories is not the same as the number of the second categories, the defect categories of the product to be detected are adjusted, and the target defect category is determined according to the adjusted product defect category and the defect category of the dataset to be detected.
[0115] In an exemplary embodiment, the adjustment of the defect categories of the product to be detected can be implemented in the following manner: according to the product defect category matching result, a first list of product defects that are included in the defect categories of the product to be detected but not in the original product defect categories, and / or a second list of product defects that are included in the original product defect categories but not in the defect categories of the product to be detected are determined; a first user prompt message is generated according to the first list of product defects and the second list of product defects, and the first user prompt message is sent to the user terminal so that the user terminal adjusts the defect categories of the product to be detected according to the first user prompt message.
[0116] Hereinafter, the specific determination process of the target defect category will be further explained and described. Specifically, in the actual application process, the user can directly search and select one or more defect categories by the defect names to be detected, or directly select a batch of defect categories through the defect category list displayed by the defect attribution layer; at the same time, when the user selects which defect categories need to be tested by the model to be detected, the network model test system will automatically intersect and associate these defect categories with the aforementioned extracted model attribute information to obtain the names of the defects to be tested this time, the number of defect training sets, etc.; at the same time, it will also obtain which defect categories are not within the scope of model training and which defects trained by the model are not tested this time, and the information will be prompted to the user, and the user further confirms the target defects to be tested this time (that is, adjusts the defect categories of the product to be detected), and finally, the target defect category is determined according to the adjusted product defect category and the defect category of the dataset to be detected.
[0117] In step S130, a target test set is determined according to the test set name and the target defect category, and a target test task is generated according to the model attribute information, the target defect category and the target test set.
[0118] In this exemplary embodiment, first, a target test set is determined according to the test set name and the target defect category. Specifically, it can be implemented in the following way: Establish a category mapping relationship between the test set name and the defect categories of the products to be detected in the target defect category; Based on the category mapping relationship, extract from the test set name the first test set name that has a mapping relationship with the defect category of the product to be detected and the second test set name that has no mapping relationship; Based on the category mapping relationship, extract from the defect categories of the products to be detected the first product defect category that has a mapping relationship with the test set name and the second product defect category that has no mapping relationship; Filter out the second test set name in the test set name, and add to the filtered test set name the test set name corresponding to the second product defect category to obtain the target test set. That is, in the actual application process, when the user selects the test set to be used this time, the test system of the model will automatically extract which test sets have corresponding target defects, which test sets do not have corresponding target defects, and which target defects do not have test sets, and then make supplements or filters based on the corresponding comparison results to obtain the target test set. Of course, it is also necessary to establish a linkage relationship between the test set - target defect - training set to facilitate the construction of the corresponding target test task.
[0119] Secondly, after the target defect category and the target test set are determined, a target test task can be generated according to the model attribute information, the target defect category, and the target test set. Among them, in the obtained target test task, the target defect category and the test set name can be presented in the form of a list, so that during the task execution process, it can be executed item by item.
[0120] In step S140, execute the target test task to obtain a target task execution result, and feedback the target task execution result to the user terminal.
[0121] In this exemplary embodiment, first, execute the target test task to obtain a target task execution result. Specifically, it can be implemented in the following way: Determine the target computing resources required to execute the target test task, and match a target execution node from the task execution nodes based on the target computing resources; Execute the target test task based on the target execution node to obtain a target task execution result. Specifically, during the matching of the target execution node, it can be implemented by the user's independent selection or by the system's automatic matching. This exemplary embodiment does not make special restrictions on this; At the same time, if it needs to be implemented by the user's independent selection, the system will automatically calculate the required resources after the user determines the name of the model to be detected, the defect category to be detected, and the test set name, and display the nodes that meet the resources for the user to select.
[0122] In an exemplary embodiment, the execution of the target test task to obtain the target task execution result can be achieved in the following manner: execute the test task of the model to be detected in the defect dimension of the data set to obtain the first sub-task execution result; and / or execute the test task of the model accuracy trend of the model to be detected on multiple different model versions to obtain the second sub-task execution result; and / or execute the test task of the model accuracy of the model to be detected on the newly added product defect types to obtain the third sub-task execution result; and / or execute the test task of the model difference of the model to be detected under different test environments to obtain the fourth sub-task execution result; and / or execute the test task of the model inference speed trend of the model to be detected on multiple different model versions to obtain the fifth sub-task execution result; and, based on the first sub-task execution result and / or the second sub-task execution result and / or the third sub-task execution result and / or the fourth sub-task execution result and / or the fifth sub-task execution result, generate the target task execution result; wherein, the test task in the defect dimension of the data set recorded herein is used to detect whether the data set of the model to be detected exists, and / or whether the data volume in the data set meets the preset requirements, and / or whether there is redundant data in the data set; the test task of the model accuracy trend on multiple different model versions recorded herein is used to detect the accuracy of the model prediction results of the models to be detected with the same name but different model versions under the same test set; the test task of the model accuracy on the newly added product defect types recorded herein is used to detect the accuracy of the model prediction results of the model to be detected on the newly added product defect types; the test task of the model difference under different test environments recorded herein is used to detect the difference between the accuracy of the model prediction results of the model to be detected in the production environment and the accuracy of the model prediction results of the model to be detected in the test environment; the test task of the model inference speed trend on multiple different model versions recorded herein is used to detect the task execution time required for the models to be detected with the same name but different model versions to execute the prediction task under the same test set.
[0123] Further, during the execution of the target test task, different test environments can be unified into a set of the same standards; for example, a unified model file processing interface, a unified algorithm service creation interface, a unified test task creation interface, and a unified test task start interface, etc., and the above interfaces all define unified input and output formats; further, when the user clicks to start the test task, a series of interfaces of the target test environment are automatically called, so as to achieve a pipeline-style evaluation execution; finally, the test results are output by calling the algorithm application program interface and other methods.
[0124] It should be noted here that during the process of testing the model, not only the defects in the data set need to be tested, but also the accuracy trend of the model needs to be tested, the accuracy of the model on the newly added product defect categories needs to be tested, the accuracy difference of the model in different environments needs to be tested, and the detection speed of different versions of the model needs to be tested. Therefore, multiple test tasks in different dimensions need to be executed, which can be determined according to actual needs. This example does not make special restrictions on this. Of course, in the actual application process, other dimensions or categories can be added to the tests required. This example is only for illustrative purposes and has no other restrictions.
[0125] In an exemplary embodiment, the target task execution result is generated according to the execution result of the first subtask and / or the execution result of the second subtask and / or the execution result of the third subtask and / or the execution result of the fourth subtask and / or the execution result of the fifth subtask, which can be achieved in the following manner: calling the first content template corresponding to the execution result of the first subtask, and generating a first content display result according to the execution result of the first subtask and the first content template; and / or calling the second content template corresponding to the execution result of the second subtask, and generating a second content display result according to the execution result of the second subtask and the second content template; and / or calling the third content template corresponding to the execution result of the third subtask, and generating a third content display result according to the execution result of the third subtask and the third content template; and / or calling the fourth content template corresponding to the execution result of the fourth subtask, and generating a fourth content display result according to the execution result of the fourth subtask and the fourth content template; and / or calling the fifth content template corresponding to the execution result of the fifth subtask, and generating a fifth content display result according to the execution result of the fifth subtask and the fifth content template; and, generating the target task execution result according to the first content display result and / or the second content display result and / or the third content display result and / or the fourth content display result and / or the fifth content display result.
[0126] In an exemplary embodiment, generating the first content display result according to the execution result of the first subtask and the first content template can be achieved in the following manner: extracting the current data set defect category included in the execution result of the first subtask, and determining the position of the current data set defect category in the first content template; according to the position of the current data set defect category in the first content template, filling the defect value corresponding to the current data set defect category into the first content template to obtain a first content filling result; according to the current data set defect category and the defect value, generating data set defect prompt information and data set defect optimization information, and filling the data set defect prompt information and the data set defect optimization information into the first content filling result to obtain the first content display result.
[0127] In an exemplary embodiment, generating a second content display result according to the execution result of the second subtask and the second content template can be achieved in the following manner: Extract the first model accuracy test results obtained by different model versions included in the execution result of the second subtask under the same data set, and generate a first accuracy change trend graph according to the first model accuracy test results; Determine the display position of the first accuracy change trend graph in the second content template, and fill the first accuracy change trend graph into the second content template according to the display position of the first accuracy change trend graph in the second content template to obtain a second content filling result; Generate a first accuracy prompt message and an optimal model version message according to the first model accuracy test results obtained by different model versions under the same data set, and fill the first accuracy prompt message and the optimal model version message into the second content filling result to obtain a second content display result.
[0128] In an exemplary embodiment, generating a third content display result according to the execution result of the third subtask and the third content template can be achieved in the following manner: Extract the second model accuracy test results of the model on the newly added product defect types included in the execution result of the third subtask, and determine the display position of the second model accuracy test results in the third content template; Fill the second model accuracy test results into the third content template according to the display position of the second model accuracy test results in the third content template to obtain a third content filling result; Generate a second accuracy prompt message according to the second model accuracy test results, and fill the second accuracy prompt message into the third content filling result to obtain a third content display result.
[0129] In an exemplary embodiment, generating a fourth content display result according to the execution result of the fourth subtask and the fourth content template can be achieved in the following manner: Extract the third model accuracy test results of the model in the production environment and the fourth model accuracy test results in the test environment included in the execution result of the fourth subtask, and obtain a model difference test result according to the third model accuracy test results and the fourth model accuracy test results; Determine the display position of the model difference test result in the fourth content template, and fill the model difference test result into the fourth content template according to the display position of the model difference test result in the fourth content template to obtain a fourth content filling result; Generate a difference prompt message according to the model difference test result, and fill the difference prompt message into the fourth content filling result to obtain a fourth content display result.
[0130] In an exemplary embodiment, generating a fifth content display result according to the execution result of the fifth subtask and the fifth content template can be achieved in the following manner: Extract the model inference speeds of the model included in the execution result of the fifth subtask on multiple different model versions, and generate a trend chart of the inference speed change based on the model inference speeds; determine the display position of the trend chart of the inference speed change in the fifth content template, and fill the trend chart of the speed change into the fifth content template according to the display position of the trend chart of the inference speed change in the fifth content template to obtain a fifth content filling result; generate inference speed prompt information based on the model inference speeds, and fill the inference speed prompt information into the fifth content filling result to obtain a fifth content display result.
[0131] So far, the testing method of the model recorded in the exemplary embodiments of the present disclosure has been fully implemented. Hereinafter, specific test processes will be further explained and illustrated with specific examples. Specifically, taking model effect evaluation, defect accuracy evaluation, algorithm ability trend analysis, etc. as examples, the specific test processes will be further explained and illustrated. Specifically:
[0132] Example 1: Evaluate the change trend of the training accuracy of a batch of algorithm models on multiple versions.
[0133] Background description: Developers need to track the algorithm training effect and longitudinally compare the accuracy trend changes of different algorithm versions. At the same time, during the actual testing process, the user side selects the names of the algorithm models M1, …, Mx, selects the version numbers to be evaluated (Vx1, Vx2, Vxn), the target test set (the same test set is used for multiple versions of the same model, and different test sets are used for different models), selects the running environment (defaulting to the current default configuration if not selected), clicks to start the algorithm evaluation task, and waits for the execution result. Finally, on the one hand, the back end of the evaluation system associates the corresponding training set and training accuracy information according to the algorithm model name selected by the user to form a model-version-training set-training accuracy association relationship, and draws a version-accuracy line chart in the form of the abscissa being the version and the ordinate being the model accuracy in the order of the versions. On the other hand, it automatically evaluates the test set and the training set to form a statistical table with the model type in the horizontal row and the difference dimension in the vertical column; among them, the obtained statistical table can be referred to Figure 7 as shown. Further, after the evaluation task is completed, the user side can view the display content of the trend chart and the statistical table.
[0134] Example 2: Evaluate the processing effect of a specified algorithm on newly added defect types (such as black pictures).
[0135] Background description: The tester verifies the black image recognition effect of the model. At the same time, during the actual test, the user selects the name M1 of the algorithm model and the algorithm version Vx1, selects the test sets SDx1 (pure black images), SDx2 (a dataset with other defects having large black blocks), and SDx3 (a defect test set not trained by this algorithm model) required for evaluation, and then executes the algorithm evaluation task and waits for the execution result. Further, the back-end of the evaluation system will perform the model evaluation process on the test sets and form a chart of the test results on each test set. Specifically, it can be referred to Figure 8 as shown, and the user can clearly see the evaluation effect of the algorithm model.
[0136] Example 3: Verify the difference between the algorithm model in the production environment and the test environment.
[0137] Background description: Regarding the large gap between the recognition effect of the algorithm model in the production environment and the expectation as feedback by the user, or the special missed detection problem in the production environment, it is necessary to conduct differential verification and analysis in a test environment with the same configuration. At the same time, during the actual test, the user selects the algorithm model name M2 and the algorithm version Vx2, selects the test set required for verification (this test set uses the images tested or retrieved in the production environment), selects the test environment with the same configuration as the target production environment, and then executes the algorithm evaluation task and waits for the execution result. The back-end of the evaluation system performs the evaluation process, feedbacks the test results of the algorithm model on this test set and forms a chart. Specifically, it can be referred to Figure 9 as shown. The user can directly compare with the effect in the production environment.
[0138] Example 4: Verify the impact of defect inference code update on the inference speed.
[0139] Background description: After the inference code is updated, compare whether there is a large delay in the processing speed of different versions of the inference code. At the same time, during the actual test, the user selects the fixed algorithm model name M3 and the algorithm version Vx3, selects the fixed test set, selects different versions of the inference environment, and then executes the algorithm evaluation task and waits for the execution result. The back-end of the evaluation system executes different inference codes according to the selection of different inference environments (such as calling different API interfaces), completes the evaluation process, feedbacks the test results of different versions of the inference code on the fixed test set and forms a chart (specifically, it can be referred to Figure 10 as shown), and the user can directly see the processing speed of different versions of the inference code (in the following table, two algorithm services such as 5983 and 5988 process the images of defects 1 to 5 respectively, and a set of comparative processing times are obtained).
[0140] It should also be noted here that the specific presentation form of the target task execution result obtained by the model testing method described in the exemplary embodiments of the present disclosure on the user side can be to display the accuracy change trend in the form of a line chart, or to display the accuracy data in the form of a table. Among them, for the specific method of displaying the accuracy change trend in the form of a line chart, reference can be made to Figure 11 as shown. For the method of displaying the accuracy data in the form of a table, reference can be made to FIGS. 12(a) and 12(b).
[0141] Furthermore, as Figure 11 shown, the model can be selected, the defect type can be selected (in the example figure, "all" is selected to identify all defects, or a single defect can be selected to view the accuracy trend), and clicking on the "draw line chart" can show the illustrated effect. In this figure, the abscissa is different model versions, and the ordinates are recall rate and mAP50 respectively; from Figure 11 it can be seen the change trend of the accuracy index with the version and the impact of the increase in training data on the model trend.
[0142] Even further, as shown in FIGS. 12(a) and 12(b), it can display the test results of all defects in an attribution layer. The specified columns display the indicators of the specific identifiers in the display area that do not meet the expected results, and the specified rows display the defect types of the specific identifiers in the display area that detect special marked defects. At the same time, based on the results shown in FIG. 12, it is very convenient to evaluate and analyze problems for users to locate, so as to achieve the purpose of improving development efficiency and testing efficiency.
[0143] Finally, based on the foregoing content, it can be known that the model testing method described in the exemplary embodiments of the present disclosure has at least the following advantages: on the one hand, the evaluators can freely select multi-dimensional testing processes (such as single-class / multi-class defects, single / multiple versions, multiple accuracy indicators), which is convenient for comprehensively evaluating the algorithm from all directions; on the other hand, the pipeline-type testing system enables the testers to quickly verify the iterative effect of the algorithm model, improve the model evaluation speed, and further improve the interception speed and interception rate of production line defects; further, the model testing method described in the exemplary embodiments of the present disclosure is applicable to both the evaluation of the end-to-end effect of the algorithm after engineering and the evaluation of the effect of model training iteration, and can quickly give the effect evaluation of the algorithm and the targeted next optimization points.
[0144] The following are the embodiments of the present disclosure apparatus, which can be used to execute the embodiments of the present disclosure method. For the details not disclosed in the embodiments of the present disclosure apparatus, please refer to the embodiments of the present disclosure method.
[0145] The exemplary embodiments of the present disclosure also provide a model testing apparatus. Specifically, refer to Figure 13As shown, the test device for the model may include a request parsing module 1310, a target defect category determination module 1320, a target test task generation module 1330, and a target test task execution module 1340. Among them:
[0146] The request parsing module 1310 can be used to receive a model evaluation request sent by the client and parse the model evaluation request to obtain the name of the model to be detected, the defect category to be detected, and the name of the test set; the target defect category determination module 1320 can be used to extract model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determine the target defect category according to the defect category to be detected and the model attribute information; the target test task generation module 1330 can be used to determine the target test set according to the name of the test set and the target defect category, and generate a target test task according to the model attribute information, the target defect category, and the target test set; the target test task execution module 1340 can be used to execute the target test task, obtain a target task execution result, and feedback the target task execution result to the client.
[0147] In an exemplary embodiment of the present disclosure, the name of the model to be detected includes the name information of the model itself and the model version number information; the defect category to be detected includes the product defect category of the industrial product that the model to be detected can detect, and / or the dataset defect category of the target test set required for detecting the model to be detected; the model attribute information includes at least one of the following: the original model training dataset and the original model validation dataset corresponding to the model to be detected, the original product defect category that the model to be detected can detect, the original data volume of the original model training dataset, and the original model accuracy of the model after training the model to be detected based on the original model training dataset.
[0148] In an exemplary embodiment of the present disclosure, determining the target defect category according to the defect category to be detected and the model attribute information includes: classifying the defect category to be detected to obtain the product defect category to be detected and the dataset defect category to be detected, and obtaining the original product defect category corresponding to the model to be detected in the model attribute information; matching the product defect category to be detected and the original product defect category to obtain a product defect category matching result, and determining the target defect category according to the product defect category matching result and the dataset defect category to be detected.
[0149] In an exemplary embodiment of the present disclosure, determining the target defect category according to the product defect category matching result and the defect category of the dataset to be detected includes: in response to the product defect category matching result indicating that all defect categories of the product to be detected are included in the original product defect categories and the number of the first categories of the defect categories to be detected is the same as the number of the second categories of the original product defect categories, taking the defect categories to be detected as the target defect category; in response to the product defect category matching result indicating that the defect categories of the product to be detected are not completely included in the original product defect categories or the number of the first categories is different from the number of the second categories, adjusting the defect categories of the product to be detected, and determining the target defect category according to the adjusted product defect categories and the defect category of the dataset to be detected.
[0150] In an exemplary embodiment of the present disclosure, adjusting the defect categories of the product to be detected includes: according to the product defect category matching result, determining a first list of product defects that are included in the defect categories of the product to be detected but not included in the original product defect categories, and / or a second list of product defects that are included in the original product defect categories but not included in the defect categories of the product to be detected; generating first user prompt information according to the first list of product defects and the second list of product defects, and sending the first user prompt information to the user terminal so that the user terminal adjusts the defect categories of the product to be detected according to the first user prompt information.
[0151] In an exemplary embodiment of the present disclosure, determining the target test set according to the test set name and the target defect category includes: establishing a category mapping relationship between the test set name and the defect categories of the product to be detected in the target defect category; based on the category mapping relationship, extracting a first test set name that has a mapping relationship with the defect categories of the product to be detected and a second test set name that has no mapping relationship from the test set name; based on the category mapping relationship, extracting a first product defect category that has a mapping relationship with the test set name and a second product defect category that has no mapping relationship from the defect categories of the product to be detected; filtering out the second test set name in the test set name, and adding a test set name corresponding to the second product defect category to the filtered test set name to obtain the target test set.
[0152] In an exemplary embodiment of the present disclosure, performing the target test task to obtain a task execution result includes: determining the target computing resources required to perform the target test task, and matching a target execution node from the task execution nodes based on the target computing resources; performing the target test task based on the target execution node to obtain a target task execution result.
[0153] In an exemplary embodiment of the present disclosure, performing the target test task to obtain a target task execution result includes: performing a test task of the model to be detected on the defect dimension of the dataset to obtain a first subtask execution result; and / or performing a test task of the model accuracy trend of the model to be detected on multiple different model versions to obtain a second subtask execution result; and / or performing a test task of the model accuracy of the model to be detected on the newly added product defect type to obtain a third subtask execution result; and / or performing a test task of the model difference of the model to be detected under different test environments to obtain a fourth subtask execution result; and / or performing a test task of the model inference speed trend of the model to be detected on multiple different model versions to obtain a fifth subtask execution result; and, generating the target task execution result according to the first subtask execution result and / or the second subtask execution result and / or the third subtask execution result and / or the fourth subtask execution result and / or the fifth subtask execution result.
[0154] In an exemplary embodiment of the present disclosure, the test task on the defect dimension of the dataset is used to detect whether the dataset of the model to be detected exists, and / or whether the data volume in the dataset meets the preset requirements, and / or whether there is redundant data in the dataset; the test task of the model accuracy trend on multiple different model versions is used to detect the accuracy of the model prediction results of the model to be detected with the same name but different model versions under the same test set; the test task of the model accuracy on the newly added product defect type is used to detect the accuracy of the model prediction results of the model to be detected on the newly added product defect type; the test task of the model difference under different test environments is used to detect the difference between the accuracy of the model prediction results of the model to be detected in the production environment and the accuracy of the model prediction results of the model to be detected in the test environment; the test task of the model inference speed trend on multiple different model versions is used to detect the task execution time required for the model to be detected with the same name but different model versions to perform the prediction task under the same test set.
[0155] In an exemplary embodiment of the present disclosure, generating the target task execution result according to the first subtask execution result and / or the second subtask execution result and / or the third subtask execution result and / or the fourth subtask execution result and / or the fifth subtask execution result includes: invoking a first content template corresponding to the first subtask execution result, and generating a first content display result according to the first subtask execution result and the first content template; and / or invoking a second content template corresponding to the second subtask execution result, and generating a second content display result according to the second subtask execution result and the second content template; and / or invoking a third content template corresponding to the third subtask execution result, and generating a third content display result according to the third subtask execution result and the third content template; and / or invoking a fourth content template corresponding to the fourth subtask execution result, and generating a fourth content display result according to the fourth subtask execution result and the fourth content template; and / or invoking a fifth content template corresponding to the fifth subtask execution result, and generating a fifth content display result according to the fifth subtask execution result and the fifth content template; and generating a target task execution result according to the first content display result and / or the second content display result and / or the third content display result and / or the fourth content display result and / or the fifth content display result.
[0156] In an exemplary embodiment of the present disclosure, generating a first content display result according to the first subtask execution result and the first content template includes: extracting the current data set defect category included in the first subtask execution result, and determining the position of the current data set defect category in the first content template; filling the defect value corresponding to the current data set defect category into the first content template according to the position of the current data set defect category in the first content template to obtain a first content filling result; generating a data set defect prompt message and a data set defect optimization message according to the current data set defect category and the defect value, and filling the data set defect prompt message and the data set defect optimization message into the first content filling result to obtain a first content display result.
[0157] In an exemplary embodiment of the present disclosure, the model evaluation request is generated in the following manner: in response to a selection operation on the model name list in the model test interface, determining the model name to be detected; in response to a selection operation on the product defect list supported by the model to be detected corresponding to the model name to be detected in the model test interface, determining the defect category to be detected; in response to a selection operation on the data set list associated with the model to be detected corresponding to the model name to be detected in the model test interface, determining the test set name; and generating a model evaluation request associated with the model to be detected according to the model name to be detected, the defect category to be detected, and the test set name.
[0158] The specific details of each module in the test device of the above model have been described in detail in the test method of the corresponding model, so they will not be elaborated here.
[0159] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0160] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0161] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as circuits, modules, or systems here.
[0162] Reference is now made to Figure 14 describe the electronic device 1400 according to this embodiment of the present disclosure. Figure 14 The electronic device 1400 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0163] As Figure 14 shown, the electronic device 1400 is presented in the form of a general-purpose computing device. The components of the electronic device 1400 may include, but are not limited to: at least one of the above processing units 1410, at least one of the above storage units 1420, a bus 1430 connecting different system components (including the storage unit 1420 and the processing unit 1410), and a display unit 1440.
[0164] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1410, so that the processing unit 1410 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 1410 can execute as Figure 1Step S110 shown in : Receive a model evaluation request sent by a client, and parse the model evaluation request to obtain the name of the model to be detected, the category of defects to be detected, and the name of the test set; Step S120: Extract model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determine the target defect category according to the category of defects to be detected and the model attribute information; Step S130: Determine the target test set according to the name of the test set and the target defect category, and generate a target test task according to the model attribute information, the target defect category, and the target test set; Step S140: Execute the target test task to obtain a target task execution result, and feedback the target task execution result to the client.
[0165] The storage unit 1420 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 14201 and / or a cache storage unit 14202, and may further include a read-only storage unit (ROM) 14203.
[0166] The storage unit 1420 may further include a program / utilities 14204 having a set (at least one) of program modules 14205. Such program modules 14205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0167] The bus 1430 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0168] The electronic device 1400 can also communicate with one or more external devices 1500 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1400, and / or communicate with any device that enables the electronic device 1400 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1450. Moreover, the electronic device 1400 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1460. As shown in the figure, the network adapter 1460 communicates with other modules of the electronic device 1400 through the bus 1430. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0169] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0170] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is further provided, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0171] The program product for implementing the above method according to the embodiments of the present disclosure can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.
[0172] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium (an exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0173] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0174] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.
[0175] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (for example, by connecting through the Internet service provider via the Internet).
[0176] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0177] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention herein disclosed. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not invented by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A method for testing a model, characterized in that, Including: Receiving a model evaluation request sent by a user terminal, parsing the model evaluation request to obtain a name of a model to be detected, a category of a defect to be detected, and a name of a test set; Extracting model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determining a target defect category according to the category of the defect to be detected and the model attribute information; Determining a target test set according to the name of the test set and the target defect category, and generating a target test task according to the model attribute information, the target defect category, and the target test set; Executing the target test task to obtain a target task execution result, and feeding back the target task execution result to the user terminal.
2. The testing method of the model according to claim 1, characterized in that The name of the model to be detected includes name information of the model itself and model version number information; The category of the defect to be detected includes a product defect category of an industrial product that the model to be detected can detect, and / or a data set defect category of a target test set required for detecting the model to be detected; The model attribute information includes at least one of the following: an original model training data set and an original model validation data set corresponding to the model to be detected, an original product defect category that the model to be detected can detect, an original data volume of the original model training data set, and an original model accuracy of the model after being trained based on the original model training data set; 3. The test method of the model according to claim 1, characterized in that Determining the target defect category according to the category of the defect to be detected and the model attribute information includes: Classifying the category of the defect to be detected to obtain a product defect category to be detected and a data set defect category to be detected, and obtaining an original product defect category corresponding to the model to be detected in the model attribute information; Matching the product defect category to be detected and the original product defect category to obtain a product defect category matching result, and determining the target defect category according to the product defect category matching result and the data set defect category to be detected.
4. The test method of the model according to claim 3, characterized in that, Determining the target defect category according to the product defect category matching result and the data set defect category to be detected includes: In response to the product defect category matching result that all of the product defect categories to be detected are included in the original product defect category and the number of the first categories of the defect categories to be detected is the same as the number of the second categories of the original product defect category, taking the defect category to be detected as the target defect category; In response to the product defect category matching result that the product defect categories to be detected are not completely included in the original product defect category, or the number of the first categories is not the same as the number of the second categories, adjusting the product defect categories to be detected, and determining the target defect category according to the adjusted product defect categories and the data set defect category to be detected.
5. The testing method of the model according to claim 4, characterized in that, Adjusting the product defect categories to be detected includes: Determining a first product defect list that is included in the product defect categories to be detected but not included in the original product defect category, and / or a second product defect list that is included in the original product defect category but not included in the product defect categories to be detected according to the product defect category matching result; Generate a first user prompt message based on the first product defect list and the second product defect list, and send the first user prompt message to the user terminal so that the user terminal adjusts the product defect categories to be detected according to the first user prompt message.
6. The testing method of the model according to claim 1, characterized in that Determine a target test set based on the test set name and the target defect category, including: Establish a category mapping relationship between the test set name and the product defect categories to be detected in the target defect category; Based on the category mapping relationship, extract a first test set name that has a mapping relationship with the product defect categories to be detected and a second test set name that has no mapping relationship from the test set name; Based on the category mapping relationship, extract a first product defect category that has a mapping relationship with the test set name and a second product defect category that has no mapping relationship from the product defect categories to be detected; Filter out the second test set name in the test set name, and add a test set name corresponding to the second product defect category to the filtered test set name to obtain the target test set.
7. The test method of the model according to claim 1, characterized in that Execute the target test task to obtain a task execution result, including: Determine the target computing resources required to execute the target test task, and match a target execution node from the task execution nodes based on the target computing resources; Execute the target test task based on the target execution node to obtain a target task execution result.
8. The testing method of the model according to claim 7, characterized in that Execute the target test task to obtain a target task execution result, including: Execute a test task on the dataset defect dimension of the model to be detected to obtain a first subtask execution result; and / or Execute a test task on the model accuracy trend of the model to be detected on multiple different model versions to obtain a second subtask execution result; and / or Execute a test task on the model accuracy of the model to be detected on new product defect types to obtain a third subtask execution result; and / or Execute a test task on the model difference of the model to be detected in different test environments to obtain a fourth subtask execution result; and / or Execute a test task on the model inference speed trend of the model to be detected on multiple different model versions to obtain a fifth subtask execution result; and Generate the target task execution result according to the first subtask execution result and / or the second subtask execution result and / or the third subtask execution result and / or the fourth subtask execution result and / or the fifth subtask execution result.
9. The test method of the model according to claim 8, characterized in that, The test task on the dataset defect dimension is used to detect whether the dataset of the model to be detected exists, and / or whether the data volume in the dataset meets the preset requirements, and / or whether there is redundant data in the dataset; The test task on the model accuracy trend of multiple different model versions is used to detect the accuracy of the model prediction results of the model to be detected with the same name but different model versions under the same test set; The test task on the model accuracy of new product defect types is used to detect the accuracy of the model prediction results of the model to be detected on new product defect types; The test task of the model differences under different test environments is used to detect the differences between the accuracy of the model prediction results of the model to be detected in the production environment and the accuracy of the model prediction results of the model to be detected in the test environment; The test task of the model inference speed trend on multiple different model versions is used to detect the task execution time required for the model to be detected with the same name but different model versions to perform the prediction task under the same test set.
10. The testing method of the model according to claim 8, characterized in that, Generating the target task execution result according to the execution result of the first subtask and / or the execution result of the second subtask and / or the execution result of the third subtask and / or the execution result of the fourth subtask and / or the execution result of the fifth subtask, including: Invoking the first content template corresponding to the execution result of the first subtask, and generating the first content display result according to the execution result of the first subtask and the first content template; and / or Invoking the second content template corresponding to the execution result of the second subtask, and generating the second content display result according to the execution result of the second subtask and the second content template; and / or Invoking the third content template corresponding to the execution result of the third subtask, and generating the third content display result according to the execution result of the third subtask and the third content template; and / or Invoking the fourth content template corresponding to the execution result of the fourth subtask, and generating the fourth content display result according to the execution result of the fourth subtask and the fourth content template; and / or Invoking the fifth content template corresponding to the execution result of the fifth subtask, and generating the fifth content display result according to the execution result of the fifth subtask and the fifth content template; and Generating the target task execution result according to the first content display result and / or the second content display result and / or the third content display result and / or the fourth content display result and / or the fifth content display result.
11. The testing method of the model according to claim 10, characterized in that, Generating the first content display result according to the execution result of the first subtask and the first content template, including: Extracting the current data set defect category included in the execution result of the first subtask, and determining the position of the current data set defect category in the first content template; Filling the defect value corresponding to the current data set defect category into the first content template according to the position of the current data set defect category in the first content template, to obtain the first content filling result; Generating the data set defect prompt information and the data set defect optimization information according to the current data set defect category and the defect value, and filling the data set defect prompt information and the data set defect optimization information into the first content filling result, to obtain the first content display result.
12. The test method of the model according to any one of claims 1-10, characterized in that, The model evaluation request is generated in the following manner: Responding to the selection operation on the model name list in the model test interface, determining the model name to be detected; Responding to the selection operation on the list of product defects supported by the model to be detected corresponding to the model name to be detected in the model test interface, determining the defect category to be detected; Responding to the selection operation on the data set list associated with the model to be detected corresponding to the model name to be detected in the model test interface, determining the test set name; Generate a model evaluation request associated with the model to be detected according to the name of the model to be detected, the category of defects to be detected, and the name of the test set.
13. A test device for a model, characterized in that, Including: A request parsing module, configured to receive a model evaluation request sent by a client, and parse the model evaluation request to obtain the name of the model to be detected, the category of defects to be detected, and the name of the test set; A target defect category determination module, configured to extract model attribute information corresponding to the model to be detected based on the name of the model to be detected, and determine a target defect category according to the category of defects to be detected and the model attribute information; A target test task generation module, configured to determine a target test set according to the name of the test set and the target defect category, and generate a target test task according to the model attribute information, the target defect category, and the target test set; A target test task execution module, configured to execute the target test task, obtain a target task execution result, and feedback the target task execution result to the client.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the test method of the model according to any one of claims 1-12.
15. An electronic device, characterized in that, Including: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the test method of the model according to any one of claims 1-12 by executing the executable instructions.