Visual inspection AI platform training and test report generation method
By performing feature screening and clustering on the training and test results of the visual inspection AI platform and designing a reasonable report format, we have solved the problems of low report generation efficiency and inconsistent formats in existing technologies. This has enabled the rapid generation of standard and easy-to-read reports and supported automatic iterative optimization of model effects.
Patent Information
- Application Number
- CN202510768607.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-16
AI Technical Summary
The existing visual inspection AI platform training and test report generation methods have problems such as high labor costs, low efficiency, inconsistent report formats, incomplete content and long generation time. Especially when the model effect is poor, report generation takes too long or the information is incomplete, affecting the platform functional experience.
By performing statistics and calculations on the training and test results, screening features and clustering them, reasonably screening the number of maps, and designing a reasonable report format, including sample display, parameters and result pages, an automated algorithm is used to generate standard and easy-to-read reports.
It enables the rapid generation of reports with standard formats, comprehensive content and easy-to-read content, reduces labor costs, improves communication efficiency, enhances model effects and supports automatic iterative optimization.
Smart Images

Figure CN120656020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual inspection technology, and in particular to a method for visual inspection AI platform training and test report generation. Background Art
[0002] Existing methods for generating training and test reports for visual inspection AI platforms include manual report writing and automatic formatting. Manual report writing involves people writing reports in a given format, filling in and pasting the data and images provided by the AI platform software into the report. This method is not only labor-intensive and inefficient, but may also result in different report writing styles due to the different report writers, resulting in no standard format for the report. When there are a large number of reports, it is necessary to understand the format before reading them, which increases the reading cost.
[0003] The automatic formatting generation method uses software instead of manual work to interact with the AI software platform and generate reports based on the given format. The reports are divided into full content reports and partial content reports according to the amount of information in the reports.
[0004] The full content report: that is, all content information is always output to the report. When the model effect is good, the output content volume is not large, the full content input information is complete, and the time consumption is not very long. However, when the model effect is very poor, such as too many over-killing and missed detections, and too many maps are required, it will be very time-consuming, resulting in a long time to generate the report, seriously affecting the platform function experience and even causing the system to crash due to the long waiting time.
[0005] You can choose to cancel the mapping to avoid abnormal situations, but this will no longer include problem samples such as missed detection, over-detection, incomplete detection, and misclassification in the report. This will lose a lot of valuable information and significantly reduce the value of using the report to analyze model problems.
[0006] Partial content report: that is, outputting partial information. When the model effect is not good, resulting in too many sample images such as over-killing and missed detection that need to be pasted, a certain number of samples are selected sequentially, evenly, or randomly for pasting, so as to avoid the problem of taking too long when the model effect is poor. However, since the number of samples selected for pasting is fixed and the appearance, characteristics, similarity and other information of the samples are not combined, the information coverage of the pasted sample images is insufficient, and it cannot provide strong support for analyzing the model effect and improving the strategy.
[0007] Therefore, reports automatically generated in a fixed format may appear rigid and difficult to read due to unreasonable or non-universal format design. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a method for generating training and test reports for a visual inspection AI platform, so as to solve the demand problem that after training and testing, the AI platform needs to quickly and automatically generate a report file with standard format, comprehensive content, easy to read and reference value.
[0009] The technical solution adopted by the present invention to solve its technical problem is: a method for training and generating test reports for a visual inspection AI platform, comprising the following steps:
[0010] S1. After the test, statistics and calculations are performed on the obtained system parameters and results;
[0011] S2. Screening the training samples; the screening includes feature structure screening and feature clustering screening;
[0012] S3. Perform sample screening on the test results; the screening includes the calculation of the number of stickers and sticker screening.
[0013] Furthermore, in step S1 of the present invention, the test result includes the output label and true label of each test sample, the output label and true label of each training sample, and the image file path of each training sample and test sample.
[0014] Furthermore, in step S1 of the present invention, the system parameters include engineering parameters, structural parameters of the model itself, and hyperparameters during training and testing.
[0015] Furthermore, in step S1 of the present invention, the calculation includes obtaining the detection results and calculating the intersection between all detections and the true annotations on the corresponding samples.
[0016] Furthermore, in step S2 of the present invention, the method of feature structure screening is:
[0017] 1) Calculate the shape hu moment of the annotation, which has 7 orders in total, forming a 7-dimensional feature;
[0018] 2) Calculate the grayscale / color histogram of the annotated part of the image, quantize the histogram into 8 levels, and form an 8-dimensional / 24-dimensional feature. For unannotated samples, that is, positive samples, fill the 8-dimensional / 24-dimensional feature with 0;
[0019] 3) Calculate the grayscale / color histogram of the image outside the annotation area, quantize the histogram into 8 levels, and form 8-dimensional / 24-dimensional features;
[0020] 4) Perform PCA dimensionality reduction on the features of the layer before the network output layer to 8 dimensions, forming 8-dimensional features;
[0021] 5) The features obtained above are concatenated to form features for training sample screening.
[0022] Furthermore, in step S2 of the present invention, the feature clustering screening method is:
[0023] a) Use the K-means clustering method to cluster the positive samples into K classes based on the features of the constructed positive samples, where K is the number of the current class sample (positive sample) maps set;
[0024] b) After clustering the positive samples, find the center of each subclass obtained by clustering the positive samples. In each subclass, select the sample closest to the center as the positive sample for display, that is, obtain K positive samples for display;
[0025] c) For each category, select samples for display according to the above method and paste them in the corresponding positions according to the set requirements.
[0026] Furthermore, in step S3 of the present invention, the map screening strategy is as follows: detect each type of map in the map, construct features according to the feature construction method, calculate the center coordinates, and divide the sample features into 5 intervals according to the distance from the center point using the mean square error formula. Each interval is allocated x candidate places, where x = T / 5, and T is the number threshold of the detection of this type. If an interval has less than x candidates, then the interval is used as the full selection interval, and the remaining places allocated by the interval are evenly distributed to other intervals with less allocated places than its sample number, thereby completing the allocation of 50 places.
[0027] When using samples for network inference, the features of the layer before the output layer are used as the network features of the samples. Taking missed detection as an example, the network feature distance between the samples in the non-selected interval and each sample in the positive samples in the training sample set and each sample in the interval is calculated. The minimum distance between each sample in the interval and all positive samples is taken as the inter-class distance of the sample. All samples in the interval are sorted from small to large according to the inter-class distance, and then samples are selected according to the allocated quota order of the interval as mapping samples.
[0028] Furthermore, the present invention also includes:
[0029] Step S4, optimization strategy: sort each detection situation according to the inter-class distance, use the sorted inter-class distance as the dependent variable to construct a distance function, select the point with the largest second-order derivative of the function as the inflection point, select the sample to the left of the inflection point as the candidate sample, cluster the candidate samples based on network features, and determine the K value according to k-means combined with the elbow rule. In each class, the sample with the closest distance to the center point within the class is used as the optimal sample. Select the optimal sample for each detection situation according to the above method, add the optimal sample to the training set for retraining, and generate the report again.
[0030] The beneficial effect of the present invention is that it solves the defects existing in the background technology, quickly and automatically generates reports with comprehensive information and standard format, and can quickly and conveniently display on-site training and information to superiors or R&D personnel and customers, providing a reference basis for subsequent work, improving the communication efficiency between the project site and all parties involved, and reducing communication costs; at the same time, it can be used for iterative optimization to improve the effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the overall structure of the report automatically generated after adopting the present invention;
[0032] Figure 2 It is a flow chart of the report generation steps of the present invention. DETAILED DESCRIPTION
[0033] The present invention will now be described in further detail with reference to the accompanying drawings and preferred embodiments. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0034] like Figure 1-Figure 2 The method for generating training and test reports for a visual inspection AI platform shown in the figure makes the report comprehensive and easy to read by designing a reasonable report format, which has great reference value. It can also limit the number of images used for display and intelligently filter the displayed images, avoiding the time-consuming execution of mapping when there are many detections, and achieving the overall optimal balance between generation speed and information volume.
[0035] The specific steps are as follows:
[0036] I. Overall Structure of the Report
[0037] like Figure 1 As shown, the report consists of three pages: sample display page, parameter page, and result page. The detailed information of each page is as follows:
[0038] 1. Sample display page:
[0039] This is the page showing the training samples. Each row contains samples from one category. The page is divided into two parts: the original sample image and the corresponding labeled image.
[0040] 2. Parameter page:
[0041] This is a page that records various types of data, such as model parameters and hyperparameters. For the models used by the AI platform to process the current project, their relevant parameters are listed separately. In this paper, the classification model and segmentation model are used;
[0042] 3. Results page:
[0043] First, a table is used to display the statistics of the training and test data of each model. Then, the result statistical table of each model is listed. For the segmentation model, it is a detection rate table, and for the classification model, it is a confusion matrix. The third part is a map of missed and falsely detected samples.
[0044] 2. Report Generation Steps
[0045] The overall steps are as follows Figure 2 As shown;
[0046] 1. Statistics and calculation of parameters and results
[0047] (1) After the test is completed, the test results are obtained:
[0048] The output label and true label of each test sample;
[0049] The output label and true label of each training sample;
[0050] The image file path of each training sample and test sample;
[0051] And system parameters:
[0052] 1) Project parameters: This refers to treating the current training and testing process as a project. Parameters that can describe and identify this project include: model version number and model generation time. These two parameters are the project parameters used in this invention, and also include various other project parameters specified according to specific needs;
[0053] 2) Model parameters, namely the structural parameters of the model itself and the hyperparameters during training and testing, include:
[0054] Model channel,
[0055] Model size: the size of the input image,
[0056] Number of iterations: that is, the number of training rounds,
[0057] Optimization times: Based on the report generated after a training test, the network is retrained using the missed and falsely detected samples shown in the report as a training set to test the network again, thereby regenerating the report. This process is called a optimization search and can be iterated multiple times, i.e., multiple optimization searches.
[0058] Whether to enable padding: When enabled, the model will expand the input sample size smaller than the model's specified size with a value of 0 around the image so that the sample reaches the specified size. If not enabled, it will be directly scaled to the specified size.
[0059] Padding size: the number of pixels to extend the edge. It is not required by default. When a value is given, the image is extended according to the given value and the extended image is scaled to the size specified by the model.
[0060] Downsampling coefficient: Not required by default. When a value is set, the network feature map and image size are downsampled to improve network execution efficiency.
[0061] Batch processing times: batch size,
[0062] Test step length: When the test image is a large image that needs to be cut, the overlap ratio when using the window sliding cut,
[0063] Minimum threshold: the detection score threshold,
[0064] Connected domain: When the intersection of the detection result and the true label is greater than this value, it is considered a correct detection.
[0065] Maximum number of defects: that is, the maximum number of connected domains detected on a sample. Set this value to n_m. When the number of defects detected on a sample is greater than n_m, sort them by area and only take the first n_m defects.
[0066] Label name and output name, minimum area and maximum area, whether to output: This group of parameters handles some special defect types. The label name is the name of the defect in the label, the output name is the name of the defect displayed on the interface, the minimum and maximum areas are the lower and upper thresholds for filtering defects based on area, and whether to output is whether to enable the processing function for this special type of defect;
[0067] The parameters obtained above are platform-related parameters based on the present invention. The specific parameters depend on the specific application platform. Fill in the corresponding positions with the obtained parameters as required.
[0068] (2) Then calculate the result parameters based on the obtained results
[0069] 1) Obtain the detection results and all connected domains in the results. Each connected domain is counted as a detection. The model workflow of the present invention is: first try to use the segmentation model to perform binary classification on the input sample, then extract the part of the segmentation result (i.e., the connected domain) corresponding to the sample image and send it to the classification network for classification. All subsequent calculation rules are calculated based on this workflow, and corresponding adjustments can be made for other workflows;
[0070] 2) Calculate the intersection between all detections and the true annotations on the corresponding samples,
[0071] If the detection overlaps with the true annotation on the corresponding sample,
[0072] If the intersection pixels are greater than or equal to the connected domain threshold (the connected domain parameter above), then for the segmentation task, it is a correct detection. For the classification task, if the category it belongs to is consistent with the labeled category, it is a correct detection. Otherwise, it is a misclassification. The number of misclassified detections is recorded as n_m;
[0073] If the intersection pixels are smaller than the connected domain threshold, it is considered an incomplete detection, and the number of incomplete detections is recorded as n_p;
[0074] If the detection has no intersection with the true annotation on the corresponding sample, it is considered an overkill, and the total number of overkills is recorded as n_o;
[0075] Then count whether the true annotation on each sample has an intersection with the detection result. If there is no intersection, it is considered a missed detection, and the total number of missed detections is recorded as n_e;
[0076] The total number of detections is recorded as N_r, and the total number of labels is recorded as N_l;
[0077] Calculate the result parameters based on the above data:
[0078] Missed detection rate = n_e / N_l;
[0079] Overkill rate = n_o / N_r;
[0080] Fill in the above obtained data and calculated result parameters into the corresponding positions on the result page;
[0081] 2. Training sample screening
[0082] (1) Feature construction
[0083] 1) Calculate the shape hu moment of the annotation, which has 7 orders in total, forming a 7-dimensional feature;
[0084] 2) Calculate the grayscale / color histogram of the annotated part of the image, quantize the histogram into 8 levels, and form an 8-dimensional / 24-dimensional feature. For unannotated samples, that is, positive samples, fill the 8-dimensional / 24-dimensional feature with 0;
[0085] 3) Calculate the grayscale / color histogram of the image outside the annotation area, quantize the histogram into 8 levels, and form 8-dimensional / 24-dimensional features;
[0086] 4) Perform PCA dimensionality reduction on the features of the layer before the network output layer to 8 dimensions, forming 8-dimensional features;
[0087] 5) Combine the features obtained above to form features for training sample screening;
[0088] (2) Feature clustering
[0089] The samples of each category are clustered using the following method:
[0090] Category refers to the label category of the sample, including positive samples and negative samples of various defect categories. Take the positive sample as an example:
[0091] 1) The features of the positive samples constructed above are clustered into K classes using the K-means clustering method, where K is the number of the set current class sample (positive sample) maps, and in this invention, K=2;
[0092] 2) After clustering the positive samples, find the centers of each subclass obtained by clustering the positive samples. In each subclass, select the sample closest to the center as the positive sample for display, that is, obtain K positive samples for display;
[0093] 3) Each category selects samples for display according to the above method and pastes them to the corresponding position. The number of clusters for all categories of samples in the present invention is K, and K is set to 2 in the present invention;
[0094] 3. Test result sample screening
[0095] (1) Calculation of the number of maps
[0096] 1) Setting the threshold for the number of textures
[0097] For samples that need to be mapped, the detection status is divided into: missed detection, over-detection, incomplete detection, and misclassification; the quantity threshold is the maximum number of maps for various detection situations.
[0098] Missed detection: T_e, default 50;
[0099] Overkill: T_o, default 20;
[0100] Incomplete detection: T_p, default 20;
[0101] Misclassification: T_m, default 20;
[0102] 2) Texture filtering
[0103] For each detection graph, if the number of detections is not greater than the corresponding threshold, all are selected. If the number exceeds the threshold, the following strategy is used to filter:
[0104] For each type of graph in the detection graph (classified by detection status), construct features according to the feature construction method in step 2, calculate the center coordinates, and divide the sample features into 5 intervals according to the distance from the center point. The distance formula is the mean square error formula. Each interval is allocated x candidate quotas, x = T / 5, T is the number threshold of the detection of this type. If an interval has less than x, then the interval is used as the full selection interval, and the remaining quotas allocated to the interval are evenly distributed to other intervals with less allocated quotas than its sample number, thus completing the allocation of 50 quotas;
[0105] When using samples to pass network inference, the features of the previous layer of the output layer are used as the network features of the samples. Taking missed detection as an example, the network feature distance between the samples in the non-selected interval and each sample in the positive samples in the training sample set and each sample in the interval is calculated. The minimum distance between each sample in the interval and all positive samples is taken as the inter-class distance of the sample. All samples in the interval are sorted from small to large according to the inter-class distance, and then samples are selected according to the allocated quota order of the interval as mapping samples.
[0106] Similarly, the inter-class distance is calculated between the over-detected sample and the classified category, the inter-class distance is calculated between the misclassified sample and the true label category, and the inter-class distance is calculated between the incomplete detection and the positive sample.
[0107] After selecting the mapping sample according to the above method, map it as required;
[0108] 3. Optimization Strategy
[0109] The report generated by the present invention supports the optimization function, that is, the training set is readjusted based on the report and the data generated during the report generation process, thereby improving the performance of the network;
[0110] Sort each detection situation according to the inter-class distance, and use the sorted inter-class distance as the dependent variable to form a distance function (that is, the inter-class distance is the y value and the x coordinate is the sorting number). Select the point with the largest second-order derivative of the function as the inflection point, select the sample to the left of the inflection point as the candidate sample, cluster the candidate samples based on network characteristics, and determine the K value according to the k-means combined with the elbow rule. In each class, the sample with the closest distance to the center point within the class is used as the optimal sample. Select the optimal sample for each detection situation according to the above method, add the optimal sample to the training set for retraining, and generate a report again. In this way, the effect changes and improvement degree before and after optimization can be compared. The above description is an automatic optimization method, which is more general. You can also manually select map samples for optimization while reading the report. At this time, the report is also a powerful tool for selecting optimal samples.
[0111] Compared with manual report writing, the present invention is a fully automatic report writing system that does not require manual intervention and labor costs, and is much more efficient than manual work. In addition, the present invention has designed a reasonable information content arrangement, making the report format standardized, comprehensive and easy to read.
[0112] Compared with automatic formatting, the present invention clusters problem samples by statistically analyzing the appearance, characteristics, and quantity of all problem samples, intelligently calculating the number of samples to be pasted through an algorithm, and further screening out representative samples so that the pasted samples can best fit the distribution of all problem samples. This ensures the quality and efficiency of report generation, while combining the advantages of both full content reports and partial content reports while overcoming their respective shortcomings.
[0113] Furthermore, the invention designs a more reasonable and universal format, which solves the problem that reports automatically generated according to a fixed format may appear stiff and difficult to read;
[0114] The invention also builds an optimization mechanism based on the strategy of test result map sorting and screening, which can automatically iterate and optimize based on the generated report. By dynamically adjusting the training set and test set, it can realize automatic sample optimization, improve the model effect, and optimize the efficiency of single training.
[0115] The above description only describes specific embodiments of the present invention. Various examples do not limit the essential content of the present invention. After reading the description, ordinary technicians in the relevant technical field can make modifications or variations to the specific embodiments described above without departing from the essence and scope of the invention.
Claims
1. A method for generating training and test reports for a visual inspection AI platform, characterized by: The following steps are included: S1. After the test, statistics and calculations are performed on the obtained system parameters and results; S2, screening training samples; filter Including feature structure screening and feature cluster screening; S3. Perform sample screening on the test results; the screening includes the calculation of the number of stickers and sticker screening.
2. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: In step S1, the test results include the output label and true label of each test sample, the output label and true label of each training sample, and the image file path of each training sample and test sample.
3. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: In step S1, the system parameters include engineering parameters, structural parameters of the model itself, and hyperparameters during training and testing.
4. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: In step S1, the calculation includes obtaining the detection results and calculating the intersection between all detections and the true annotations on the corresponding samples.
5. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: In step S2, the method of feature structure screening is: 1) Calculate the shape hu moment of the annotation, which has 7 orders in total, forming a 7-dimensional feature; 2) Calculate the grayscale / color histogram of the annotated part of the image, quantize the histogram into 8 levels, and form an 8-dimensional / 24-dimensional feature. For unannotated samples, that is, positive samples, fill the 8-dimensional / 24-dimensional feature with 0; 3) Calculate the grayscale / color histogram of the image outside the annotation area, quantize the histogram into 8 levels, and form 8-dimensional / 24-dimensional features; 4) Perform PCA dimensionality reduction on the features of the layer before the network output layer to 8 dimensions, forming 8-dimensional features; 5) The features obtained above are concatenated to form features for training sample screening.
6. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: In step S2, the feature clustering screening method is: a) Use the K-means clustering method to cluster the positive samples into K categories based on the features of the constructed positive samples, where K is the number of positive sample maps of the current class; b) After clustering the positive samples, find the center of each subclass obtained by clustering the positive samples. In each subclass, select the sample closest to the center as the positive sample for display, that is, obtain K positive samples for display; c) For each category, select samples for display according to the above method and paste them in the corresponding positions according to the set requirements.
7. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: In step S3, the map screening strategy is as follows: detect each type of map in the map, construct features according to the feature construction method, calculate the center coordinates, and divide the sample features into 5 intervals according to the distance from the center point, using the mean square error formula. Each interval is allocated x candidate places, x = T / 5, T is the number threshold of the detection of this type. If an interval has less than x candidates, then the interval is used as the full selection interval, and the remaining places allocated by the interval are evenly distributed to other intervals with less allocated places than its sample number, thereby completing the allocation of 50 places. When using samples for network inference, the features of the layer before the output layer are used as the network features of the samples. Taking missed detection as an example, the network feature distance between the samples in the non-selected interval and each sample in the positive samples in the training sample set and each sample in the interval is calculated. The minimum distance between each sample in the interval and all positive samples is taken as the inter-class distance of the sample. All samples in the interval are sorted from small to large according to the inter-class distance, and then samples are selected according to the allocated quota order of the interval as mapping samples.
8. The method for generating a training and test report for a visual inspection AI platform according to claim 1, wherein: Also includes, Step S4, optimization strategy: sort each detection situation according to the inter-class distance, use the sorted inter-class distance as the dependent variable to construct a distance function, select the point with the largest second-order derivative of the function as the inflection point, select the sample to the left of the inflection point as the candidate sample, cluster the candidate samples based on network features, and determine the K value according to k-means combined with the elbow rule. In each class, the sample with the closest distance to the center point within the class is used as the optimal sample. Select the optimal sample for each detection situation according to the above method, add the optimal sample to the training set for retraining, and generate the report again.