A Method for Evaluating the Accuracy of a Classifier Based on the Silhouette Coefficient

Through the method of clustering profile coefficients and polynomial regression fitting, the accuracy of the classifier on the label-free data set is predicted, which solves the problems of unreliable and high overhead of evaluation methods in the prior art, and achieves efficient and accurate prediction of accuracy.

CN116628574BActive Publication Date: 2025-06-20HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310654826.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2025-06-20
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

The existing label-free classifier evaluation methods have unreliability and high overhead problems, especially the confidence-based methods rely on threshold selection and are unreliable. The methods based on the previous task require retraining the classifier, which increases overhead and may reduce the real performance on the test set.

Method used

Using a method based on clustering contour coefficients, the features of the synthetic data set are clustered, the contour coefficients are calculated, and the relationship between the contour coefficients and accuracy is fitted using polynomial regression, thereby predicting the accuracy of the classifier on the labelless data set.

Benefits of technology

This method does not require retraining the classifier, improves efficiency, and significantly improves prediction accuracy through strong linear relationships, avoiding the unreliability and high overhead problems of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116628574B_ABST
    Figure CN116628574B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and provides a method for evaluating the accuracy of a classifier based on the silhouette coefficient, which is used to evaluate the accuracy of a classification model to be tested for classifying data to be tested, including: preparing a dataset in picture format; extracting the features of each data through the classifier of the classification model to be tested; calculating the accuracy of the classifier for classifying the dataset; clustering the features of each data in the dataset to obtain the silhouette coefficient; fitting the relationship between the silhouette coefficient and the accuracy through polynomial regression; using the evaluation of the classification model to be tested to extract the features of the data to be tested, calculating the silhouette coefficient, and calculating the accuracy of the classification model to be tested for classifying the data to be tested based on the polynomial and the silhouette coefficient. Compared with the existing method, the present application can improve the efficiency and prediction accuracy without retraining the classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly relates to a method for evaluating the accuracy of a classifier based on the silhouette coefficient of clustering. Background Art

[0002] With the development of deep learning and artificial intelligence, various classifiers are applied in production and life. Evaluating the performance of classifiers, that is, predicting the accuracy of classifiers in real environments, becomes increasingly important. Since there are differences in the performance of classifiers between the training set and the real environment, in order to evaluate the performance of classifiers in the real environment, we need to sample a test set in the real environment. Since the sampled test set has no labels and it is extremely expensive to manually add annotations to the test set. How to predict the accuracy of classifiers without test set labels has been a research hotspot in recent years. The task of predicting the accuracy of classifiers without labels predicts the accuracy of the classifier on the unlabeled test set by learning the feature differences of different data sets on the classifier. This task synthesizes a labeled data set by performing image transformation on the labeled data set to more efficiently predict the model accuracy without relying on test set labels. Implementing a method for accurately predicting the accuracy of classifiers is of great significance for evaluating the accuracy of classifiers in real environments and understanding the generalization ability of classifiers.

[0003] The existing methods for evaluating classifiers without labels can be divided into two strategies: confidence-based prediction methods and previous-task-based methods. The confidence-based prediction method directly compares a preset threshold with the maximum probability value output by the softmax activation function of the classifier. If the maximum probability value output by the softmax activation function of the classifier is greater than the threshold, it is considered that the classifier classifies the sample correctly. The previous-task-based method proves that there is a linear relationship between the classifier accuracy and the previous-task accuracy. This method adds a previous-task module after the convolutional layer of the classifier and predicts the classifier accuracy according to the accuracy of the previous task. For example, some recent research works use rotation prediction or out-of-distribution detection score to predict the classifier accuracy. Generally speaking, the existing methods mainly perform image transformation on the labeled data set to generate a large number of synthetic data sets and learn the feature differences of the synthetic data sets on the classifier to predict the accuracy of the classifier on the unlabeled data set.

[0004] However, in the existing methods, the disadvantage of the confidence-based method is that the effect of this method depends on the selection of the threshold. Since the unlabeled test set is often invisible, the threshold we selected may not be effective in different environments. This method has been proven to be unreliable. The disadvantage of the previous-task-based method is that this method needs to retrain the classifier after adding the previous-task module. Fine-tuning the classifier will not only increase the cost but also change the parameters of the classifier. The change of the classifier parameters may cause its true performance on the test set to decline. Summary of the Invention

[0005] To solve the above problems, the present invention provides a method for evaluating the accuracy of a classifier based on the silhouette coefficient.

[0006] This method is used to evaluate the accuracy of a classification model to be tested for classifying data to be tested, and is characterized by including the following steps:

[0007] Step 1: Prepare a basic data set, where each sample in the basic data set is a labeled picture;

[0008] Step 2: Based on the basic data set, generate a synthetic data set, where each sample in the synthetic data set has a label;

[0009] Step 3: Extract the features of each data in the synthetic data set through the classifier of the classification model to be tested;

[0010] Step 4: Calculate the accuracy of the classifier of the classification model to be tested for classifying the synthetic data set;

[0011] Step 5: Cluster the features of each data in the synthetic data set to obtain the silhouette coefficient SC1;

[0012] Step 6: Use polynomial regression to fit the relationship between the silhouette coefficient and the accuracy;

[0013] Step 7: Use the classification model to be tested to extract features from the data to be tested, calculate the silhouette coefficient, and calculate the accuracy of the classification model to be tested for classifying the data to be tested based on the polynomial and the silhouette coefficient SC2.

[0014] Further, in step 2, generating the synthetic data set based on the basic data set specifically means performing various image transformations on the samples in the basic data set to obtain new pictures, and forming the synthetic data set with the new pictures.

[0015] Preferably, the image transformations include: rotation, cropping, contrast adjustment, color channel inversion, hue adjustment, and noise addition.

[0016] Further, step 4 specifically includes: inputting the pictures in the synthetic data set into the classifier of the classification model to be tested to obtain predicted labels, and comparing the labels of the pictures in the synthetic data set with the predicted labels to obtain the accuracy of the classifier of the classification model to be tested for classifying the synthetic data set.

[0017] Further, the polynomial in step 6 is a cubic polynomial of one variable.

[0018] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0019] This application proves through experiments that there is a strong linear relationship between the silhouette coefficient of feature clustering of a dataset on a classifier and the accuracy rate of the classifier on the dataset, and fits the relationship between the two through polynomial regression to predict the accuracy rate of the classifier on an unlabeled dataset. At the same time, since the number of label types in the unlabeled test set is the same as that in the synthetic dataset, the number of clustering centers of features does not need to be calculated additionally using other algorithms. Compared with the existing methods, this application can improve the efficiency and prediction accuracy without retraining the classifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a method for evaluating the accuracy rate of a classifier based on the silhouette coefficient provided by an embodiment of the present invention;

[0021] Figure 2 is a linear relationship diagram between the accuracy rate and the silhouette coefficient of the classifier provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Before elaborating on the technical solutions of the various embodiments of the present invention, the nouns and terms involved are explained. In this specification, components with the same name or the same reference numeral represent similar or identical structures, and are for illustrative purposes only.

[0023] This application is funded by the Collaborative Innovation Project of Anhui Provincial Universities, and the project number of the funded project is GXXT-2022-043 of the Collaborative Innovation Project of Anhui Provincial Universities.

[0024] This application proposes a method for predicting the accuracy rate of a classifier based on the silhouette coefficient. Since it is impossible to directly evaluate the accuracy rate of the classifier on an unlabeled test set, this method uses the silhouette coefficient of feature clustering of the dataset on the classifier as the feature difference of the dataset and learns the feature difference of the synthetic dataset through polynomial regression to predict the accuracy rate of the classifier on the unlabeled test set. The overall flowchart of this method is as Figure 1 shown: Given a classifier, a basic dataset, and an unlabeled test set, first generate a set of synthetic datasets through image transformation in the basic dataset, and obtain the features of the synthetic dataset and the accuracy rate of the classifier on the synthetic dataset. Then, obtain the silhouette coefficient by clustering the dataset features. Learn the relationship between the two through polynomial regression. Use this polynomial regression to predict the accuracy rate of the unlabeled test set, that is, extract the features of the unlabeled test set and calculate the clustering silhouette coefficient, and input it into the polynomial regression to predict the accuracy rate of the classifier on the unlabeled test set.

[0025] 1. Generation of the synthetic dataset

[0026] It is unrealistic to use a large number of real datasets to study the relationship between the silhouette coefficient of dataset feature clustering and the accuracy of the classifier on the dataset. Therefore, a basic dataset is prepared first.

[0027]

[0028] Among them is the picture in the basic dataset, is the picture corresponding label, N represents the number of samples in the basic dataset, 1 ≤ n ≤ N. For the basic dataset D train Apply three of the six methods of rotation, cropping, adjusting contrast, inverting color channels, adjusting hue, and adding noise to each picture in it to generate 600 synthetic datasets D j , Among them represents the nth picture in the synthetic dataset D j , represents the corresponding label. Since the images in the basic dataset all have labels and the image transformation does not affect the labels of the pictures in the basic dataset, the labels of the synthetic dataset are also known.

[0029] 2. Feature extraction and the accuracy of the classifier on the synthetic dataset

[0030] Use the set of features of all pictures in the jth synthetic dataset on the classifier f to represent the corresponding feature h j of the synthetic dataset D j :

[0031]

[0032] Among them is the output of the penultimate layer of the classifier for the nth picture in the synthetic dataset D j . Since the labels of the synthetic dataset are known, the accuracy ACC j of the classifier on the synthetic dataset can be directly obtained by comparing the picture label with the predicted label . Among them, the predicted label is the output obtained by inputting the picture into the classifier f.

[0033] 3. Feature clustering and silhouette coefficient

[0034] Use the K-MEANS algorithm to cluster the image features h jPerform clustering: First, arbitrarily select K image features as the initial clustering centers; for the remaining other image features, assign them to the class that is most similar to them according to their distances from these clustering centers; then calculate the clustering center of each newly obtained cluster (the mean of all objects in the cluster); continuously repeat this process until the standard measure function starts to converge. Since the number of label types in the unlabeled test set is the same as that in the synthetic data set, the number of preset clustering centers K in K-MEANS can be directly set to the number of label types in the synthetic data set.

[0035] For the silhouette coefficient S(v) of a point v in a class, it can be expressed as:

[0036]

[0037] where a(v) is the distance from point v to all other points in the class it belongs to, and b(v) is the minimum value among the average distances from point v to other points in a class that does not contain it. The overall silhouette coefficient SC of the clustering is expressed as:

[0038]

[0039] 4. Polynomial regression

[0040] For a classifier, the more correct the image features extracted by the classifier are, the higher the classification accuracy. At the same time, good image features are beneficial to the clustering effect, that is, the more correct the classifier features are, the larger the clustering silhouette coefficient. In addition, this application also proves through a large number of experimental results that there is a connection between the silhouette coefficient of feature clustering of the data set on the classifier and the accuracy of the classifier on the data set, as Figure 2 shown: The linear relationship between the accuracy of the classifier on the data set and the silhouette coefficient after feature clustering of the data set. The vertical axis is the accuracy of the classifier on the data set, and the horizontal axis is the silhouette coefficient. The Pearson coefficient between the two is 0.92, showing a strong linear relationship.

[0041] This application uses polynomial regression to fit the relationship between the silhouette coefficient of feature clustering on the classifier and the accuracy of the classifier on the data set. The input of the polynomial regression is the silhouette coefficient of feature clustering of the data set on the classifier, and the output is the accuracy of the classifier on the data set. Since the linear relationship between the input and output is strong, this application uses a cubic polynomial of one variable to fit the relationship between the two:

[0042]

[0043] where ACC j is the accuracy of the classifier on the synthetic data set, j the silhouette coefficient SC j of the clustering of the corresponding feature hj is the input, and A, B, C, and D are all parameters to be learned for polynomial regression.

[0044] 5. Accuracy of the classifier on the unlabeled test set

[0045] After the polynomial regression training is completed, the silhouette coefficient of the features of the unlabeled test set can be input into the polynomial to predict the accuracy of the classifier on the unlabeled test set. First, the unlabeled test set is input into the classifier f to extract the features of each image on the classifier, then it is clustered and the silhouette coefficient is calculated. Finally, the silhouette coefficient of the clustering is input into the polynomial, and the resulting output is the prediction of the accuracy of the classifier on the unlabeled test set.

[0046] 6. Experiments

[0047] The classifier object evaluated in this application is ResNet-50, and CIFAR-10 is used as the basic dataset. The CIFAR-10 dataset contains 60,000 32x32 color images, divided into 10 categories, with 6,000 images in each category. There are 50,000 training images and 10,000 test images. The unlabeled test sets for evaluation are the CIFAR-10.1 dataset and 30 CIFAR-C datasets. CIFAR-10.1 is a subset of the TinyImages dataset and contains 2,000 new test images. The CIFAR-C dataset is 30 synthetic datasets generated by applying image transformations to the CIFAR-10 test set.

[0048] The experimental evaluation metric is the root mean square error, which represents the square root of the ratio of the sum of the squares of the deviations between the predicted values and the true values to the number of observations. Therefore, the root mean square error can intuitively represent the error between the predicted accuracy and the true accuracy of this method. The smaller the root mean square error, the better the method.

[0049] Table 1 Experimental results

[0050]

[0051]

[0052] The experimental results are shown in Table 1. The results show that the method based on the silhouette coefficient proposed in this application can better predict the accuracy of the classifier on the unlabeled dataset than the method based on confidence and the method based on the previous task. Since there is no need to retrain the classifier, the method proposed in this application is much more efficient than the method for evaluating the unlabeled classifier based on the previous task.

[0053] The following combines a specific embodiment to explain the present invention.

[0054] There may be a difference in the accuracy of the license plate recognition system between the training environment (bright, unobstructed) and the actual deployment environment (interferences such as rainy days and heavy fog). This application can predict the accuracy of the license plate recognition system before its deployment, providing a basis for the deployment, adjustment, and optimization of the license plate recognition system.

[0055] For example, it is necessary to predict the accuracy of a VGG-16 classifier trained on the HyperLRP training set in a real unlabeled environment. First, use image transformation on the HyperLRP training set to generate 600 synthetic data sets. Then, extract the features of the images in each synthetic data set on the VGG-16 classifier and calculate the silhouette coefficient for the clustering of the image features in each data set. After that, use polynomial regression to learn the relationship between the clustering silhouette coefficient and the accuracy of VGG-16 on the synthetic data set. Finally, extract samples in the real environment as the unlabeled test set, extract the image features of the unlabeled test set, and input the clustering silhouette coefficient into the polynomial regression to predict the accuracy of the VGG-16 classifier in the real unlabeled environment.

[0056] This application proposes a method for evaluating the accuracy of a classifier based on the clustering silhouette coefficient. For a given basic data set and classifier, this method generates a set of synthetic data sets through image transformation, and uses the clustering silhouette coefficient of the image features in the synthetic data set as the feature difference between data sets, and predicts the accuracy of the classifier on the unlabeled data set by learning a polynomial regression. A large number of experiments show that the method proposed in this application is significantly better than the current methods. In addition, the classifier evaluation method proposed in this application can also be widely applied to image classification to efficiently evaluate the performance of the classifier.

[0057] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for evaluating the accuracy of a classifier based on the silhouette coefficient, which is used to evaluate the accuracy of a classification model to be tested for classifying data to be tested, characterized in that, It includes the following steps: Step 1: Prepare a basic dataset, where each sample in the basic dataset is a labeled picture; Step 2: Generate a synthetic dataset based on the basic dataset, and each sample in the synthetic dataset has a label; Step 3: Extract the features of each data in the synthetic dataset through the classifier of the classification model to be tested; Step 4: Calculate the accuracy of the classifier of the classification model to be tested in classifying the synthetic dataset; Step 5: Cluster the features of each data in the synthetic dataset to obtain the silhouette coefficient , and the specific steps are as follows: Using the K-MEANS algorithm, the image features of the dataset are clustered: First, arbitrarily select K image features as the initial cluster centers; for the remaining other image features, according to their distances from these cluster centers, they are respectively assigned to the class that is most similar to them; then calculate the cluster center of each newly obtained cluster, that is, the mean of all objects in this cluster; continuously repeat this process until the standard measure function starts to converge; since the number of label types in the unlabeled test set is the same as the number of label types in the synthetic dataset, the number of preset cluster centers K in K-MEANS is directly set to the number of label types in the synthetic dataset; For the points in the class The silhouette coefficient Is expressed as: ; wherein, is the distance from a point to other points in all classes it belongs to, and is the minimum value among the average distances from a point to other points in a certain class that does not contain it; the overall silhouette coefficient of clustering is expressed as: ; Step 6: Use polynomial regression to fit the relationship between the silhouette coefficient and the accuracy. The specific steps are as follows: Use a cubic polynomial of one variable to fit the relationship between the two; ; Among them, the accuracy rate of the classifier on the synthetic dataset is the output, and the synthetic dataset corresponding features silhouette coefficient of the clustering is the input, , , , are all parameters to be learned for polynomial regression; Step 7: Use the classification model to be tested to extract features from the data to be tested, calculate the silhouette coefficient, and calculate the accuracy of the classification model to be tested for classifying the data to be tested based on the polynomial and the silhouette coefficient , and calculate the accuracy of the classification model to be tested for classifying the data to be tested.

2. The method for evaluating the accuracy of a classifier based on the silhouette coefficient according to claim 1, characterized in that, In Step 2, generating the synthetic dataset based on the basic dataset specifically means performing various image transformations on the samples in the basic dataset to obtain new pictures, and using the new pictures to form the synthetic dataset.

3. The method for evaluating the accuracy of a classifier based on the silhouette coefficient according to claim 2, characterized in that, The image transformations include: rotation, cropping, contrast adjustment, color channel inversion, hue adjustment, and noise addition.

4. The method for evaluating the accuracy of a classifier based on the silhouette coefficient according to claim 1, characterized in that, Step 4 specifically includes: inputting the pictures in the synthetic dataset into the classifier of the classification model to be tested to obtain predicted labels, and comparing the labels of the pictures in the synthetic dataset with the predicted labels to obtain the accuracy of the classifier of the classification model to be tested in classifying the synthetic dataset.

5. The method for evaluating the accuracy of a classifier based on the silhouette coefficient according to claim 1, characterized in that, The polynomial in Step 6 is a cubic polynomial of one variable.

Citation Information

Patent Citations

  • Finger tip point extraction method based on pixel classifier and ellipse fitting

    CN105046199A

  • New intention discovery method and device based on clustering, equipment and storage medium

    CN114510567A