Neural Network Specific Misclassification Repair Method Based on Perturbation Feature Recovery
Through the neural network-specific misclassification repair method based on perturbation feature recovery, the problem of accuracy reduction caused by insufficient distinction capabilities of specific categories and global parameter adjustment is solved, and efficient misclassification repair and accuracy maintenance are achieved.
Patent Information
- Application Number
- CN202510146970.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The problem of insufficient discrimination ability among specific categories and significant decrease in accuracy caused by global adjustment of unconstrained parameters in the prior art.
Through a neural network-specific misclassification repair method based on perturbation feature recovery, the perturbation feature is repaired using the feature mapping model, and the repaired features are verified by the validator, and repaired only on the suspicious subset filtered in the original data set, narrowing the scope of decision boundary adjustment.
A better repair effect was achieved, the accuracy level was maintained without a significant decrease, the decision-making boundaries between other categories were protected, and the accuracy of the original model was protected to the greatest extent.
Smart Images

Figure CN119625467B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of misclassification repair, and in particular to a method for repairing specific misclassifications of a neural network based on restoring perturbed features. Background Art
[0002] With the development of deep neural networks, image classification has become the core of many practical applications and is of great significance in the field of deep learning. However, misclassifications in classification models may lead to incorrect decisions, and the potential hazards brought about by the large-scale application of neural networks have increasingly attracted people's attention. A misclassification of a neural network model refers to a situation where the predicted class of a given input image does not match the actual class of the image.
[0003] In misclassifications, a specific misclassification refers to a situation where the model misclassifies a sample belonging to a certain specific class as another specific class. Specific misclassification instances / samples are all clean samples on the original dataset. In a medical diagnosis task, misdiagnosing a "pre-cancer" patient as "healthy" may cause the patient to miss the best treatment opportunity, leading to the deterioration of the disease. Nevertheless, the harms caused by different types of specific misclassifications can vary greatly. Some specific misclassifications can cause serious consequences, while others are relatively minor.
[0004] In recent years, researchers have developed a series of repair techniques for DNN (Deep Neural Networks) models, but there are still two main limitations: insufficient ability to distinguish specific categories and a significant decrease in accuracy caused by unconstrained global adjustment of parameters. For example, mainly targeting misclassifications caused by clean or adversarial inputs, for misclassifications caused by adversarial samples, a pre-trained model trained on the original training set through adversarial training and a training on an additionally generated adversarial sample dataset are used to enhance the robustness of the model. However, adversarial training heavily relies on additional synthetic samples, which requires a large amount of time and computing resources. And it is difficult to improve the generalization ability of the model on the clean dataset through retraining based on sample synthesis for samples on the clean dataset (i.e., the original dataset). Therefore, in the absence of additional data, the general misclassifications caused by clean inputs are effectively solved by learning the overall weights of the reference model. This unconstrained global parameter adjustment has a significant impact on the decision boundary of the model but is prone to introducing new misclassifications. By locating the neurons that cause errors and performing neuron-level fine-tuning to achieve repair, the adjustment range of the model parameters is narrowed, and the number of misclassifications introduced during the repair process can be effectively controlled. However, this method still affects the original classification performance (i.e., accuracy) of the model during the repair process, resulting in a significant decrease in accuracy. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present invention aim to provide a method for repairing specific misclassifications of a deep neural network based on perturbed feature recovery, so as to solve the problems of insufficient discrimination ability between specific categories and the decline in the accuracy of misclassification repair with unconstrained global adjustment of parameters in the prior art.
[0006] The object of the present invention is mainly achieved through the following technical solutions:
[0007] On the one hand, the embodiments of the present invention provide a method for repairing specific misclassifications of a neural network based on perturbed feature recovery, including the following steps:
[0008] Use the neural network model trained by the original image dataset to perform initial classification on the dataset of images to be classified. The neural network model includes an initial feature extractor for extracting the features of the images to be classified and a classifier for classifying using the features.
[0009] Screen out the image sample data classified as category P in the dataset of images to be classified as the suspected set.
[0010] Based on the suspected set and the trained feature mapping model, perform perturbed feature repair on the latent features extracted by the initial feature extractor to obtain repaired features.
[0011] Use the trained validator to verify whether the repaired features correspond to specific misclassifications. If so, repair the image sample data misclassified as category P in the suspected set to obtain the repaired classification result. Among them, specific misclassification refers to misclassifying the image sample belonging to category T into category P.
[0012] Further, performing the perturbed feature repair includes:
[0013] Use the feature mapping model to repair the perturbed features in the latent features extracted by the initial feature extractor from the suspected set into the features extracted by the dedicated feature extractor. Among them, the dedicated feature extractor is trained using the image samples of specific misclassification-related categories.
[0014] Update the repaired perturbed features into the latent features of the suspected set to obtain the repaired features.
[0015] Further, using the trained validator to verify whether the repaired features correspond to specific misclassifications includes:
[0016] Input the repaired features into the trained validator.
[0017] According to the fact that the image sample data misclassified as category P in the suspected set has the features of category T, reclassify the repaired features.
[0018] If the repaired feature corresponds to the category T of the specific misclassification, the label is the positive class and repair is performed; otherwise, the label is the negative class and the result of the initial classification is retained.
[0019] Further, the repaired classification result is to repair the image sample data labeled as the positive class by the trained validator to the category T.
[0020] Further, training the validator includes:
[0021] Relabel the original image data set to obtain two types of image samples; among them, the two types of image samples include labeling the image sample data of the category T as the positive class and labeling the sample data classified as other categories in the original image data set as the negative class;
[0022] Using the two types of image samples as the input data set, training the validator based on a binary classifier to obtain a classification result including the positive class and the negative class.
[0023] Further, the binary classifier is a classifier structure based on the neural network.
[0024] Further, training to obtain the dedicated feature extractor includes:
[0025] In the original image data set, extract the image samples of the categories related to the specific misclassification as the retraining data set; among them, the categories related to the specific misclassification include the category T and the category P;
[0026] Freeze the classifier of the neural network and retrain the initial feature extractor based on the retraining data set to obtain the dedicated feature extractor.
[0027] Further, training the feature mapping model includes:
[0028] Based on the initial feature extractor, extract the latent features of the retraining data set, and use the perturbation features in the latent features as the input sample set;
[0029] Based on the dedicated feature extractor, extract the dedicated features of the retraining data set, and use the features corresponding to the perturbation dimensions in the dedicated features as the output sample set; where the perturbation dimension is the dimension corresponding to the perturbation feature.
[0030] Further, train the feature mapping model based on a multi-layer perceptron and a mean square error loss function.
[0031] Further, the neural network includes a convolutional neural network, and the initial feature extractor includes a convolutional layer and a pooling layer.
[0032] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0033] 1. The present invention proposes to identify disturbance features based on a specific misclassified suspicion set, restore the disturbance features through a feature mapping model, effectively eliminate redundant information, and repair the misclassified features in combination with a verifier, which not only enables the original model to achieve a better repair effect, but also keeps the accuracy level from decreasing significantly;
[0034] 2. By performing repair only on the suspected subset (i.e., the suspected set) screened in the original image dataset, the scope of decision boundary adjustment is narrowed, thereby protecting the decision boundaries trained between other categories from interference, and protecting the accuracy of the original model from being severely reduced due to decision boundary adjustment to the greatest extent;
[0035] 3. The training of the feature mapping model is carried out with the help of a dedicated feature extractor. The initial feature extractor is retrained on a subset containing only specific misclassified related categories to obtain a dedicated feature extractor, which is used to improve the model's ability to distinguish between category T and category P;
[0036] 4. The model classification performance is enhanced through pluggable feature recovery mapping and validators without directly adjusting the model weights, avoiding additional computational and storage overhead.
[0037] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the entire drawings, the same reference symbols represent the same components;
[0039] Figure 1 It is a flow chart of a method for repairing specific misclassification of a deep neural network based on perturbation feature recovery according to an embodiment of the present invention;
[0040] Figure 2 It is a technical solution framework diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0042] A specific embodiment of the present invention discloses a method for repairing specific misclassifications of a neural network based on perturbation feature recovery, as follows Figure 1 shown, including the following steps:
[0043] Step S1: Use the neural network model trained with the original image dataset to perform initial classification on the dataset to be classified. The neural network model includes an initial feature extractor for extracting the features of the image to be classified and a classifier for classifying using the features.
[0044] Step S2: Screen out the image sample data classified as category P in the dataset to be classified as the suspected set.
[0045] Step S3: Based on the suspected set and the trained feature mapping model, perform perturbation feature repair on the latent features extracted by the initial feature extractor to obtain the repaired features.
[0046] Step S4: Use the trained validator to verify whether the repaired features correspond to specific misclassifications. If so, repair the image sample data misclassified as category P in the suspected set to obtain the repaired classification result. Among them, specific misclassification refers to misclassifying the image sample belonging to category T into category P.
[0047] Through the above method, by screening the suspected set in the original dataset for repair, adjusting the local decision boundary for relevant categories, using the feature mapping model to repair the latent features extracted by the initial feature extractor, enhancing the discrimination ability of specific categories, using the validator to correct the potential specific misclassifications in the suspected set, realizing the correction of specific misclassifications in the multi-classification model, and at the same time reducing the accuracy loss caused by the repair process to ensure the accuracy of the repair.
[0048] Specifically, in step S1, the neural network can be a convolutional neural network. The initial feature extractor is responsible for extracting the machine-understandable representation (i.e., latent features) from the original dataset, usually including convolutional layers and pooling layers, etc. The classifier is responsible for distinguishing the extracted features and outputting the finally predicted category, usually composed of fully connected layers. In fact, the output of the feature extractor can be regarded as the input of the classifier. The occurrence of specific misclassifications can be attributed to the failure of the machine learning model to accurately distinguish its true category and its mispredicted category, called the relevant category. Represent the specific misclassification as , where category T and category P represent the true label and the mispredicted label of the specific misclassification respectively, that is, misclassifying the image sample belonging to category T into category P. Obviously, for a given specific misclassification, its occurrence is only related to these two categories.
[0049] Most existing methods only fix specific misclassifications by directly adjusting the parameters of the pre-trained model, without enhancing the model's recognition ability for the corresponding classes. In order to ensure the overall classification accuracy of the multi-class neural network model, its feature extractor needs to extract general features applicable to distinguish all classes, which can be regarded as a redundant feature extractor (RFE), including redundant features that describe the differences between other classes, and cannot extract all the differences required to effectively separate two specific classes, thus restricting the distinction between specific classes; on the other hand, the unconstrained global adjustment of the pre-trained model parameters will destroy the precise classification boundaries that have been trained between other classes before. This interference often causes samples that were originally correctly classified (i.e., samples with correct predictions but low prediction confidence) to be misclassified near the decision boundary, thus introducing new misclassifications.
[0050] Based on the above problems, the present invention proposes an efficient method for repairing specific misclassifications in a neural network graph (Specific Misclassification Immune Repair, SMiR), as Figure 2 shown. Among them, the dashed arrow indicates the framework process of the repair model construction stage, and the solid arrow indicates the framework process of the forward inference stage. SMiR includes three main mechanisms: perturbation feature localization and recovery, suspect set screening, and validator construction.
[0051] Specifically, in step S2, to solve the problem of global adjustment of model parameters, SMiR restricts the adjustment of the decision boundary, divides the original data set according to the original prediction output of the model, and adjusts the object of the repair operation in a local subset containing the target specific misclassification.
[0052] Exemplarily, for a specific misclassification situation , samples predicted as P by the original model are identified as potential misclassified samples to be repaired. The image instance samples among them contain parts that truly belong to the label P and parts that are misclassified as P. Therefore, the set composed of these samples is called the suspect set, denoted by . The remaining samples are regarded as the trusted set, denoted by .
[0053] Assume that the dataset D to be classified is class-balanced, and the target misclassification to be repaired is , then most of the samples in the screened suspect set have the true class of . It can be approximately considered that the suspect set only contains approximately in the dataset Dtotal samples (where C is the total number of categories in the dataset). In other words, through the screening operation on the suspicious set, the possibility of introducing misclassifications in the remaining samples is eliminated (for example, this ratio is 90% for the CIFAR-10 dataset). In addition, the number of samples that need to undergo feature recovery is significantly reduced, significantly reducing the computational cost.
[0054] It should be noted that a specific misclassification is caused by the model's decision boundary wrongly favoring the incorrect class P over the correct class T. Repairing it is essentially a compensatory adjustment to the decision boundary. However, unconstrained adjustment of the entire decision boundary will inevitably damage the originally good classification boundaries between other classes. Since the trusted set does not contain specific misclassification instance samples, no additional operations are performed on it. By only applying the repair operation to the suspicious set. The decline in the classification performance of the model for other classes is avoided.
[0055] Specifically, in step S3, the located perturbed features are restored through the trained feature mapping model to reduce the adverse effects of redundant information therein. Repairing the perturbed features of the extracted latent features includes:
[0056] S311. Using the feature mapping model to repair the perturbed features in the latent features of the suspicious set extracted by the initial feature extractor into features extracted by a dedicated feature extractor (DFE); among them, the features severely affected by redundant information are regarded as perturbed features.
[0057] S312. Replace the features corresponding to the dimensions in the latent features of the suspicious set with the repaired perturbed features, and keep the features of other dimensions unchanged to obtain the updated latent features, which are the repaired features.
[0058] Among them, training the dedicated feature extractor using the image samples of the relevant classes of specific misclassifications includes:
[0059] S321. In the original image dataset of the neural network model, extract the image sample data of the relevant classes of specific misclassifications as the retraining dataset ( that is, a subset of the image sample data in the original image dataset that only contains the image sample data of class T and class P);
[0060] S322. Freeze the classifier component in the neural network model, copy the initial feature extractor in the neural network model, and retrain the RFE component based on the retraining dataset to obtain the DFE.
[0061] Further, training the feature mapping model includes:
[0062] Extracting the latent features of the retraining dataset based on the initial feature extractor, and using the perturbed features in the latent features as the input sample set;
[0063] Extracting the dedicated features of the retraining dataset based on the dedicated feature extractor, and using the features corresponding to the perturbed dimensions in the dedicated features as the output sample set; wherein, the perturbed dimension is the dimension corresponding to the perturbed feature.
[0064] Exemplarily, a lightweight feature mapping model is trained based on a multi-layer perceptron and a mean square error loss function. In the forward inference stage, the trained lightweight feature mapping model is used to separately recover the localization dimension of the perturbed features.
[0065] It should be noted that a dedicated feature extractor is used to specifically extract the features of relevant categories in a given misclassification, so as to solve the problem that the features extracted by the feature extractor contain redundant features, which restricts the distinction between two specific categories. The features extracted by DFE are not directly applied to the model, but are combined with the features extracted by RFE to identify and recover those severely perturbed feature dimensions. DFE is only used for the localization and recovery of perturbed features in the repair process. Therefore, additional computational and storage overheads are avoided in the forward inference process.
[0066] Further, a model interpretability technique (Deep Learning Important FeaTures, DeepLift) is adopted to quantify the contribution values of the features of each dimension to the correct classification of specific category samples; compare the contribution values of the features of each dimension extracted by the two feature extractors for categories T and P before and after noise reduction; according to the characteristic that the higher the increase in the contribution value of the features of each dimension extracted by DFE compared to RFE, the more the feature dimension can contribute to the correct identification of the corresponding category after denoising, locate the features with a large increase in contribution after denoising as perturbed features; wherein, DFE can be regarded as the result of RFE after noise reduction for categories T and P.
[0067] Specifically, in step S4, in order to correct a specific misclassification, using a validator to verify whether the recovered sample features correspond to the specific misclassification to be repaired includes:
[0068] S41. Input the repaired features into the trained validator;
[0069] S42. Reclassify the repaired features according to whether the image sample data misclassified as category P in the suspicion set has the features of category T.
[0070] S43. If the feature after repair corresponds to the specific misclassified class T, the label is a positive class and repair is performed; otherwise, the label is a negative class and the result of the initial classification is retained.
[0071] It should be noted that since all the suspicious samples of the input validator are predicted as class P, the validator only needs to check whether the true label of the input sample is T, that is, whether it is the specific misclassification to be repaired; if so, repair is performed. The repair action is to repair the image sample data labeled as a positive class by the validator to class T, and the obtained classification result is the classification result after repair.
[0072] Furthermore, training the validator includes:
[0073] Relabel the original image data set, relabel the samples with the original label of T as positive classes (labeled as 1), and relabel the non-T samples as negative classes (labeled as 0) to obtain two types of image samples;
[0074] Use the two types of image samples as the input data set and train the validator based on a binary classifier , and obtain the classification results corresponding to the two classes; among them, determine the architecture of the binary classifier based on the classifier structure of the neural network.
[0075] Exemplarily, the m-dimensional feature recovered from the suspicious set is used as the input of the validator. For any data pair (x, y) from the original suspicious set , the output of the validator is as follows:
[0076] ,
[0077] Among them, the samples with an output of 1 will be rejected by the validator, which indicates that the corresponding samples are identified as specific misclassifications. Correct the predicted labels of these rejected inputs to T. For the samples not rejected (the validator output is 0), it does not affect the original prediction result of the model. Denote the repaired model as , then for any sample pair (x, y) from the original data set D, the final predicted label is as follows:
[0078] ,
[0079] By converting the multi-class suspicious set into a one-vs-rest (OVR) binary classification data set . The misclassification situation in the original model is equivalent to the misclassification situation in the validator . With the help of the validator, the repair of is transformed into the repair of Repair.
[0080] It should be noted that most of the computational overhead of SMiR is consumed in the construction phase of the repair model. The only additional operations required in the forward inference phase are to recover the perturbed features in the suspicion set and to verify and repair the recovered features using the validator.
[0081] In addition, SMiR can be applied to software defect prediction for misclassification repair. Existing software defect prediction proposes to mainly classify CWE (Common Weakness Enumeration) defects into five categories: input validation and data integrity issues, computation and code injection issues, authentication and authorization issues, security configuration and management issues, and buffer issues.
[0082] According to the above theory, by classifying the above five types of software defect labels according to their severity, the following defect type ranking can be obtained:
[0083] Computation and code injection issues: The most serious, because code injection can allow an attacker to remotely execute malicious code, directly threatening system security and data integrity.
[0084] Authentication and authorization issues: The severity is second only to the above. Such issues may lead to unauthorized access, posing a direct threat to data and system functions.
[0085] Input validation and data integrity issues: Such issues can lead to malicious data being input into the system, possibly causing abnormal system behavior or data corruption.
[0086] Security configuration and management issues: Although they may not be directly exploited to execute code, improper configuration and management can be used to bypass security measures.
[0087] Buffer issues: Usually related to system stability and data security, they may cause the system to crash or be used to execute code. However, compared with the previous issues, the complexity and conditional limitations for their exploitation are more, and the severity is the lowest.
[0088] Label each type of software defect in order of severity according to the above sorting and predict the classification of software defects. For example, label the most severe calculation and code injection problems as 4 and the least severe buffer problems as 0. Then the label order is exactly consistent with the increasing order of severity. This means that any category with a larger label value misclassified as a category with a smaller label value, that is, misclassifying a category with a higher severity as a category with a lower severity, belongs to a safety-critical misclassification. Misclassifying "calculation and code injection problems" as "buffer problems" is a safety-critical misclassification because calculation and code injection problems allow attackers to directly run malicious code, causing serious damage to the system. If misclassified as buffer problems, these serious security threats may not be recognized and prevented in time, resulting in higher security risks.
[0089] For the misclassification of "calculation and code injection problems" as "buffer problems", using SMiR to repair this specific misclassification includes:
[0090] Regard "calculation and code injection problems" as category T and "buffer problems" as category P; use the software function dataset as the training set of the neural network model, and complete the repair of the specific misclassification according to the method described in SMiR. Among them, the software function dataset mainly comes from open-source projects with sufficient activity on the GitHub website. For example, Apache JMeter is a load testing tool with 123,000 lines of code. When performing multi-label software defect prediction, all Java functions in the open-source project software code can be statically analyzed, and the static code metrics among them are extracted as function features. For defective functions, they are labeled according to the CWE defect category, and for non-defective functions, they are uniformly labeled as 0. Through the above process, a multi-label defect prediction dataset for open-source projects is established, where each function corresponds to one piece of data. The static code metrics of the function evaluate the complexity, number of lines of code, etc. of the function as data input, and the corresponding defect information of the function is used as output. The neural network is trained on the constructed dataset to achieve the classification of software defect categories.
[0091] It should be noted that the above severity sorting is not strictly established, but is based on the historical experience of software testers and developers. To verify the superiority and effectiveness of the present invention in repairing specific misclassifications, the following experiments were carried out:
[0092] Three widely used datasets were used to evaluate the repair of specific misclassifications. In the experiment, the training dataset was used to guide the model repair, and the test dataset was used to evaluate the repair effect.
[0093] Fashion-MNIST (FM): FM is an image dataset containing ten categories of clothing. It has 60,000 training samples and 10,000 test samples, and each sample is a 32×32 grayscale image.
[0094] CIFAR10 (C10): C10 is an image dataset containing ten categories, covering vehicles and common animals, etc. It contains 50,000 training samples and 10,000 test samples, and each image is a 32×32 pixel RGB three-channel color picture.
[0095] SVHN: SVHN is an image dataset composed of house number images, where each image depicts a number from 0 to 9, and each image is a 32×32 pixel RGB three-channel color picture. This dataset contains 73,257 training samples and 26,032 test samples.
[0096] Based on 5 different pre-trained DNN models, for the FM dataset, FM - SimpleNet (fully connected model) is adopted, and for the C10 dataset, C10 - SimpleNet and C10 - CNN (both are convolutional neural network CNN models) are used. Specifically, SVHN has the same image resolution and data structure as the C10 dataset, so two convolutional neural network (CNN) models (SVHN - SimpleNet and SVHN - CNN) are trained, and their architectures are the same as the models applied to C10.
[0097] In this experiment, the Adam optimizer with weight decay of is applied during training. The initial learning rate is , and during a total of 200 training epochs, it decays by 0.1 times every 50 epochs, and the batch size is 256. After training is completed, the model slice of the training epoch with the highest test accuracy among the 200 epochs is selected as the final training result. As shown in Table 1, the detailed information of each model is as follows (in the number of parameters, K represents thousands and M represents millions):
[0098] Table 1
[0099]
[0100] Under a given specific misclassification , the repair is evaluated using the Repair Rate (RR) and the Introduced Misclassification (IM), which is the ratio of the number of repaired specific misclassification samples to the number of original specific misclassification samples. Let Indicates the case of misclassification of the original misclassified sample set. Then RR is expressed as:
[0101] ,
[0102] IM is defined as the increase in the overall misclassification situation after model repair. This metric is used to evaluate the degree of damage to the model's classification performance during the repair process. The calculation formula is:
[0103] ,
[0104] where represent the misclassified sample sets of the original model and the repaired model respectively.
[0105] To verify the effectiveness of the proposed method, a comparative experiment between SMiR and the baseline method was designed. The heuristic neural network repair method Arachne was used as the baseline method. SMiR and the baseline method were applied to repair the three most common (i.e., the highest frequency) specific misclassifications in each model (for example, top1 in Table 1 represents the most common misclassification), aiming to minimize the bias caused by potential random factors and obtain more stable results.
[0106] It should be noted that the repair process of specific misclassifications will inevitably damage the overall accuracy of the model. Therefore, simply comparing the repair rate RR or the number of introduced misclassifications IM alone cannot comprehensively reflect the performance of the repair method. For the sake of rigor, it is considered that only when a repair method can achieve a high repair rate and reduce the situation of introduced misclassifications, it is better than other methods. However, it is difficult to directly compare these two negatively correlated metrics. For example, it is difficult to determine whether a method that achieves a high repair rate but introduces significantly more misclassifications is better than a method that achieves a low repair rate but introduces very few misclassifications.
[0107] To be able to comprehensively compare these two metrics, the results of the baseline method Arachne (the only existing technique specifically for eliminating specific misclassifications) were used as a benchmark. Then, the parameters of the present invention were adjusted to achieve a similar repair rate. Each experimental setting was repeated five times with different random seeds, and the average results were reported. The specific experimental settings are as follows:
[0108] SMiR: The SMiR validator was trained using the Stochastic Gradient Descent (SGD) optimizer, with a batch size of 1024, applied the cosine learning rate scheduler, and set different initial learning rates according to the different complexities of the model: for FM - SimpleNet for C10 - SimpleNet , for C10 - CNN, , for SVHN - SimpleNet, , for SVHN - CNN, . After 100 training epochs, the learning rate was uniformly reduced to .
[0109] Arachne: Adopts the same data partitioning strategy as this experiment to perform specific misclassification repair. As shown in Table 2, it is a comparison table of experimental results:
[0110] Table 2
[0111]
[0112] Among them, RR represents the repair rate (presented as a percentage %), and IM represents the misclassification cases introduced during the repair process. The table cells with a light gray background indicate that the corresponding method achieved the best performance in both RR and IM metrics. The bold results mean that the corresponding method has improved in at least one of the RR or IM metrics.
[0113] As can be seen from Table 2, SMiR achieved the best average performance on all five models. For example, in the C10 - SimpleNet model, SMiR outperformed Arachne, with its repair rate increasing by 8.4% and the misclassification cases introduced decreasing by 37.4%. However, on the two SVHN - based models, Arachne only corrected 8.9% and 5.1% of the errors, which is significantly lower than the 30.3% and 22.5% achieved by SMiR. This is mainly because the evaluation during the search process relies on a sufficient number of specific misclassification instances. However, the high training accuracy of these two models results in the scarcity of such instances in the training set, thus causing inaccurate guidance for the patch search process. However, SMiR can improve the model's recognition ability for specific classes through perturbation feature recovery, thus enhancing the misclassification repair effect in the absence of specific misclassification samples.
[0114] Among the 15 specific misclassification cases investigated in the experiment, SMiR achieved the best performance in 13 of them. Even in the two exceptional cases, SMiR also showed excellent performance. When repairing the 3rd most frequent error of the SVHN - CNN model, Arachne achieved the lowest number of misclassification cases introduced (3.8), but only corrected 0.6% of the misclassification cases. In contrast, SMiR corrected more than 37 times as many misclassification cases at the cost of only 179% more misclassification cases introduced. For the 2nd most frequent error of the C10 - CNN model, Arachne introduced the fewest misclassification cases (-7.0), but only achieved a repair rate of 28.0%.
[0115] In summary, SMiR achieves a higher repair rate in most cases, with minimal damage to the standard accuracy. Even when the training set lacks sufficient specific misclassified instances to be repaired, SMiR can always maintain a high repair rate.
[0116] Compared with the prior art, a method for repairing specific misclassifications of a neural network based on perturbed feature recovery provided in this embodiment screens a specific misclassification suspicion set from the original image dataset, uses the trained feature mapping model to complete the repair of the features, and uses a validator to perform the final repair on the samples after feature repair, improving the repair accuracy in the case of specific misclassifications. On the one hand, the method uses a dedicated extractor to locate the perturbed features and trains the feature mapping model using the located perturbed features, enhancing the model's ability to distinguish between class T and class P; on the other hand, by screening only a subset from the original dataset for repair, the range of decision boundary adjustment is narrowed, maximizing the protection of the accuracy of the original model.
[0117] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0118] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A neural network specific misclassification repair method based on perturbation feature recovery, characterized in that: The steps include: Using a neural network model trained with the original image data set to perform initial classification on the image data set to be classified, the neural network model includes an initial feature extractor for extracting features of the image to be classified and a classifier for classification using the features; Screening out image sample data classified into category P from the to-be-classified image data set as a suspect set; Based on the suspicion set and the trained feature mapping model, performing perturbation feature repair on the potential features extracted by the initial feature extractor to obtain repaired features; Use the trained verifier to verify whether the repaired features correspond to specific misclassifications. If yes, repair the image sample data in the suspected set that is misclassified as category P to obtain the repaired classification results; where specific misclassification refers to misclassifying image samples belonging to category T into category P; Among them, the features that are seriously affected by redundant information are regarded as disturbance features; Performing the disturbance feature repair includes: Using the feature mapping model, the perturbed features in the potential features of the suspicion set extracted by the initial feature extractor are repaired into features extracted by a dedicated feature extractor; wherein the dedicated feature extractor is obtained by training with image samples of specific misclassified related categories; Updating the repaired disturbance features to the potential features of the suspicion set to obtain the repaired features; Use the trained validator to verify whether the repaired features correspond to specific misclassifications, including: Inputting the repaired features into the trained verifier; Reclassifying the restored features according to whether the image sample data misclassified as category P in the suspicion set has features of category T; If the repaired feature corresponds to the specific misclassified category T, the label is positive and repair is performed; otherwise, the label is negative and the initial classification result is retained; Training the validator includes: Relabeling the original image data set to obtain two categories of image samples; wherein the two categories of image samples include labeling the image sample data of category T as positive class, and labeling the sample data classified into other categories in the original image data set as negative class; Taking the two types of image samples as input data sets, training the verifier based on a binary classifier to obtain classification results including the positive class and the negative class; The dedicated feature extractor obtained by training includes: Extracting image samples of specific misclassification-related categories from the original image data set as a retraining data set; wherein the specific misclassification-related categories include the category T and the category P; The classifier of the neural network is frozen, and the initial feature extractor is retrained based on the retraining data set to obtain the dedicated feature extractor.
2. A neural network specific misclassification repair method based on perturbation feature recovery according to claim 1, characterized in that: The repaired classification result is to repair the image sample data whose label of the trained verifier is positive class to category T.
3. A neural network specific misclassification repair method based on perturbation feature recovery according to claim 1, characterized in that: The binary classifier is a classifier structure based on the neural network.
4. A neural network specific misclassification repair method based on perturbation feature recovery according to claim 1, characterized in that: Training the feature mapping model includes: Extracting potential features of the retraining data set based on the initial feature extractor, and taking perturbation features in the potential features as an input sample set; Based on the dedicated feature extractor, the dedicated features of the retraining data set are extracted, and the features corresponding to the perturbation dimension in the dedicated features are used as the output sample set; wherein the perturbation dimension is the dimension corresponding to the perturbation feature.
5. A neural network specific misclassification repair method based on perturbation feature recovery according to claim 4, characterized in that: The feature mapping model is trained based on a multi-layer perceptron and a mean square error loss function.
6. A neural network specific misclassification repair method based on perturbation feature recovery according to claim 1, characterized in that: The neural network includes a convolutional neural network, and the initial feature extractor includes a convolutional layer and a pooling layer.
Citation Information
Patent Citations
Image classification method, parameter training method and image classification device
CN115049869A