Construction method of neural network specific misclassification repair model
By building a neural network specific misclassification repair model, using special feature extractors and feature mapping models for feature repair, the problem of imbalance in the existing technology of specific categories' distinction ability and original classification accuracy is solved, and efficient specific misclassification repair and model classification accuracy are achieved.
Patent Information
- Application Number
- CN202510146968.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
When building classification models in the prior art, the ability to distinguish specific categories from the original classification accuracy is unbalanced, making it difficult to effectively solve specific misclassification problems.
By building a neural network-specific misclassification repair model, using a special feature extractor to locate distorted features, training a feature map model for feature repair, and correcting potential specific misclassification through a validator to improve the model's ability to distinguish specific categories.
It realizes that while maintaining the accuracy of the original classification of the model, the repair effect of specific misclassification is improved, the model's ability to identify specific categories is improved, and unnecessary computing and storage costs are reduced.
Smart Images

Figure CN120071102A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of misclassification repair, and particularly to a method for constructing a neural network specific misclassification repair model. Background Art
[0002] In the past few years, deep neural networks (DNNs) have attracted great attention from practitioners due to their excellent performance in many cutting-edge applications. Among them, classification tasks are of great significance in the field of deep learning. Especially with the development of deep neural networks, image classification has become the core of many practical applications. In autonomous driving, accurately identifying image information such as road signs, vehicles, and pedestrians is a key factor to ensure driving safety.
[0003] However, the potential hazards brought about by the large-scale application of neural networks have also increasingly attracted people's attention. Specific misclassification of a neural network model refers to the situation where the model misclassifies an input sample of one category as another category, making its predicted category not conform to the actual category of the sample. The harms caused by different types of specific misclassifications vary a lot, but the consequences of some specific misclassifications are relatively serious. For example, in the traffic sign recognition task, misidentifying a "no entry" sign as a "speed limit (45 km / h)" sign may lead to traffic accidents.
[0004] In order to maintain the reliability of DNN models in practical applications, a series of repair technologies for DNN models have emerged. For example, QuoTe and DeepPatch aim to alleviate misclassifications caused by adversarial samples, including through adversarial training on the original training set and training on the generated adversarial sample dataset of perturbations. However, such repair technologies require additional synthetic samples, which will consume a large amount of time and computing resources. In the case of no additional data, such as Apricot by learning the overall weights of its reference model, and cost-sensitive learning techniques by giving more attention to minority classes to reduce the dominance of majority classes, or by giving greater weights to specific misclassified samples of the type to be repaired in retraining to reduce the corresponding specific misclassifications, without improving the model's recognition ability for the corresponding classes. But there is still a problem of imbalance between the specific category discrimination ability and the original classification accuracy of the model. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present invention aim to provide a method for constructing a neural network specific misclassification repair model to solve the problem of imbalance between the specific category discrimination ability and the original classification accuracy of the classification model constructed by the prior art.
[0006] The object of the present invention is mainly achieved by the following technical solutions:
[0007] On the one hand, an embodiment of the present invention provides a method for constructing a neural network specific misclassification repair model, including the following steps:
[0008] Input the original training set into a general feature extractor for feature extraction, and input the extracted multi-dimensional features into a classifier for image classification. After training, an initial neural network model based on the general feature extractor and the classifier is obtained; wherein, the original training set is a set of labeled image samples;
[0009] Screen out sample data related to specific misclassification from the original training set to obtain a specific training set, and train a dedicated feature extractor based on the specific training set; wherein, the specific misclassification refers to misclassifying a sample belonging to the T class label as the F class label;
[0010] Compare the features output by the dedicated feature extractor and the general feature extractor to locate the distorted features; use the distorted features as input to construct a feature mapping model for repairing the distorted features of the multi-dimensional features extracted by the general feature extractor.
[0011] Train a validator based on the multi-dimensional features corresponding to the re-labeled original training set. The validator is used to repair the sample data misclassified as the F class label to the T class label.
[0012] Further, locating the distorted features includes:
[0013] Quantify the contribution of each dimension feature output by the dedicated feature extractor and the general feature extractor to correct classification respectively to obtain a first group of contribution vectors and a second group of contribution vectors;
[0014] Calculate the difference between the first group of contribution vectors and the second group of contribution vectors to obtain a contribution difference vector containing the relative contribution differences of each dimension feature;
[0015] According to the characteristic that the greater the relative contribution difference, the more severely the corresponding dimension feature is affected by redundant information, locate the dimension of the distorted feature severely affected by redundant information;
[0016] Locate the feature corresponding to the dimension in the multi-dimensional features output by the feature extractor to obtain the distorted feature.
[0017] Further, locating the dimension of the distorted feature includes:
[0018] Set the total number of perturbation dimensions d;
[0019] Take the first d dimensions with larger relative contribution differences in the contribution difference vector as the dimensions of the distorted features.
[0020] Further, based on the model interpretability technique DeepLift, the contributions of the features in each dimension output by the dedicated feature extractor and the feature extractor to the correct classification are quantified respectively, and the sum of the contributions of all the dimensional features in each group of contribution vectors is 1.
[0021] Further, training the feature mapping model based on the located distorted features includes:
[0022] Extracting multi-dimensional features respectively by using the dedicated feature extractor and the feature extractor based on the specific training set to obtain a first group of features and a second group of features;
[0023] Taking the distorted features in the first group of features as the input sample set for training the feature mapping model; extracting the corresponding features in the second group of features according to the dimension of the distorted features as the output sample set for training the feature mapping model.
[0024] Further, the dedicated feature extractor is only used for the location of the distorted features and the training of the feature mapping model, and the feature mapping model is directly used for feature recovery in the forward inference stage.
[0025] Further, training to obtain the dedicated feature extractor includes:
[0026] Based on the architecture of the initial neural network remaining unchanged, retraining the general feature extractor by using the specific training set and the cross-entropy loss function for training the initial neural network to obtain the dedicated feature extractor, wherein the unchanged architecture includes the classifier configuration remaining unchanged.
[0027] Further, training the feature mapping model based on a multi-layer perceptron and the mean square error loss function.
[0028] Further, training the verifier includes:
[0029] Relabeling the original training set, labeling the samples with the T-class label in the specific misclassifications as the first class, and labeling the sample data with other class labels in the original training set as the second class;
[0030] Training the verifier with the labeled sample set to obtain the corresponding classification result.
[0031] Further, the neural network includes a convolutional neural network and a support vector machine.
[0032] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0033] 1. The present invention proposes to identify specific misclassified distorted features based on a dedicated feature extractor. To effectively eliminate redundant information, a feature mapping model is trained based on the located distorted features to achieve the restoration of the distorted features. A verifier is trained based on the restored features to repair the misclassifications. This not only enables the original model to achieve the repair effect of specific misclassifications but also maintains the accuracy of the original classification of the model.
[0034] 2. Retrain the original general feature extractor on a subset that only contains categories related to specific misclassifications to obtain a dedicated feature extractor, which is used for the localization and repair of perturbed features, improving the model's discrimination ability for categories related to specific misclassifications.
[0035] 3. Instead of directly adjusting the weights of the model, use the dedicated feature extractor to locate the distorted features and train a feature mapping model for feature repair, reducing unnecessary computational and storage costs.
[0036] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the content specifically pointed out in the specification and the drawings. Description of the Drawings
[0037] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components.
[0038] Figure 1 It is a flowchart of a method for constructing a neural network specific misclassification repair model according to an embodiment of the present invention;
[0039] Figure 2 It is a framework diagram of the technical solution according to an embodiment of the present invention;
[0040] Figure 3 It is a schematic diagram of the distorted feature localization and feature restoration according to an embodiment of the present invention. Detailed Embodiments
[0041] Next, the preferred embodiments of the present invention will be specifically described with reference to the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, not for limiting the scope of the present invention.
[0042] A specific embodiment of the present invention discloses a method for constructing a neural network specific misclassification repair model, as Figure 1 shown, including the following steps:
[0043] Step S1: Input the original training set into a general feature extractor for feature extraction, and input the extracted multi-dimensional features into a classifier for image classification. After training, an initial neural network model based on the general feature extractor and the classifier is obtained. Among them, the original training set is a set of labeled image samples; the set of labeled image samples can be a set of image samples of various road signs such as no entry, speed limit, turning, different vehicles and pedestrians, and the corresponding labels include road signs, vehicles, and pedestrians.
[0044] Step S2: Screen out the sample data related to specific misclassification in the original training set to obtain a specific training set, and train a dedicated feature extractor based on the specific training set. Among them, the specific misclassification refers to misclassifying a sample belonging to the T class label as the F class label.
[0045] Step S3: Compare the features output by the dedicated feature extractor and the general feature extractor to locate the distorted features. Using the distorted features as input, construct a feature mapping model for repairing the distorted features in the multi-dimensional features extracted by the general feature extractor.
[0046] Step S4: Train a validator based on the multi-dimensional features corresponding to the re-labeled original training set. The validator is used to repair the sample data misclassified as the F class label into the T class label.
[0047] Through the above method, the dimension of the distorted features is located by using a dedicated feature extractor, enhancing the ability to distinguish specific categories. The distorted features are repaired by constructing a feature mapping model, and a validator is constructed to correct potential specific misclassifications after repairing the perturbed features, ensuring that the overall classification accuracy of the model is not reduced while improving the accuracy of repairing specific misclassifications.
[0048] It should be noted that the specific misclassification is a certain misclassification prediction that occurs in the neural network model. If the relevant categories of the specific misclassification are represented by the T class label and the F class label, then the corresponding specific misclassification refers to misclassifying a sample belonging to the T class label as the F class label. As Figure 2 shown, the repair model includes training a dedicated feature extractor with a feature training set screened from the original training set, locating the distorted features based on the multi-dimensional features output by the dedicated feature extractor and the original general extractor, and training a feature mapping model with the distorted features to complete the repair of the distorted features in the multi-dimensional features extracted by the general extractor. Then, the repaired features are input into the validator to correct potential specific misclassifications.
[0049] Specifically, in step S1, the neural network, as an end-to-end machine learning model, mainly consists of two components: a feature extractor and a classifier. Multidimensional features are extracted from the input original training set through a general feature extractor; the multidimensional features output by the general feature extractor are differentiated, predicted for classification, and the predicted class labels are output through a classifier; under the guidance of the classical cross-entropy loss, the initial neural network is trained; the neural network model is a convolutional neural network or a support vector machine.
[0050] It should be noted that the general feature extractor in the multi-classification model is expected to capture the differences between all classes in a limited number of feature dimensions. This general feature extraction mode learned from all classes helps the model achieve the highest overall classification performance. Usually, it cannot capture all the differences required to effectively separate two specific classes, and is doped with many other redundant features. These redundant features are extracted to describe the difference information between other classes, but they hinder the distinction between these two specific classes, resulting in the inability to accurately extract the features of specific classes. For any two specific classes, the extracted latent features are mixed with the redundant features initially learned to distinguish other classes. Eliminating these redundant information can prevent them from hindering the distinction of specific classes.
[0051] To address the limitation of the insufficient discrimination ability between specific classes of the model, a dedicated feature extractor for specifically extracting the features of a given misclassified class is trained using samples related to specific misclassification. In step S2, the specific steps for training the dedicated feature extractor include:
[0052] Step S21: In the original training set, extract the sample data of the classes related to specific misclassification (i.e., the samples with T class labels and F class labels) as the specific training set.
[0053] Step S22: Based on the unchanged architecture of the initial neural network, freeze the classifier component, and retrain the general feature extractor using the specific training set and the cross-entropy loss function for training the initial neural network to obtain the dedicated feature extractor.
[0054] Among them, the dedicated feature extractor can be regarded as a general feature extractor that removes redundant information. Compared with the initial general feature extractor trained to maximize the overall accuracy, it has a higher classification accuracy on the corresponding specific two classes and better captures the features for distinguishing relevant classes.
[0055] Although the features extracted by the dedicated feature extractor improve the model's recognition ability for classes T and F, they also weaken the initial neural network model's recognition ability for other classes except T and F. In addition, introducing an additional dedicated feature extractor not only increases the storage cost, but also increases the computational cost due to the corresponding processing of feature extraction. To solve this problem, identify and restore those severely perturbed feature dimensions among the features extracted by the feature extractor in the initial neural network model architecture, and use the dedicated feature extractor only for locating the distorted features during the repair process and training the feature mapping model, rather than directly applying the features extracted by the dedicated feature extractor to the model. Among them, the feature dimensions severely affected by redundant information are regarded as distorted features (Distorted Features, DF).
[0056] Specifically, in step S3, based on the multi-dimensional features extracted from classes T and F by the dedicated feature extractor and the feature extractor respectively, compare the features output by the two extractors, and locate the distorted features by measuring the contribution values of each dimension feature in the feature to the correct classification of specific class samples before and after removing redundancy. As Figure 3 shown, the specific steps for locating distorted features include:
[0057] Step S311: Quantify the contribution of each dimension feature output by the dedicated feature extractor and the redundant feature extractor to the correct classification respectively, to obtain the relative contribution of each dimension feature, and the sum of the contributions of all feature dimensions is equal to 1;
[0058] Exemplarily, use the model interpretability technique (Deep Learning Important FeaTures, DeepLift) to quantify the contribution of each feature dimension to the correct classification on a specific training set. Define the contribution of the features extracted by the general feature extractor as the first group of contribution vectors, denoted as c R =[c 1 R ,c 2 R ,…,c n R , where represents the relative contribution of the feature dimension i in the general feature extractor. Similarly, define the contribution vector of the features extracted by the dedicated feature extractor as the second group of contribution vectors, denoted by c D .
[0059] Step S312: To measure the perturbation degree of each feature dimension, calculate the difference between the first group of contribution vectors and the second group of contribution vectors, to obtain the relative contribution difference of each dimension feature, that is, calculate the reduction value of the contribution degree of each feature dimension before and after removing redundant features.
[0060] Exemplarily, for a specific dimension i, Indicates the increased value of the contribution of the features extracted by the dedicated feature extractor compared to the features extracted by the general feature extractor.
[0061] Step S313: According to the characteristic that the greater the relative contribution difference, the more severely the corresponding dimensional feature is affected by redundant information, locate the dimension of the distorted feature severely affected by redundant information;
[0062] It should be noted that Δc i The higher the relative contribution difference, the more the feature dimension can contribute to the correct identification of the corresponding category after removing redundant information, which means that the redundant information has a more severe impact on dimension i. According to this characteristic, the feature dimensions with a large increase in contribution after redundancy removal are regarded as the dimensions severely polluted by redundant information in the classification of a specific category, and the features corresponding to the dimensions are the distorted features.
[0063] Step S314: Locate the feature corresponding to the dimension in the multi-dimensional features output by the general feature extractor to obtain the distorted feature.
[0064] Furthermore, screen out the dimensions that meet the preset requirement for the number of distorted feature dimensions among all relative contributions as the dimensions of the distorted features. This includes setting the total number of perturbation dimensions d; taking the first d dimensions with larger relative contribution differences in the contribution difference vector as the dimensions of the distorted features.
[0065] Exemplarily, let the contribution values of the features extracted by the general feature extractor and the dedicated feature extractor be c R =[0.1, 0.1, 0.2, 0.1, 0.1, 0.2] and c D =[0.1, 0.4, 0.1, 0.1, 0.2, 0.1]. Through Δc = c D -c R It is easy to obtain the contribution difference as Let the preset threshold value of the distorted feature dimension be set as d = 2, then locate the 2 dimensions with the highest contribution degree difference and identify them as the perturbed feature dimensions DF, which are dimension 2 and dimension 5 in turn (emphasized by the symbol *).
[0066] Furthermore, feature restoration is used to eliminate the redundant information that affects the identification of a specific category on the perturbed dimension. Actually, it constructs a feature mapping from dimension d to dimension d, that is, R d -R d . Restore the features corresponding to the located distorted feature dimensions through the feature mapping model to reduce the adverse effects of redundant information therein. Training the feature mapping model based on the located distorted features includes:
[0067] Step S321: Extract multi-dimensional features respectively using the dedicated feature extractor and the redundant feature extractor based on the specific training set, obtaining a first set of features and a second set of features;
[0068] Step S322: Use the distorted features in the first set of features as the input sample set for training the feature mapping model; extract the corresponding features in the second set of features according to the dimension of the distorted features as the output sample set for training the feature mapping model;
[0069] Step S323: Train the feature mapping model based on a multi-layer perceptron and a mean squared error loss function.
[0070] Exemplarily, a multi-layer perceptron (MLP) model with two hidden layers is used as the feature mapping model. Each hidden layer contains 128 neurons, and the input and output dimensions are determined by the number d of localization dimensions. In the forward inference stage, the trained lightweight feature mapping model is used to separately recover the localization dimensions of the distorted features.
[0071] In the above manner, not only the advantages and effective information of the features extracted by the general feature extractor and the dedicated feature extractor are retained, but also additional computational and storage overheads are avoided during the forward inference process, and the feature mapping model is directly used for feature recovery in the forward inference stage.
[0072] Specifically, the training of the validator in step S4 includes:
[0073] Step S41: Re-label the original training set, label the samples with the T-class label in the specific misclassification as the first class, and label the sample data with other class labels in the original training set as the second class;
[0074] Step S42: Train the validator with the labeled sample set to obtain the corresponding classification results.
[0075] Exemplarily, a binary classifier with the same structure as the classifier of the initial neural network model is used to train the validator, and the validator is used to repair the sample data misclassified as the F-class label to the T-class label to obtain the classification results for the corresponding two classes.
[0076] It should be noted that the sample features after feature recovery are used as the input and transmitted to the validator, rather than directly inputting the samples into the validator. The validator evaluates and checks whether the input belongs to the specific misclassification, that is, whether it is the first class; if so, it is repaired to the T-class label. In the forward inference stage, only the trained feature mapping model is used to recover the distorted features, and the trained validator is used to verify the recovered features to complete the repair of the specific misclassification.
[0077] Exemplarily, an initial classification is performed on the image to be classified based on the initial neural network; images classified as F-class labels are extracted from the image to be classified; the distorted features in the multi-dimensional features corresponding to the images are input into the trained feature mapping model for feature repair to obtain the repaired multi-dimensional features; then the repaired multi-dimensional features are input into the trained validator for repairing specific misclassification results, and the image data misclassified as F-class labels is repaired to T-class labels.
[0078] The repair model constructed by the method can be applied to the repair of software defect misclassification. Most existing software defect prediction studies tend to focus on distinguishing whether a given software module has software defects, that is, discriminating whether the software has defects (labeled as 1) or no defects (labeled as 0). There are also those who use CWE (Common Weakness Enumeration) as the fine-grained model prediction label, formulating the traditional binary classification defect prediction task as a multi-label classification problem for defect categories. Xing Ying et al. mainly divide CWE defects into five categories of defects: input validation and data integrity issues, calculation and code injection issues, authentication and authorization issues, security configuration and management issues, and buffer issues.
[0079] According to the above theory, dividing the above five types of software defect labels according to the severity level, the following defect type ranking can be obtained:
[0080] Calculation and code injection issues: The most serious, because code injection can allow attackers to remotely execute malicious code, directly threatening system security and data integrity.
[0081] Authentication and authorization issues: The severity is second only to the above. Such issues may lead to unauthorized access, posing a direct threat to data and system functions.
[0082] Input validation and data integrity issues: Such issues will cause malicious data to be input into the system, which may lead to abnormal system behavior or data corruption.
[0083] Security configuration and management issues: Although it may not be directly exploited to execute code, improper configuration and management can be used to bypass security measures.
[0084] Buffer issues: Usually related to system stability and data security, may cause the system to crash or be used to execute code, but compared with the previous issues, the complexity and conditional restrictions for its exploitation are more, and the severity is the lowest.
[0085] Classify software defects of each type in the order of severity according to the above sorting to complete software defect classification. That is, label the most serious calculation and code injection problems as 4, and label the least serious buffer problems as 0. Then the label order is exactly consistent with the increasing order of severity. This means that any category with a larger label value being misclassified as a category with a smaller label value, that is, misclassifying a category with a higher severity as a category with a lower severity, belongs to a safety-critical misclassification.
[0086] Misclassifying "calculation and code injection problems" as "buffer problems" is a safety-critical misclassification because calculation and code injection problems allow attackers to directly run malicious code, causing serious damage to the system. If misclassified as buffer problems, it may lead to these serious security threats not being recognized and prevented in a timely manner, thus bringing higher security risks. Therefore, this type of misclassification should be repaired first. Relatively speaking, misclassifying "buffer problems" as "calculation and code injection problems", although equally inaccurate, usually does not directly result in immediate system damage and is relatively more tolerable, with a lower level of urgency.
[0087] For the misclassification of "calculation and code injection problems" as "buffer problems", the repair of this specific misclassification based on the specific classification repair model of the neural network includes:
[0088] Regard "calculation and code injection problems" as the T class label and "buffer problems" as the F class label; use the software function dataset as the training set of the initial neural network model, and build a model according to the above method to complete the repair of the specific misclassification. Among them, the software function dataset mainly comes from open-source projects with sufficient activity on the GitHub website. For example, Apache JMeter is a load testing tool with 123,000 lines of code. When performing multi-label software defect prediction, all Java functions in the open-source project software code can be statically analyzed, and the static code metrics among them are extracted as function features. For defective functions, they are labeled according to the CWE defect category, and for non-defective functions, they are uniformly labeled as 0. Through the above process, a multi-label defect prediction dataset for open-source projects is established, where each function corresponds to one piece of data. The static code metrics of the function evaluate the complexity, number of lines of code, etc. of the function, which are used as data inputs, and the corresponding defect information of the function is used as the output. Train the initial neural network based on the above data inputs to achieve the classification of software defect categories.
[0089] Compared with the prior art, a method for constructing a neural network specific misclassification repair model provided by this embodiment locates distorted features using a dedicated extractor and trains a feature mapping model using the located distorted features to complete the repair of the features. By restoring the distorted features, the recognition ability of the model for specific categories is improved, and the final repair is performed in combination with a validator, achieving an improvement in the repair accuracy of specific misclassification cases with minimal damage to the standard accuracy of the model's classification.
[0090] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0091] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for constructing a neural network specific misclassification repair model, characterized in that: The steps include: Inputting the original training set into a universal feature extractor for feature extraction, and inputting the extracted multi-dimensional features into a classifier for image classification, and obtaining an initial neural network model based on the universal feature extractor and the classifier after training; wherein the original training set is a set of labeled image samples; Filtering sample data related to specific misclassification in the original training set to obtain a specific training set, and training a dedicated feature extractor based on the specific training set; wherein the specific misclassification refers to misclassifying samples belonging to class T labels into class F labels; Comparing the features output by the dedicated feature extractor and the general feature extractor to locate the distorted features; using the distorted features as input, constructing a feature mapping model for repairing the distorted features of the multidimensional features extracted by the general feature extractor; A verifier is trained based on the multidimensional features corresponding to the relabeled original training set, and the verifier is used to repair sample data misclassified as F-type labels to T-type labels.
2. The method for constructing a neural network specific misclassification repair model according to claim 1, characterized in that: Locating the distortion feature includes: quantifying the contribution of the features of each dimension output by the dedicated feature extractor and the general feature extractor to the correct classification respectively, to obtain a first group of contribution vectors and a second group of contribution vectors; Calculate the difference between the first group of contribution vectors and the second group of contribution vectors to obtain a contribution difference vector containing the relative contribution difference of each dimensional feature; According to the characteristic that the larger the relative contribution difference is, the more seriously the corresponding dimension feature is affected by redundant information, the dimension of distorted features that are seriously affected by redundant information is located; The feature corresponding to the dimension is located in the multi-dimensional features output by the universal feature extractor to obtain the distortion feature.
3. The method for constructing a neural network specific misclassification repair model according to claim 2, characterized in that: The dimensions for locating the distortion features include: Set the total number of perturbation dimensions d; The dimensions of the first d largest relative contribution differences in the contribution difference vector are used as the dimensions of the distortion feature.
4. The method for constructing a neural network specific misclassification repair model according to claim 2, characterized in that: Based on the model interpretability technology DeepLift, the contribution of each dimensional feature output by the dedicated feature extractor and the general feature extractor to the correct classification is quantified respectively, and the sum of the contributions of all dimensional features in each group of contribution vectors is 1.
5. The method for constructing a neural network specific misclassification repair model according to claim 2, characterized in that: Training the feature mapping model based on the localized distortion features includes: Based on the specific training set, respectively, the dedicated feature extractor and the universal feature extractor are used to extract multidimensional features to obtain a first set of features and a second set of features; The distorted features in the first group of features are used as an input sample set for training a feature mapping model; and corresponding features in the second group of features are extracted according to the dimensions of the distorted features as an output sample set for training a feature mapping model.
6. The method for constructing a neural network specific misclassification repair model according to claim 5, characterized in that: The dedicated feature extractor is only used for locating the distortion features and training the feature mapping model, and the trained feature mapping model is directly used for feature recovery in the forward reasoning stage.
7. The method for constructing a neural network specific misclassification repair model according to claim 6, characterized in that: The dedicated feature extractor obtained by training includes: Based on the unchanged architecture of the initial neural network, the general feature extractor is retrained using the specific training set and the cross entropy loss function for training the initial neural network to obtain the dedicated feature extractor, wherein the unchanged architecture includes that the classifier configuration remains unchanged.
8. The method for constructing a neural network specific misclassification repair model according to claim 5, characterized in that: The feature mapping model is trained based on a multi-layer perceptron and a mean square error loss function.
9. A method for constructing a neural network specific misclassification repair model according to any one of claims 1 to 8, characterized in that: Training the validator includes: Relabeling the original training set, labeling the samples with T class labels in the specific misclassification as the first class, and labeling the sample data with other class labels in the original training set as the second class; The verifier is trained with the labeled sample set to obtain the corresponding classification result.
10. The method for constructing a neural network specific misclassification repair model according to claim 1, characterized in that: The neural network includes a convolutional neural network and a support vector machine.