Enhancement Method, Device, Computer Equipment and Storage Medium for Training Data
By screening and enhancing the training data of the convolutional neural network model, using the residual network to extract false features and generate an enhanced training data set, the problem of insufficient generalization ability of the model is solved and the robustness and data diversity of the model are improved.
Patent Information
- Application Number
- CN202210358348.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-06
AI Technical Summary
The existing convolutional neural network classification model lacks generalization ability, resulting in low model robustness. The existing methods cannot fundamentally improve the model by increasing the amount of training data.
By obtaining the training data set, input it into the classification model for pre-training, filtering out false features with a greater impact than the preset value, using the residual network to extract features, perform data augmentation, generate an enhanced training data set and retrain the model.
Improve the robustness of the model, reduce the changes in prediction results caused by environmental or posture changes, and enhance the diversity of training data.
Smart Images

Figure CN114743067B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of model training, and particularly to a method, device, computer device, and storage medium for augmenting training data. Background Art
[0002] Classification methods based on convolutional neural networks currently show good performance in multiple tasks, but their generalization ability is still limited. In different works, the performance of the model may vary greatly, which affects the user's trust in the model. Currently, to solve the problem of incorrect model recognition results, mostly increasing the amount of training data is used, and the model cannot be fundamentally improved, resulting in low robustness of the model. Summary of the Invention
[0003] The main purpose of this application is to provide a method, device, computer device, and storage medium for augmenting training data, aiming to solve the problem that the robustness of the model is low due to the existence of incorrect data in the training data that significantly affects the model results.
[0004] To achieve the above invention purpose, this application proposes a method for augmenting training data, the method comprising:
[0005] Obtain a training data set;
[0006] Input the training data set into a classification model for pre-training, and obtain the original images and classification data with incorrect classifications;
[0007] Input the original images into a preset residual network to obtain the features of the original images;
[0008] Calculate the mutual information of the features according to the classification data, and screen out the false features in the features with an influence degree greater than a preset value according to the mutual information;
[0009] Perform data augmentation on the original images according to the false features to obtain target images;
[0010] Generate an augmented training data set according to the target images, and re-train the classification model based on the augmented training data set.
[0011] Further, after screening out the false features in the features with an influence degree greater than a preset value according to the mutual information, it further comprises:
[0012] Input the false features into a decision tree for training;
[0013] Obtain the leaf node with the highest error rate in the decision tree, and determine the target false features according to the leaf node.
[0014] Further, the data augmentation of the original image according to the false feature to obtain a target image includes:
[0015] Normalize the target false feature and enlarge the normalized target false feature to the same size as the original image to obtain a heat map;
[0016] Overlay and fuse the heat map with the original image to perform data augmentation on the original image to obtain a target image.
[0017] Further, the overlay and fusion of the heat map with the original image to perform data augmentation on the original image to obtain a target image includes:
[0018] Perform masking processing on the area corresponding to the false feature in the original image according to the heat map to perform data augmentation on the original image to obtain a target image.
[0019] Further, the calculation of the mutual information of the feature according to the classification data includes:
[0020] Obtain the wrong label data in the classification data;
[0021] Calculate the mutual information between each feature and the wrong label data to obtain the mutual information of each feature.
[0022] Further, the input of the original image into a preset residual network to obtain the features of the original image includes:
[0023] Input the original image into a preset residual network, and the residual network includes a Resnet50 network;
[0024] Extract the high-dimensional features of the original image based on the last convolutional layer of the Resnet50 network.
[0025] Further, after generating an augmented training dataset according to the target image and retraining the classification model based on the augmented training dataset, it further includes:
[0026] Obtain the accuracy rate of the augmented training dataset;
[0027] When the accuracy rate of the augmented training dataset is greater than a preset value, output the classification model.
[0028] This application also provides an apparatus for augmenting training data, and the apparatus includes:
[0029] A data acquisition module, configured to acquire a training dataset;
[0030] A pre-training module for inputting the training data set into a classification model for pre-training to obtain original images with classification errors and classification data;
[0031] A feature extraction module for inputting the original images into a preset residual network to obtain the features of the original images;
[0032] A false feature module for calculating the mutual information of the features according to the classification data and screening out false features with an influence degree greater than a preset value in the features according to the mutual information;
[0033] A data augmentation module for augmenting the original images according to the false features to obtain target images;
[0034] A re-training module for generating an augmented training data set according to the target images to re-train the classification model based on the augmented training data set.
[0035] The present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the enhancement method of the training data described in any one of the above is implemented.
[0036] The present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the enhancement method of the training data described in any one of the above is implemented.
[0037] This application example provides a method for enhancing the training data of a classification model based on error attribution. First, a training data set is obtained, and then the unprocessed training data set is input into the classification model for pre-training. By pre-training, each parameter in the classification model is corrected so that the classification model can identify the classification corresponding to the images in the training data set with the highest accuracy. During the pre-training process, the classification model makes mistakes in classifying some images in the training data set. The original images with classification errors are obtained, and the classification data of the original images are obtained. The original images are input into a preset residual network. The residual network can effectively identify different images and extract features in the images that can provide more information for image classification, so as to obtain the features of the original images. Then, the influence degree of the features on the misclassification of the original images is characterized by mutual information. The mutual information of the features is calculated according to the classification data, and the false features with an influence degree greater than a preset value are screened out from the features according to the mutual information. The original images are data-augmented according to the false features to obtain target images, that is, an enhanced training data set is generated according to the target images and the correct classification of the original images. Then, the enhanced training data set is input into the classification model again for training, so as to re-train the classification model based on the enhanced training data set. The classification model can more accurately adjust the parameters of the classification model according to the enhanced training data set, so as to accurately classify the images. By tracing the results of the recognition errors of the classification model, false features related to the predicted wrong results are found, and the original images are enhanced based on the false features, so as to obtain a feature-confused data set, that is, an enhanced training data set. Then, the classification model is trained based on the enhanced training data set to improve the diversity of the training data of the classification model and reduce the change of the prediction results caused by environmental or pose changes, thereby improving the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic flowchart of an embodiment of the method for enhancing the training data of this application;
[0039] Figure 2 It is a schematic flowchart of an embodiment of determining target false features of this application;
[0040] Figure 3 It is a schematic flowchart of an embodiment of data-augmenting the original images of this application;
[0041] Figure 4 It is a schematic flowchart of an embodiment of superimposing and fusing the heat map and the original image to data-augment the original image to obtain a target image of this application;
[0042] Figure 5Schematic flowchart of an embodiment for calculating the mutual information of the features of this application;
[0043] Figure 6 Schematic flowchart of an embodiment for obtaining the features of the original image of this application;
[0044] Figure 7 Schematic flowchart of an embodiment after generating an enhanced training dataset according to the target image of this application and retraining the classification model based on the enhanced training dataset;
[0045] Figure 8 Schematic structural diagram of an embodiment of the training data enhancement device of this application;
[0046] Figure 9 Schematic block diagram of an embodiment of the computer device of this application.
[0047] The realization, functional features and advantages of the purpose of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0048] In order to make the purpose, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0049] Refer to Figure 1 , an embodiment of this application provides a method for enhancing training data. The method for enhancing training data includes steps S101-S106, and the detailed description of each step of the method for enhancing training data is as follows.
[0050] S101. Obtain a training dataset.
[0051] This embodiment is applied to the data enhancement scenario of model training. In this scenario, a classification model based on a convolutional neural network is currently applied in multiple tasks, but the generalization ability of this classification model is still limited. In different works, the performance of the classification model may vary greatly, and the accuracy rate of the classification model varies greatly, which affects the user's trust in the model. In this embodiment, during the training of the classification model, data enhancement is performed on the clusters with a relatively high error rate in the training set, and then retraining is performed based on the enhanced training set. Specifically, first, a training dataset is obtained. In one implementation manner, the training set contains several images of the same classification. For example, 100 images classified as "black-footed albatross", and the training set also contains images of different classifications. For example, it not only contains 100 images classified as "black-footed albatross", but also contains 100 images classified as "alpaca".
[0052] S102. Input the training data set into the classification model for pre-training to obtain the original images with classification errors and classification data.
[0053] In this embodiment, after obtaining the training data set, it is first necessary to train the classification model. The process of inputting the unprocessed training data set into the classification model for training is defined as pre-training, that is, inputting the training data set into the classification model for pre-training. By pre-training, each parameter in the classification model is corrected so that the classification model can identify the classification corresponding to the images in the training data set with the highest accuracy. During the pre-training process, the classification model cannot correctly identify and classify some images in the training data set, that is, the classification of some images in the training data set by the classification model is incorrect. At this time, obtain the images with classification errors by the classification model, define the images with classification errors as the original images, and obtain the classification data of the original images, including obtaining the correct classification data of the original images and the incorrect classification data identified by the model.
[0054] S103. Input the original images into a preset residual network to obtain the features of the original images.
[0055] In this embodiment, after inputting the training data set into the classification model for pre-training to obtain the original images with classification errors and classification data, input the original images into a preset residual network. The preset residual network can, through pre-training, identify the features in the images that can provide information for image classification, so as to identify and obtain the features of the original images based on the residual network. In one implementation, the residual network is trained on the CUB2011 data set. The cross-entropy loss function is used during the training process to measure the difference information between two probability distributions, and the Adam optimizer is used to optimize the initial parameters in the residual network. And set the learning rate of the residual network to be adaptively adjusted according to the result accuracy. Among them, the initial learning rate is 0.1. The trained residual network can effectively identify different images and extract the features in the images that can provide more information for image classification.
[0056] S104. Calculate the mutual information of the features according to the classification data, and screen out the spurious features in the features with an influence degree greater than a preset value according to the mutual information.
[0057] In this embodiment, after the original image is input into a preset residual network to obtain the features of the original image, when the features included in the misclassified original image are extracted, it is necessary to determine the degree of influence of the features in the original image on the misclassification. Specifically, the mutual information is used to characterize the degree of influence of the features on the misclassification of the original image. Mutual information is used to evaluate the amount of information contributed by the occurrence of one event to the occurrence of another event. That is, the mutual information of the features is calculated according to the classification data, and the false features with an influence degree greater than a preset value are screened out from the features. In one implementation, the top 10 features with the highest influence degree are selected from the features as false features. The false features are important features for the classification model to misclassify the original image.
[0058] S105. Perform data augmentation on the original image according to the false features to obtain a target image.
[0059] In this embodiment, after calculating the mutual information of the features according to the classification data and screening out the false features with an influence degree greater than a preset value from the features according to the mutual information, data augmentation is performed on the original image according to the false features to obtain a target image. In one implementation, the false features perform data augmentation on the corresponding area of the original image. The data augmentation includes the contrast and saturation of the corresponding area of the original image, replacing the corresponding area of the original image with other images, etc., so as to obtain a target image.
[0060] S106. Generate an augmented training dataset according to the target image to retrain the classification model based on the augmented training dataset.
[0061] In this embodiment, after performing data augmentation on the original image according to the false features to obtain a target image, an augmented training dataset is generated according to the target image, that is, an augmented training dataset is generated according to the target image and the correct classification of the original image, and then the augmented training dataset is input into the classification model again for training to retrain the classification model based on the augmented training dataset. Since the false features in the original image are augmented, the classification model can more accurately adjust the parameters of the classification model according to the augmented training dataset, so as to accurately classify the image. By tracing the results of the misidentifications of the classification model, false features related to the mispredicted results are found, and the original image is augmented based on the false features, so that a feature-confused dataset, that is, an augmented training dataset, can be obtained, and then the classification model is trained based on the augmented training dataset to improve the diversity of the training data of the classification model and reduce the change of the prediction results caused by environmental or pose changes, thereby improving the robustness of the model.
[0062] This embodiment provides a method for enhancing the training data of a classification model based on error attribution. First, a training data set is obtained, and then the unprocessed training data set is input into the classification model for pre-training. By pre-training, each parameter in the classification model is corrected, so that the classification model can identify the classification corresponding to the image in the training data set with the highest accuracy. During the pre-training process, the classification model makes mistakes in classifying some images in the training data set. The original images with classification errors are obtained, and the classification data of the original images are obtained. The original images are input into a preset residual network. The residual network can effectively identify different images and extract features in the images that can provide more information for image classification, so as to obtain the features of the original images. Then, the influence degree of the features on the classification error of the original images is characterized by mutual information. The mutual information of the features is calculated according to the classification data, and the false features with an influence degree greater than a preset value are screened out from the features according to the mutual information. The original images are data-enhanced according to the false features to obtain target images, that is, an enhanced training data set is generated according to the target images and the correct classification of the original images. Then, the enhanced training data set is input into the classification model again for training, so as to re-train the classification model based on the enhanced training data set. The classification model can more accurately adjust the parameters of the classification model according to the enhanced training data set, so as to accurately classify images. By tracing the results of the classification model with recognition errors, false features related to the predicted wrong results are found, and the original images are enhanced based on the false features, so that a feature-confused data set, that is, an enhanced training data set, can be obtained. Then, the classification model is trained based on the enhanced training data set to improve the diversity of the training data of the classification model and reduce the change of the prediction results caused by environmental or pose changes, thereby improving the robustness of the model.
[0063] In one embodiment, as Figure 2 shown, after screening out the false features with an influence degree greater than a preset value from the features according to the mutual information, the method further includes steps S201-202:
[0064] S201, input the false features into a decision tree for training;
[0065] S202, obtain the leaf node with the highest error rate in the decision tree, and determine the target false feature according to the leaf node.
[0066] In this embodiment, after screening out the false features with an influence degree greater than a preset value among the features according to the mutual information, in order to more accurately screen out the features that affect the prediction error of the classification model, the false features are input into a decision tree for training, and then the leaf node with the highest error rate in the decision tree is obtained. That is, each of the false features is input into the decision tree, and the data and features are divided based on the decision tree, that is, the classification data and the false features are divided. The classification data is used as the root node, and each false feature is used as a leaf node for training. Then, the error rate of each leaf node is obtained, and then the leaf node with the highest error rate is selected. The target false feature is determined according to the leaf node, so as to screen out the false feature that has the greatest influence on the wrong prediction result of the classification model. Subsequently, data augmentation is performed on the original image according to the false feature. By tracing the result of the classification model's misrecognition, false features related to the wrong prediction result are found, and the original image is augmented based on the false feature, thereby improving the diversity of the training data.
[0067] In one embodiment, as Figure 3 shown, the data augmentation of the original image according to the false feature to obtain a target image includes steps S301 - S302:
[0068] S301, normalize the target false feature, and scale the normalized target false feature to the same size as the original image to obtain a heat map;
[0069] S302, superimpose and fuse the heat map and the original image to perform data augmentation on the original image to obtain a target image.
[0070] In this embodiment, during the process of performing data augmentation on the original image according to the false feature to obtain a target image, the target false feature is normalized. In one implementation, the target false feature is normalized to the interval [0, 1], and then the normalized target false feature is scaled. Specifically, the normalized target false feature is scaled to the same size as the original image to obtain a heat map, so as to accurately determine the position of the target false feature in the original image. Then, the heat map and the original image are superimposed and fused to perform data augmentation on the original image to obtain a target image, thereby improving the diversity of the training data.
[0071] In one embodiment, as Figure 4 shown, the superimposing and fusing the heat map and the original image to perform data augmentation on the original image to obtain a target image further includes step S401:
[0072] S401. Mask the area corresponding to the false feature in the original image according to the heat map to perform data augmentation on the original image, and obtain a target image.
[0073] In this embodiment, in the process of superimposing and fusing the heat map and the original image to perform data augmentation on the original image and obtain a target image, mask the area corresponding to the false feature in the original image according to the heat map to perform data augmentation on the original image and obtain a target image. In one implementation, perform Gaussian blur processing on the area corresponding to the false feature in the original image according to the heat map, so as to perform data augmentation on the original image. The masked false feature can reduce the influence on the classification result of the classification model, thereby improving the diversity of training data.
[0074] In one embodiment, as Figure 5 shown, calculating the mutual information of the feature according to the classification data further includes steps S501 - S502:
[0075] S501. Obtain the mislabeled data in the classification data;
[0076] S502. Calculate the mutual information between each feature and the mislabeled data to obtain the mutual information of each feature.
[0077] In this embodiment, in the process of calculating the mutual information of the feature according to the classification data, obtain the mislabeled data in the classification data. For several original images belonging to the same classification result, due to errors in the recognition process of the classification model, the misclassification results obtained for these original images are diverse. Convert each misclassification result into mislabeled data, and then calculate the mutual information between each feature and the mislabeled data to obtain the mutual information of each feature. Based on the mutual information, screen the features, which can accurately and efficiently determine the false features related to the mispredicted results, and perform enhancement on the original image based on the false features, thereby improving the diversity of training data.
[0078] In one embodiment, as Figure 6 shown, inputting the original image into a preset residual network to obtain the feature of the original image further includes steps S601 - S602:
[0079] S601. Input the original image into a preset residual network, and the residual network includes a Resnet50 network;
[0080] S602. Extract the high-dimensional feature of the original image based on the last convolutional layer of the Resnet50 network.
[0081] In this embodiment, in the process of inputting the original image into a preset residual network to obtain the features of the original image, the original image is input into the preset residual network, and the residual network includes a Resnet50 network. Then, high-dimensional features of the original image are extracted based on the last convolutional layer of the Resnet50 network. Feature extraction is performed on the original image through multiple convolutional layers of the Resnet50 network, and only the high-dimensional features of the original image extracted by the last convolutional layer of the Resnet50 network are retained. This can effectively extract the features in the original image that provide information for the classification result, thereby reducing the number of extracted features, reducing the computational amount, and improving the accuracy of feature extraction.
[0082] In one embodiment, as Figure 7 shown, after generating an augmented training dataset according to the target image and retraining the classification model based on the augmented training dataset, the method further includes steps S701 - S702:
[0083] S701, obtaining the accuracy rate of the augmented training dataset;
[0084] S702, when the accuracy rate of the augmented training dataset is greater than a preset value, outputting the classification model.
[0085] In this embodiment, after generating an augmented training dataset according to the target image and retraining the classification model based on the augmented training dataset, the accuracy rate of the augmented training dataset is obtained. When the accuracy rate of the augmented training dataset is greater than a preset value, the classification model is output. When the accuracy rate of the augmented training dataset is lower than the preset value, other false features can be screened again and then the target image and the augmented training dataset are regenerated to reduce the change in the prediction result caused by environmental or pose changes, thereby improving the robustness of the model.
[0086] Referring to Figure 8 , the present application further provides an apparatus for augmenting training data, including:
[0087] A data acquisition module 101, configured to acquire a training dataset;
[0088] A pre-training module 102, configured to input the training dataset into a classification model for pre-training to obtain the original images with classification errors and classification data;
[0089] A feature extraction module 103, configured to input the original image into a preset residual network to obtain the features of the original image;
[0090] The false feature module 104 is used to calculate the mutual information of the features according to the classification data, and screen the false features in the features with an influence degree greater than a preset value according to the mutual information;
[0091] The data augmentation module 105 is used to perform data augmentation on the original image according to the false features to obtain a target image;
[0092] The retraining module 106 is used to generate an augmented training data set according to the target image, and retrain the classification model based on the augmented training data set.
[0093] As described above, it can be understood that each component of the training data augmentation device proposed in this application can implement the functions of any one of the above-mentioned training data augmentation methods.
[0094] In one embodiment, after screening the false features in the features with an influence degree greater than a preset value according to the mutual information, it further includes:
[0095] Input the false features into a decision tree for training;
[0096] Obtain the leaf node with the highest error rate in the decision tree, and determine the target false feature according to the leaf node.
[0097] In one embodiment, performing data augmentation on the original image according to the false features to obtain a target image includes:
[0098] Normalize the target false feature, and enlarge the normalized target false feature to the same size as the original image to obtain a heat map;
[0099] Superimpose and fuse the heat map with the original image to perform data augmentation on the original image to obtain a target image.
[0100] In one embodiment, superimposing and fusing the heat map with the original image to perform data augmentation on the original image to obtain a target image includes:
[0101] Perform masking processing on the area corresponding to the false feature in the original image according to the heat map to perform data augmentation on the original image to obtain a target image.
[0102] In one embodiment, calculating the mutual information of the features according to the classification data includes:
[0103] Obtain the error label data in the classification data;
[0104] Calculate the mutual information between each of the features and the error label data to obtain the mutual information of each of the features.
[0105] In one embodiment, the step of inputting the original image into a preset residual network to obtain the features of the original image includes:
[0106] Input the original image into a preset residual network, where the residual network includes a Resnet50 network;
[0107] Extract high-dimensional features of the original image based on the last convolutional layer of the Resnet50 network.
[0108] In one embodiment, after generating an augmented training dataset according to the target image and retraining the classification model based on the augmented training dataset, the method further includes:
[0109] Obtain the accuracy rate of the augmented training dataset;
[0110] When the accuracy rate of the augmented training dataset is greater than a preset value, output the classification model.
[0111] Refer to Figure 9 , in an embodiment of the present application, a computer device is further provided. The computer device may be a mobile terminal, and its internal structure may be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, a display device, and an input device connected through a system bus. Among them, the network interface of the computer device is used to communicate with an external terminal through a network connection. The display device of the computer device is used to display an offline application. The input device of the computer device is used to receive user input in the offline application. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium. The non-volatile storage medium stores an operating system, a computer program, and a database. The database of the computer device is used to store original data. When the computer program is executed by the processor, it implements a method for augmenting training data.
[0112] The above-mentioned processor executes the above-mentioned method for augmenting training data, and the method includes: obtaining a training dataset; inputting the training dataset into a classification model for pre-training to obtain the original images with classification errors and classification data; inputting the original images into a preset residual network to obtain the features of the original images; calculating the mutual information of the features according to the classification data, and screening out false features with an influence degree greater than a preset value in the features according to the mutual information; performing data augmentation on the original images according to the false features to obtain target images; generating an augmented training dataset according to the target images, and retraining the classification model based on the augmented training dataset.
[0113] The computer device provides a method for enhancing the training data of a classification model based on error attribution. First, a training data set is obtained, and then the unprocessed training data set is input into the classification model for pre-training. By pre-training, each parameter in the classification model is corrected so that the classification model can identify the classification corresponding to the images in the training data set with the highest accuracy. During the pre-training process, the classification model makes mistakes in classifying some images in the training data set. The original images with classification errors and the classification data of the original images are obtained. The original images are input into a preset residual network. The residual network can effectively identify different images and extract features in the images that can provide more information for image classification, so as to obtain the features of the original images. Then, the influence degree of the features on the misclassification of the original images is characterized by mutual information. The mutual information of the features is calculated according to the classification data, and the spurious features with an influence degree greater than a preset value are screened out from the features according to the mutual information. The original images are data-augmented according to the spurious features to obtain target images, that is, an augmented training data set is generated according to the target images and the correct classifications of the original images. Then, the augmented training data set is input into the classification model again for training to re-train the classification model based on the augmented training data set. The classification model can more accurately adjust the parameters of the classification model according to the augmented training data set, so as to accurately classify the images. By tracing the results of the classification model with recognition errors, spurious features related to the predicted wrong results are found, and the original images are augmented based on the spurious features, so that a feature-confused data set, that is, an augmented training data set, can be obtained. Then, the classification model is trained based on the augmented training data set to improve the diversity of the training data of the classification model and reduce the change of the prediction results caused by environmental or pose changes, thereby improving the robustness of the model.
[0114] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, a method for enhancing training data is implemented, including the steps of: obtaining a training data set; inputting the training data set into a classification model for pre-training to obtain original images with classification errors and classification data; inputting the original images into a preset residual network to obtain the features of the original images; calculating the mutual information of the features according to the classification data, and screening out spurious features with an influence degree greater than a preset value from the features according to the mutual information; performing data augmentation on the original images according to the spurious features to obtain target images; generating an augmented training data set according to the target images to re-train the classification model based on the augmented training data set.
[0115] The computer-readable storage medium provides a method for enhancing the training data of a classification model based on error attribution. First, a training data set is obtained, and then the unprocessed training data set is input into the classification model for pre-training. By pre-training, each parameter in the classification model is corrected so that the classification model can identify the classification corresponding to the images in the training data set with the maximum accuracy. During the pre-training process, the classification model makes mistakes in classifying some images in the training data set. The original images with classification errors are obtained, and the classification data of the original images are obtained. The original images are input into a preset residual network. The residual network can effectively identify different images and extract features in the images that can provide more information for image classification, so as to obtain the features of the original images. Then, the influence degree of the features on the classification error of the original images is characterized by mutual information. The mutual information of the features is calculated according to the classification data, and the false features with an influence degree greater than a preset value in the features are screened according to the mutual information. The original images are data-augmented according to the false features to obtain target images, that is, an enhanced training data set is generated according to the target images and the correct classifications of the original images. Then, the enhanced training data set is input into the classification model again for training, so as to retrain the classification model based on the enhanced training data set. The classification model can more accurately adjust the parameters of the classification model according to the enhanced training data set, so as to accurately classify the images. By tracing the results of the classification model with recognition errors, false features related to the predicted error results are found, and the original images are enhanced based on the false features, so that a feature-confused data set, that is, an enhanced training data set, can be obtained. Then, the classification model is trained based on the enhanced training data set to improve the diversity of the training data of the classification model and reduce the change of the prediction results caused by environmental or pose changes, thereby improving the robustness of the model.
[0116] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0117] It should be noted that in this document, the terms "including", "comprising", or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, apparatus, article, or method. Without further limitation, an element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method that includes such element.
[0118] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A method for augmenting training data, characterized in that, The method includes: Obtain a training data set; Input the training data set into a classification model for pre-training to obtain the original images with classification errors and classification data; Input the original images into a preset residual network to obtain the features of the original images; Calculate the mutual information of the features according to the classification data, and filter out the spurious features in the features with an influence degree greater than a preset value according to the mutual information; Perform data augmentation on the original images according to the spurious features to obtain target images; Generate an augmented training data set according to the target images to re-train the classification model based on the augmented training data set; After filtering out the spurious features in the features with an influence degree greater than a preset value according to the mutual information, it further includes: Input the spurious features into a decision tree for training; Obtain the leaf node with the highest error rate in the decision tree, and determine the target spurious features according to the leaf node; Performing data augmentation on the original images according to the spurious features to obtain target images includes: Normalize the target spurious features, and magnify the normalized target spurious features to the same size as the original images to obtain heat maps; Overlay and fuse the heat maps with the original images to perform data augmentation on the original images to obtain target images.
2. The method for augmenting training data according to claim 1, wherein, Overlaying and fusing the heat maps with the original images to perform data augmentation on the original images to obtain target images includes: Perform masking processing on the area corresponding to the spurious features in the original images according to the heat maps to perform data augmentation on the original images to obtain target images.
3. The method for augmenting training data according to claim 1, wherein Calculating the mutual information of the features according to the classification data includes: Obtain the error label data in the classification data; Calculate the mutual information between each feature and the error label data to obtain the mutual information of each feature.
4. The method for augmenting training data according to claim 1, wherein Inputting the original images into a preset residual network to obtain the features of the original images includes: Input the original images into a preset residual network, and the residual network includes a Resnet50 network; Extract the high-dimensional features of the original images based on the last convolutional layer of the Resnet50 network.
5. The method for augmenting training data according to claim 1, wherein After generating an augmented training data set according to the target images to re-train the classification model based on the augmented training data set, it further includes: Obtain the accuracy rate of the augmented training data set; When the accuracy rate of the augmented training data set is greater than a preset value, output the classification model.
6. An enhancement device for training data, which is used to implement the method according to any one of claims 1-5, characterized in that, The device includes: A data acquisition module for obtaining a training data set; A pre-training module for inputting the training data set into a classification model for pre-training to obtain the original images with classification errors and classification data; A feature extraction module for inputting the original images into a preset residual network to obtain the features of the original images; A spurious feature module for calculating the mutual information of the features according to the classification data, and filtering out the spurious features in the features with an influence degree greater than a preset value according to the mutual information; A data augmentation module, configured to perform data augmentation on the original image according to the false features to obtain a target image; A retraining module, configured to generate an augmented training data set according to the target image, so as to retrain the classification model based on the augmented training data set.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method for augmenting the training data according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method for augmenting the training data according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Image recognition method and device, medium and electronic equipment thereof
CN112580544A
Error correction method and device for training data, equipment and storage medium
CN112766387A
Feature selection method and device based on conditional mutual information, equipment and storage medium
CN113761026A