An Explanation Result Correction Method for Deep Discriminant Models Based on Expert Interactive Learning
Through expert interactive learning and loss function optimization, the problem of lack of user knowledge in the existing technology of interpreting results is solved, and the interpretability and user trust of the model are improved.
Patent Information
- Application Number
- CN202310061482.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-01-17
AI Technical Summary
Existing interpretable methods ignore user knowledge and cannot correct them when interpreting errors, and the explanation results lack user knowledge, resulting in low user trust in AI decision-making models.
Through expert interactive learning, image data is obtained and divided into predictive training sets and interpreted training sets. Resnet18 network model is used to train and add correction modules to generate interpretable heat maps, combine expert knowledge to annotate and correct the error areas, and use loss function to optimize the model to interpret the results.
The interpretability and prediction capabilities of the network model are improved, the interpretability capabilities are quantified, and the user's trust in the model is enhanced.
Smart Images

Figure CN116246144B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of interpretable technologies, and particularly relates to a method for correcting the interpretation result of a deep discriminant model based on expert interactive learning. Background Art
[0002] With the rapid development of technology and the continuous improvement of the quality of life, artificial intelligence (AI) driven by deep neural networks (DNN) is ubiquitous in the entire human society. However, when in use, especially when medical diagnoses and important decisions need to be made, the AI decision-making model cannot only provide "suggestions", but also needs to provide the "suggestions" and the corresponding logical processes. Otherwise, the AI decision-making model cannot gain the "trust" of the users who use it.
[0003] Currently, most interpretable technologies are interpretable technologies based on gradient-based visualization methods, model distillation methods that use white-box models to explain black-box models, and intrinsic methods that believe that the model needs to provide decision-making bases while providing decisions. These interpretable methods enable users to understand the basis of model decisions to a certain extent.
[0004] However, these interpretable methods ignore the role of users' prior knowledge in the training process. The interpretation methods proposed by current interpretable technologies cannot meet users' interpretation needs because these interpretation methods are for the model to interpret the model rather than for the model to interpret for users.
[0005] The problems of the prior art are:
[0006] Current interpretable methods provide interpretations through end-to-end learning of the network model. Although this gives the model a certain degree of interpretability, when there are interpretation errors, they cannot be corrected based on users' knowledge. Moreover, due to the limitations of end-to-end learning, the interpretability obtained by the model does not include users' knowledge, resulting in the lack of users' knowledge in the interpretable results. Therefore, current interpretable discriminant models face the challenges of how to correct them when there are interpretation errors and how to handle the knowledge feedback from users as the training process progresses. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a method for correcting the interpretation result of a deep discriminant model based on expert interactive learning, including:
[0008] S1: Obtain the image to be interpreted, divide the image into a prediction training set and an interpretation training set, and perform data preprocessing on the image;
[0009] S2: Train the ResNet18 network model using the preprocessed prediction training set, and add a correction module to the trained ResNet18 network model to obtain an interpretation and discrimination model;
[0010] S3: Input the images in the preprocessed interpretation training set into the interpretation and discrimination model, and obtain an interpretable heatmap showing the importance of each pixel for the model's decision by calculating the gradient value of each pixel in the image and its feature map in the model;
[0011] S4: Interact with the interpretable heatmap using expert knowledge. When the expert judges that the interpretation of the interpretable heatmap is correct, do not operate on the image. When the expert judges that the interpretation of the interpretable heatmap is incorrect, label the incorrect area and select the area that the model should correctly interpret for annotation;
[0012] S5: Input the annotation information into the correction module of the interpretation and discrimination model to correct the incorrect interpretation result;
[0013] S51: Mask the marked area with incorrect interpretation and correct the incorrect interpretation using a loss function;
[0014] S52: Use gradient region detection technology to extract the area that the model pays the most attention to from the marked area that the model should correctly interpret, and combine with the correct interpretation area provided by the user to design a loss function to correct the incorrect interpretation;
[0015] S6: Perform gradient region detection on the corrected interpretation result, obtain the number of pixels in the gradient region and the number of pixels included in the area provided by the user, calculate the evaluation index based on the number of pixels in the gradient region and the number of pixels included in the area provided by the user, set a threshold, and repeat S5 until the evaluation index is less than or equal to the set threshold, then the interpretation correction is completed.
[0016] Advantages of the present invention:
[0017] The present invention trains on a network model using an active learning training strategy and realizes the interaction between expert knowledge and the network model by combining an interpretable method; and through the interaction between the user and the network model, the correction of the interpretable result of the network model is realized, thereby improving the interpretable ability of the network model and a certain prediction ability; in addition, the proposed interpretable evaluation method quantifies the interpretable ability of the network model and can accurately evaluate the interpretable ability of the current network model. Finally, based on these optimization effects, the user's trust in the network model is improved, and the network model can more accurately judge its decision-making basis by combining expert knowledge. Description of the Drawings
[0018] Figure 1Flowchart of the method for correcting the interpretation result of the interpretation and discrimination model based on expert knowledge interactive learning of the present invention;
[0019] Figure 2 Schematic diagram of the technical process of the gradient region detection technology of the present invention. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] A method for correcting the interpretation result of a deep discrimination model based on expert interactive learning, as Figure 1 shown, includes:
[0022] S1: Obtain the image to be interpreted, divide the image into a prediction training set and an interpretation training set, and perform data preprocessing on the image;
[0023] S2: Train the resnet18 network model with the preprocessed prediction training set, and add a correction module to the trained resnet18 network model to obtain an interpretation and discrimination model;
[0024] S3: Input the images in the preprocessed interpretation training set into the interpretation and discrimination model, and obtain an interpretable heat map of the importance of each pixel to the model decision by calculating the gradient value of each pixel in the image and its feature map in the model;
[0025] S4: Interact with the interpretable heat map through expert knowledge. When the expert judges that the interpretation of the interpretable heat map is correct, no operation is performed on the image; when the expert judges that the interpretation of the interpretable heat map is incorrect, mark the incorrect interpretation area and select the area that the model should correctly interpret for marking;
[0026] S5: Input the annotation information into the correction module of the interpretation and discrimination model to correct the incorrect interpretation result;
[0027] S51: Mask the marked incorrect interpretation area and correct the incorrect interpretation through the loss function;
[0028] S52: Extract the area that the model most concerns in the marked area that the model should correctly interpret by using the gradient region detection technology, and design a loss function to correct the incorrect interpretation in combination with the correct interpretation area provided by the user;
[0029] S6: Detect the gradient regions for the corrected interpretation results. Calculate the evaluation metric based on the number of pixels in the obtained gradient regions and the number of pixels in the region provided by the user. Set a threshold, and repeat S5 until the evaluation metric is less than or equal to the set threshold, then the interpretation correction is completed.
[0030] The threshold: First, calculate the interpretable evaluation metric for each sample in the entire dataset. Then, use the 75th percentile of the interpretable evaluation metrics of the dataset samples as the threshold. Finally, compare the remaining sample metrics with this threshold. When the samples with metrics less than or equal to this threshold can basically be considered to have correct interpretations for the samples.
[0031] Download the dataset from the natural image dataset ImageNet. A total of 10 categories of natural images are obtained. From the natural image dataset ImageNet, 1450 images of each of the 10 categories are obtained. The categories are cat, dog, chair, baseball, horse, shark, bird, pot, frog, and car. Reset the image size to the input size of the resnet18 network model, which is 224×224. Perform image enhancement on the images with the reset size to obtain the preprocessed images. First, perform data processing on the prediction training set. For the natural image dataset, first reset the image size to the network model input size of 224×224 and perform data augmentation. Then, use the processed prediction training set as the input to the resnet18 network model and train the network model.
[0032] Train the resnet18 network model using the prediction training set to obtain the trained network model, including:
[0033] Input the images in the preprocessed prediction training set into the resnet18 network model. Reduce the dimension of the images through a convolutional layer with a convolutional kernel of 7×7, 64 channels, and a stride of 2. Extract the features of the input images through 2 convolutional layers with convolutional kernels of 3×3 and 64, 128, 256, and 512 channels. Perform image classification through an average pooling layer, a fully connected layer, and a softmax layer to complete the training of the resnet18 network model.
[0034] For the ResNet18 network model, the model contains five convolutional blocks A, B, C, D, and E. Among them, convolutional block A includes a convolutional layer with a 7×7 kernel, 64 channels, and a stride of 2, which is used to receive the input image and reduce its dimension; convolutional block B uses residual connections and contains a max-pooling layer and two residual blocks. At the same time, each residual block contains two convolutional layers with 3×3 kernels and 64 channels; convolutional blocks C, D, and E are similar to convolutional block B, but the number of channels is different, which are 128, 256, and 512 respectively; convolutional blocks B, C, D, and E are used to extract the features of the input image; in addition, the ResNet18 network model also includes an average pooling layer, a fully connected layer, and a softmax layer after the five convolutional blocks. These three layers are mainly used to perform image classification on the image features extracted from the convolutional blocks through the fully connected layer and the softmax layer.
[0035] By calculating the gradient value of each pixel in the image and its feature map in the model, an interpretable heatmap showing the importance of each pixel to the model's decision is obtained, including:
[0036]
[0037]
[0038] Among them, L heatMap represents the interpretable heatmap showing the importance of each pixel to the model's decision, A represents a feature layer in the network, k represents the k-th channel in feature layer A, represents the weight for A k ; c represents the class, y c represents the score predicted by the network for class c, represents the data at the position of ij in channel k of feature layer A, and Z represents the product of the width and height of feature layer A.
[0039] The interpretable heatmap includes: regions irrelevant to the model's decision and regions relevant to the model's decision; at the same time, there will be cases where the model's interpretation is incorrect, that is, the model provides incorrect basis for the decision. The specific manifestation is that regions irrelevant to the model's decision are judged as regions relevant to the model's decision. Therefore, the first interactive method for the interpretable result is to select the regions where the model's interpretation is incorrect, and the second interactive method for the interpretable result is to select the regions where the model should provide correct interpretations.
[0040] Two correction methods are provided for the regions with incorrect model interpretations. The first method is to add counterexamples, replacing the incorrect parts in the interpretable heatmap judged as incorrect by the user with counterexamples. The second method is to provide a mask for the incorrect parts in the interpretable heatmap and design a new loss function, which is used to correct the incorrect interpretations.
[0041] The second correction method masks the regions with incorrect interpretations and corrects the incorrect interpretations through the loss function, including:
[0042]
[0043] Among them, θ represents the model parameters, X represents the input feature vector, y represents the label, A represents the mask, represents the predicted value, λ1 and λ2 represent the first and second regularization coefficients, N represents the number of samples, n represents the nth sample, K represents the number of labels, k represents the kth label, D represents the dimension of the input feature vector, d represents the dth dimension of the input feature vector, A nd represents the mask of the dth dimension of the nth sample, represents the gradient of the dth dimension of the nth sample, i represents the ith layer of the model, represents the square of the model parameters of the ith layer.
[0044] The gradient region detection technique is used to extract the regions that the model pays the most attention to. As Figure 2 shown, including:
[0045] For the regions that the model should correctly interpret and are marked, calculate the gradient value of each pixel in the region, set the 95th percentile of the overall gradient value of the region as the threshold, perform auxiliary operations to set the gradient values less than the threshold to 0, and extract the pixel region in the region where the gradient value is greater than the set threshold, which is the region that the model pays the most attention to.
[0046] Extract the regions that the model pays the most attention to, and combine with the correct interpretation regions provided by the user to design a loss function to correct the incorrect interpretations, including:
[0047]
[0048] Among them, g outside represents the gradient value of each pixel of the incorrect interpretation provided by the model outside the region provided by the user, g inside represents the gradient value within the region provided by the user, p grad represents the number of pixels in the interpretable region provided by the model obtained by gradient region detection, p mask represents the number of pixels included in the region provided by the user, p grad ∩p mask represents the number of pixels in the intersection region between the two regions, p grad∪p mask Represents the total number of pixels in two regions Represents the classification loss of the sample, output represents the predicted value of the network model, and y represents the label value of the sample
[0049] Calculate evaluation metrics, including:
[0050]
[0051] Among them, I represents the evaluation metric, and p grad Represents the number of pixels in the interpretable region provided by the model obtained by gradient region detection, and p mask Represents the number of pixels included in the region provided by the user, and p grad ∩p mask Represents the number of pixels in the intersection region between two regions, and p grad ∪p mask Represents the total number of pixels in two regions
[0052] First, calculate the interpretable evaluation metrics for each sample in the entire dataset. Then, use the 75th percentile of the interpretable evaluation metrics of the dataset samples as the threshold. Put the metrics of all samples in the dataset into a list, sort them from smallest to largest, then multiply the length of the list by 0.75 to get an n. Then, the nth number in the list is the 75th percentile of this list. Compare the metrics of the remaining samples with this threshold. When the metrics of the samples less than or equal to the threshold can be basically considered that the interpretation of the sample is correct
[0053] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents
Claims
1. A method for correcting the interpretation result of a deep discriminant model based on expert interactive learning, characterized in that, Including: S1: Obtain the image to be explained, divide the image into a prediction training set and an explanation training set, and perform data preprocessing on the image; S2: Train the resnet18 network model with the preprocessed prediction training set, and add a correction module to the trained resnet18 network model to obtain an explanation discrimination model; S3: Input the images in the preprocessed explanation training set into the explanation discrimination model, and obtain an interpretable heatmap of the importance of each pixel to the model decision by calculating the gradient value of each pixel in the image and its feature map in the model; S4: Interact with the interpretable heatmap through expert knowledge. When the expert judges that the explanation of the interpretable heatmap is correct, no operation is performed on the image; when the expert judges that the explanation of the interpretable heatmap is incorrect, mark the incorrect explanation area and select the area that the model should correctly explain for marking; S5: Input the annotation information into the correction module of the explanation discrimination model to correct the incorrect explanation result; S51: Mask the marked incorrect explanation area and correct the incorrect explanation through a loss function; S52: Use the gradient region detection technique to extract the area that the model is most concerned about from the marked area that the model should correctly explain, and combine it with the correct explanation area provided by the user to design a loss function to correct the incorrect explanation; S6: Through gradient region detection of the corrected explanation result, obtain the number of pixels in the gradient region and the number of pixels included in the area provided by the user, calculate the interpretable evaluation index based on the number of pixels in the gradient region and the number of pixels included in the area provided by the user, set a threshold, and repeat S5 until the evaluation index is less than or equal to the set threshold, then the explanation correction is completed.
2. The method for correcting the interpretation result of a deep discrimination model based on expert interactive learning according to claim 1, characterized in that, Process the image, including: Reset the image size to the input size of the resnet18 network model, which is 224×224, and perform image enhancement on the resized image to obtain the preprocessed image.
3. A method for correcting the interpretation result of a deep discrimination model based on expert interactive learning according to claim 1, characterized in that Train the resnet18 network model with the prediction training set to obtain the trained network model, including: Input the images in the preprocessed prediction training set into the resnet18 network model, perform image dimensionality reduction through a convolutional layer with a 7×7 convolutional kernel, 64 channels, and a stride of 2, extract the features of the input image through 2 convolutional layers with 3×3 convolutional kernels, 64, 128, 256, and 512 channels, and perform image classification through an average pooling layer, a fully connected layer, and a softmax layer to complete the training of the resnet18 network model.
4. The method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, wherein Obtain an interpretable heatmap of the importance of each pixel to the model decision by calculating the gradient value of each pixel in the image and its feature map in the model, including: Among them, L heatMap represents an interpretable heatmap indicating the importance of each pixel for the model's decision. A represents a feature layer in the network, and k represents the k-th channel in the feature layer A. represents the weight for A k c represents the category, and y c represents the score predicted by the network for the category c. represents the data at the position of ij in the k-th channel of the feature layer A. Z represents the product of the width and height of the feature layer A. 5. A method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, characterized in that The interpretable heatmap includes: regions irrelevant to the model decision and regions relevant to the model decision.
6. The method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, wherein Mask the incorrect explanation area and correct the incorrect explanation through a loss function, including: Among them, θ represents the model parameters, X represents the input feature vector, y represents the label, A represents the mask, represents the predicted value, λ1 and λ2 represent the first and second regularization coefficients, N represents the number of samples, n represents the nth sample, K represents the number of labels, k represents the kth label, D represents the dimension of the input feature vector, d represents the dth dimension of the input feature vector, A nd represents the dth dimension mask of the nth sample, represents the gradient of the dth dimension of the nth sample, i represents the ith layer of the model, represents the square of the model parameters of the ith layer.
7. A method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, characterized in that Use the gradient region detection technique to extract the area that the model is most concerned about, including: For the regions that should be correctly interpreted by the marked model, calculate the gradient value of each pixel in the region, set the 95th percentile of the overall gradient value of the region as the threshold, and extract the pixel region in the region where the gradient value is greater than the set threshold, which is the region that the model pays the most attention to.
8. A method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, characterized in that Extract the region that the model pays the most attention to, and combine it with the correctly interpreted region provided by the user to design a loss function to correct the wrong interpretation, including: Among them, g outside represents the gradient value of each pixel of the wrong explanation provided by the out-of-region model provided by the user, and g inside represents the gradient value within the region provided by the user, and p grad represents the number of pixels in the interpretable region provided by the model detected by the gradient region, and p mask represents the number of pixels contained in the region provided by the user, and p grad ∩p mask represents the number of pixels in the intersection region between the two regions, and p grad ∪p mask represents the total number of pixels in the two regions, represents the classification loss of the sample, output represents the predicted value of the network model, and y represents the label value of the sample.
9. The method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, characterized in that Calculate evaluation metrics, including: Among them, I represents the evaluation index, p grad represents the number of pixels in the interpretable region provided by the model obtained by gradient region detection, p mask represents the number of pixels included in the region provided by the user, p grad ∩p mask represents the number of pixels in the intersection region between the two regions, p grad ∪p mask represents the total number of pixels in the two regions.
10. The method for correcting the interpretation result of a deep discriminant model based on expert interactive learning according to claim 1, characterized in that The threshold, including: calculate the interpretable evaluation metrics of all the obtained images, and set the 75th percentile of the interpretable evaluation metrics of all the images as the threshold.