An interpretable convolutional neural network image classification method based on binary tree structure embedding
By embedding interpretable modules based on binary tree structure in deep convolutional neural networks, the problem of poor interpretability in image classification tasks is solved, and the analysis and quantification of classification decision paths and credibility are realized, which improves the interpretability and stability of the task.
Patent Information
- Application Number
- CN202211561810.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-07
AI Technical Summary
Deep convolutional neural networks have poor interpretability in image classification tasks, which are difficult to provide sufficient information and security, which limits their application in fields such as autonomous driving systems and facial recognition systems.
A method of image classification of interpretable convolutional neural networks based on binary tree structure embedding is proposed. By embedding interpretable modules in convolutional neural networks, the features are processed using binary tree structures, and the classification decision path and classification credibility are extracted.
It improves the interpretability of image classification tasks, can analyze classification decision paths, quantify the credibility of classification decisions, and judges the stability of sample classification decisions, which improves the understanding and control of deep neural network classification decisions.
Smart Images

Figure CN115775337B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning interpretability, and specifically relates to an interpretable convolutional neural network image classification method based on binary tree structure embedding. Background Art
[0002] In recent years, deep convolutional neural networks have been deployed and applied in many fields, such as image processing, natural language processing, speech recognition, etc., and their excellent performance in some of these fields has reached or even exceeded that of humans. However, the learning and decision-making processes of the hidden layers of deep convolutional neural networks are still considered to be "black box" structures, mainly because their high complexity and non-linearity result in low transparency and poor interpretability. In many fields, there are concerns about the application of deep convolutional neural network models because the models themselves cannot provide sufficient information and security. For example, application systems built around artificial intelligence (AI) usually involve and affect many fields, such as autonomous driving systems, face recognition systems, auxiliary medical systems, etc. Considering the above challenges, the usefulness and security of these artificial intelligence systems will be limited due to the difficulty of understanding, interpreting, and controlling them. Therefore, deep convolutional neural network models are expected to make greater progress theoretically to improve people's understanding of them.
[0003] Image classification is one of the three basic tasks in computer vision. It classifies different types of targets with the smallest classification error based on the information and labels in the image. Deep learning first demonstrated great potential in image classification tasks, and the local connectivity and translational invariance of convolutional networks perfectly match the characteristics of image data. While the accuracy of various novel deep models has improved in classification tasks, problems such as fairness, generalization, and robustness have also emerged in deep models, which also pose higher requirements for interpretable classification models.
[0004] The main existing interpretation methods for neural network-based image classification are as follows: First, visualization methods. Visualization methods achieve an intuitive interpretation by presenting the feature knowledge learned by the neural network in image classification tasks to humans, which helps to understand and explain the working mechanism of the classification decision when the deep network performs classification tasks; Second, surrogate models. A surrogate model is a low-complexity and well-interpretable alternative model that mimics the original image classification model, retaining the original excellent performance while reducing complexity; Third, feature space decoupling. The feature decoupling method refers to the semantic separation of the feature expressions learned from the classification network model. By controlling some specific encoding modules, different semantic features of different classes in a classification task are learned, and image classification is achieved based on these features, which inherently contains a certain degree of interpretability. Summary of the Invention
[0005] To solve the problems existing in the above prior art, the present invention proposes an interpretable convolutional neural network image classification method based on the embedding of a binary tree structure, and the method includes:
[0006] Obtain the original image and preprocess the original image;
[0007] Input the preprocessed original image into a convolutional neural network embedded with an interpretable module with a binary tree structure after training;
[0008] Use the first convolutional layer of the convolutional neural network model to extract the first feature;
[0009] Use the interpretable module with a binary tree structure of the convolutional neural network model to process the first feature, and extract the second feature and the neuron activation values of each binary tree branch;
[0010] Use the second convolutional layer of the convolutional neural network model to process the first feature, and extract the third feature;
[0011] Use the third convolutional layer of the convolutional neural network model to fuse the second feature and the third feature, and extract the fourth feature;
[0012] Use the fully connected layer of the convolutional neural network model to classify the fourth feature to obtain the classification result of the original image, and use the neuron activation values of each binary tree branch to obtain the classification decision path and classification credibility of the original image.
[0013] Advantages of the present invention:
[0014] The present invention embeds an interpretable module into a traditional deep convolutional neural network, ensuring the accuracy of the image classification task and improving the interpretability of the classification decision of the classification network. First, by embedding an interpretable module in a conventional convolutional neural network, the accuracy of image classification remains at a relatively high level; second, the present invention can analyze the classification decision path of the convolutional neural network performing the classification task. In addition to the correctness of the classification decision, the error link of the misclassified decision can also be obtained by comparing with the standard decision path. By calculating the activation values of the two-sided branch neural networks, the classification decision credibility of a binary tree branch of a feature can be quantified, that is, for the classification of an image, the trustworthiness of each decision can be intuitively seen; by comprehensively comparing the decision path of a sample classification decision with the standard decision path, the trustworthiness of the classification decision of the deep neural network for this sample can be quantified; third, by adding noise perturbation to the input training samples, the change of activation values and the deviation of the decision path can be analyzed to judge the stability of a classification decision for a sample. Figure 2 By adding noise perturbation to the input training samples, the change of activation values and the deviation of the decision path can be analyzed to judge the stability of a classification decision for a sample. Description of the Drawings
[0015] Figure 1 It is a flowchart of an interpretable convolutional neural network image classification method based on binary tree structure embedding according to an embodiment of the present invention;
[0016] Figure 2 It is a schematic diagram of an interpretable module embedded in the neural network of the present invention;
[0017] Figure 3 It is a feature of the present invention Figure 2 Binary tree branch structure diagram;
[0018] Figure 4 It is a schematic diagram of a sample passing through all decision branches of the interpretable module in an embodiment of the present invention. Specific implementation manners
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] The image classification method provided by the embodiments of the present invention can be applied to a computer application environment. Specifically, the image classification method is applied in an image classification system, and the image classification system includes a client and a server. The client communicates with the server through a network to solve the problems of low accuracy and weak interpretability of image classification. Among them, the client, also known as the user side, refers to a program that provides local services corresponding to the server. The client can be installed on, but not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0021] Figure 1 It is a flowchart of an interpretable convolutional neural network image classification method based on binary tree structure embedding according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0022] 101. Obtain an original image and preprocess the original image;
[0023] In an embodiment of the present invention, it can be understood that the preprocessing instruction can be sent by the user through devices such as mobile terminals and computers, or can be an original image automatically generated after the user inputs the original image resolution, the number of original image samplings, and the initial image with the original image resolution. Among them, the original image resolution refers to the target value that the user or the image classification model specifies to convert images with various different resolutions into images with the same resolution; the number of original image samplings refers to the number of times of scale conversion (such as zooming times) that need to be performed on the initial image during the image classification preprocessing process; the original image can be images in different application scenarios. Exemplarily, the original image can be an ID photo, a pathological photo, etc., and the original image resolution is the image resolution of the initial image. Exemplarily, assuming that in an application scenario, an ID photo of 300*300 needs to be converted into an ID photo of 4*4 after five scale conversions, where 300*300 is the initial image resolution, the number of original image samplings is five, and the target image resolution is 4*4.
[0024] For the convenience of description, in this embodiment, the original image can be divided into a training image and a test image. The training image contains classification labels, and the test image does not contain classification labels; the training image is used to train the neural network model, and the test image is used to test the neural network model and obtain the corresponding test results, and this test result can be used for the classification interpretation of the test image.
[0025] 102. Input the preprocessed original image into a convolutional neural network embedded with an interpretable module having a binary tree structure after training;
[0026] In an embodiment of the present invention, after preprocessing the original image, the original image can be input into a convolutional neural network embedded with an interpretable module having a binary tree structure after training for classification and recognition; among them, in some embodiments, the structure of the convolutional neural network embedded with an interpretable module having a binary tree structure is as follows:
[0027] Select a convolutional block in ResNet or DenseNet or other common classification networks as the embedding entry, use the feature information extracted by the previous convolutional block as the input of this embedding module, and as the output of the next convolutional block. The embedding module contains multiple feature Figure 2 branches of the binary tree, and the feature Figure 2 The number of branches of the binary tree is determined by the number of classification categories in the classification task.
[0028] Figure 2 is a schematic diagram of the embedding interpretable module in the neural network of the present invention; as Figure 2As shown, the convolutional neural network embedded with the interpretable module of the binary tree structure at least includes a first convolutional layer, a second convolutional layer, and a third convolutional layer; each convolutional layer contains multiple convolutional blocks, so each convolutional layer can be connected through the corresponding convolutional blocks; wherein the interpretable module of the binary tree structure is located between the first convolutional layer and the third convolutional layer and is juxtaposed with the second convolutional layer; it shows that the embedding position of the neural network is outside the backbone of the neural network.
[0029] In the embodiment of the present invention, during the training process of the convolutional neural network embedded with the interpretable module of the binary tree structure, it is mainly divided into two pre-training and re-training processes, which are specifically as follows:
[0030] Input the training images into the convolutional neural network embedded with the interpretable module of the binary tree structure for pre-training. By minimizing the total loss function, when the pre-training ends, the neuron activation values of each binary tree branch and the classification decision path of the training images are obtained; input the training images superimposed with perturbations into the convolutional neural network embedded with the interpretable module of the binary tree structure for re-training, and when the re-training ends, the neuron activation values of each binary tree branch and the classification decision path of the training images are obtained; by comparing the corresponding neuron activation values and classification decision paths in the pre-training and re-training processes, when the comparison value is less than the preset threshold, output the convolutional neural network embedded with the interpretable module of the binary tree structure after stabilization; otherwise, continue to use the training images to re-perform pre-training.
[0031] Since in the process of constructing the interpretable module, at each time of feature Figure 2 forking, a loss function is used to control the decoupling of the category information corresponding to the two branches. Assume that the feature map to be further extracted at a certain time is x, and the convolutional layer extraction functions for decoupling the two branches are abstracted into functions g(x,θ1), g(x,θ2); then assume that the average activation levels of the two feature maps are f1(x) and f2(x) respectively, then there are:
[0032] f1(x) = AVG(g(x,θ1))
[0033] f2(x) = AVG(g(x,θ2))
[0034] In the above two formulas, the function g() represents the linear and non-linear operations in the convolutional layer, and AVG() represents the global average activation value of the neurons in the last layer of feature maps, which can be regarded as how much information a neural node allows to pass through. If f1(x) is greater than a relatively large constant, it can be considered that the node on the left participates in the information decoupling of this category (similarly for f2(x)); if f1(x) is less than a relatively small constant close to 0, it can be considered that the node on the left does not participate in the information decoupling of this category (similarly for f2(x)). This patent sets two loss functions for each of the two branches of the feature map branch each time, which are used to specify that a certain branch decouples the information corresponding to a category. Therefore, the loss values of the neural networks on both sides of the branch are respectively expressed as:
[0035]
[0036] Among them, f i1 (x) represents the average activation level of the left-branch neural network of the feature map x at the i-th binary tree branch; f i2 (x) represents the average activation level of the right-branch neural network of the feature map x at the i-th binary tree branch; δ represents a constant greater than 0. Usually, when the average activation value of a neuron is greater than a certain constant, we consider that the neuron participates in the decoupling of category information. Through the loss function; J i1 (θ) controls the left-branch neural network to decouple a part of the category information, and through the loss function J i2 (θ) controls the right-branch neural network to decouple another part of the category information.
[0037] Suppose that in a certain feature map branch, the two sides decouple the information of the i-th category and the j-th category respectively (note that the i-th category and the j-th category are not necessarily a single category. As the number of branches increases, the number of categories will become smaller and smaller). When the input sample belongs to the i-th category, the value of f i1 (x) will tend to a constant, while the value of f i2 (x) will tend to 0, so as to control the convolutional neural networks on both sides of the branch to decouple the information of different categories.
[0038] The original loss function L0 for classification is expressed as:
[0039]
[0040] Among them, y is the label value, is the predicted value; J i1 (θ) and J i2 (θ) will be added as additional loss function terms to the total loss function L, and the calculation formula of the total loss function L is:
[0041]
[0042] Among them, L0 represents the classification loss function; λ represents the balance factor, and λ is used to balance the classification performance and interpretability performance of the network. According to the general neural network training method, when the training cycle reaches the set threshold, the training stage ends; the weights of the network structure are saved and the testing stage is entered. After multiple binary tree branches in an interpretable convolutional neural network, a total of N branch neural networks will be obtained, and the loss value of each of them will be added to the final loss function L.
[0043] Since the label of each sample is fixed, then the standard decision path of each sample is fixed. Therefore, at each feature Figure 2 binary tree branch, it can be determined whether the network decision is correct according to whether it is consistent with the standard decision path.
[0044] 103. Extract the first feature by using the first convolutional layer of the convolutional neural network model;
[0045] In the embodiment of the present invention, similar to the traditional classification model, in this embodiment, the preprocessed original image needs to be input into the neural network model, and the first feature is extracted by the first convolutional layer in the backbone network, where the first feature does not mean only one feature, and it may be a set of multiple feature maps. Here, for the convenience of description, it is referred to as the first feature.
[0046] 104. Process the first feature by using the interpretable module of the binary tree structure of the convolutional neural network model, and extract the second feature and the neuron activation values of each binary tree branch;
[0047] In the embodiment of the present invention, each feature map of the first feature is input into the interpretable module of the binary tree structure, so that one side branch neural network decouples the information of the first half of the categories; the other side branch neural network decouples the information of the second half of the categories; until the last layer of branch neural network decouples the feature information of a single category, and the feature information of each single category is fused to obtain the second feature; and the average activation level of each feature map on each side of each layer of branch neural network is calculated.
[0048] In the embodiment of the present invention, Figure 3 is a feature of the present invention Figure 2 binary tree branch structure diagram, as Figure 3 shown, the figure records the single feature in the interpretable module Figure 2In the case of the fork tree branches, assuming that the feature information extracted each time is x, when branching, let one side branch neural network decouple the information of the first half of the categories, and the extracted feature information is g(x, θ1), and the other side branch neural network decouples the information of the second half of the categories, and the extracted feature information is g(x, θ2). During the process of decoupling separately, ensure that the sum of the number of feature maps of the two side branch neural networks after each branch is the same as that before the branch, so as to ensure the structural equivalence between the embedding module and the backbone module. For the two side branch neural networks after each branch, the feature information obtained after passing through multiple convolutional layers will also perform binary tree branching, and the branching rule is the same as above. Until after multiple branchings, the neural network on each side has decoupled the feature information of a single category, then the features in the interpretable module Figure 2 The decoupling process of the fork tree branches ends.
[0049] 105. Use the second convolutional layer of the convolutional neural network model to process the first feature and extract the third feature;
[0050] In the embodiment of the present invention, similar to the traditional classification model, in this embodiment, it is necessary to continue to perform convolutional operation processing on the first feature extracted by the first convolutional layer in the backbone network. The third feature does not represent only one feature, and it may be a set of multiple feature maps. Here, for the convenience of description, it is referred to as the third feature.
[0051] 106. Use the third convolutional layer of the convolutional neural network model to fuse the second feature and the third feature and extract the fourth feature;
[0052] In the embodiment of the present invention, considering that after the feature maps are operated by the interpretable module, at the last layer of the interpretable module, the branches of each feature map will form multiple isolated leaf node feature information. Therefore, it is necessary to fuse the feature information of all categories through a convolutional layer to obtain the second feature, and then fuse it with the third feature of the backbone network.
[0053] 107. Use the fully connected layer of the convolutional neural network model to classify the fourth feature to obtain the classification result of the original image, and use the neuron activation values of each binary tree branch to obtain the classification decision path and classification credibility of the original image.
[0054] In the embodiment of the present invention, according to the relative magnitudes of the average activation levels of the feature maps of the first feature in each side branch neural network, the classification decision path of the original image is determined; combining the ratio of the average activation levels and the classification decision path, the classification credibility of the decision of each side branch neural network is calculated, and the classification credibility of the decisions of all branch neural networks is comprehensively evaluated to obtain the classification credibility of the original image.
[0055] In the embodiments of the present invention, the fused fourth feature can continue to be operated and transformed in the subsequent convolutional layers of the convolutional neural network until the final fully connected layer to obtain the corresponding classification result. In addition, the output of the convolutional neural network modified by the present invention further includes the neuron activation values at each branch point in the interpretable module after the original image to be measured passes through the entire convolutional neural network, thereby obtaining the classification decision path and classification credibility of the original image, further improving the classification effect of the image.
[0056] Considering that the evaluation of only giving correct or incorrect for sample decisions is too single, in order to quantify the credibility of decisions, we assume that the more consistent the network decision path is with the standard decision path, the more credible the decision is. Therefore, by comparing the relative magnitudes and ratios of the average activation values of the neurons on both sides of the binary tree branch for each feature Figure 2 to determine the credibility of each branch decision. Assume that the average activation values of the neurons on both sides of the binary tree branch of each feature map x are f i1 (x) and f i2 (x), and use the ratio of the larger value and the smaller value of the two values as the credibility index r of the i-th branch decision:
[0057]
[0058] That is, the greater the difference between the two numerical values, the more credible this decision is.
[0059] And based on the comprehensive evaluation of the credibility indices of all branch decisions when a sample passes through the interpretable module, the credibility R of the network's classification decision for a sample can be obtained:
[0060]
[0061] where i represents the serial number of all the feature Figure 2 binary tree branches of a sample, R represents the classification credibility of the original image, R finds the maximum credibility index among them, r represents the classification credibility of each feature map of the first feature in the neural network decision of each side branch, f i1 (x) represents the average activation level of the left-side branch neural network when the feature map x undergoes the i-th binary tree branch; f i2 (x) represents the average activation level of the right-side branch neural network when the feature map x undergoes the i-th binary tree branch; i ∈ {1, 2,..., N}; N represents the serial number of the branch neural network of the binary tree structure.
[0062] Figure 4 is a schematic diagram of all decision branches of a sample of the present invention passing through the interpretable module. Taking the ten-classification task on the Cifar-10 dataset as an example, the class numbers range from 0 to 9. Feature Figure 2The binary tree branches need to be divided into three layers in total. By the third time, there are already nodes that only responsible for the decoupling of one category, so no further branching can be carried out; each convolutional block contains multiple convolutional layers. In the first binary tree branch, nodes 2 and 3 are responsible for decoupling the information of the first five categories (categories 0-4) and the last five categories (categories 5-9) respectively; in the second layer of the binary tree branch, nodes 4 and 5 are responsible for decoupling the information of the first three categories (categories 0-2) and the last two categories (categories 3-4) in categories 0-4, and nodes 6 and 7 follow the same pattern; in the third layer of the binary tree branch, nodes 8 and 9 are responsible for decoupling the information of the first two categories (categories 0-1) and the last one category (category 2) in categories 0-2, and nodes 10 and 11 are responsible for decoupling category 3 and category 4 in categories 3-4, and subsequent nodes follow the same pattern.
[0063] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium can include: ROM, RAM, magnetic disk or optical disk, etc.
[0064] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An interpretable convolutional neural network image classification method based on binary tree structure embedding, characterized in that, Including: Obtain the original image and preprocess the original image; Input the preprocessed original image into a convolutional neural network embedded with an interpretable module having a binary tree structure that has been trained; Extract first features using the first convolutional layer of the convolutional neural network model; Process the first features using the interpretable module of the binary tree structure of the convolutional neural network model to extract second features and neuron activation values of each binary tree branch, specifically including: inputting each feature map of the first features into the interpretable module of the binary tree structure, decoupling the information of the first half of the categories by one side branch neural network; decoupling the information of the second half of the categories by the other side branch neural network; until the last layer of branch neural network decouples the feature information of a single category, fusing the feature information of each single category to obtain second features; and calculating the average activation level of each feature map on each side of each layer of branch neural network; The structure of the convolutional neural network embedded with the interpretable module having a binary tree structure includes selecting a convolutional block in the classification network as the embedding entry, using the feature information extracted by the previous convolutional block as the input of this embedding module and as the output of the next convolutional block, and the embedding module includes multiple branches of the feature map binary tree, and the number of branches of the feature map binary tree is determined by the number of classification categories in the classification task; Process the first features using the second convolutional layer of the convolutional neural network model to extract third features; Fuse and process the second features and the third features using the third convolutional layer of the convolutional neural network model to extract fourth features; Perform classification processing on the fourth features using the fully connected layer of the convolutional neural network model to obtain the classification result of the original image; Obtain the classification decision path and classification credibility of the original image using the neuron activation values of each binary tree branch, specifically including: determining the classification decision path of the original image according to the relative magnitudes of the average activation levels of each feature map of the first features on each side branch neural network; combining the ratio of the average activation levels and the classification decision path, calculating the classification credibility of the decision of each side branch neural network, and comprehensively evaluating the classification credibility of the decisions of all branch neural networks to obtain the classification credibility of the original image.
2. The interpretable convolutional neural network image classification method based on binary tree structure embedding according to claim 1, characterized in that The calculation formula for the classification credibility of the original image is expressed as: Among them, R represents the classification credibility of the original image, r represents the classification credibility of each feature map of the first feature in the branch neural network decision on each side, and f i1 (x) represents the average activation level of the left branch neural network when the feature graph x branches at the i-th binary tree; f i2 (x) represents the average activation level of the right branch neural network when the feature graph x branches for the i-th binary tree; i∈{1,2,…,N}; N represents the number of branch neural networks.
3. A method for classifying images of an interpretable convolutional neural network embedded with a binary tree structure according to claim 1, wherein the training process of the convolutional neural network with an interpretable module embedded with a binary tree structure includes inputting training images into the convolutional neural network with an interpretable module embedded with a binary tree structure for pre-training. By minimizing the total loss function, at the end of pre-training, the neuron activation values of each binary tree branch and the classification decision path of the training images are obtained; the training images are input into the convolutional neural network with an interpretable module embedded with a binary tree structure after being superimposed with perturbations for re-training, and at the end of re-training, the neuron activation values of each binary tree branch and the classification decision path of the training images are obtained; by comparing the corresponding neuron activation values and classification decision paths in the pre-training and re-training processes, when the comparison value is less than a preset threshold, the convolutional neural network with an interpretable module embedded with a binary tree structure after stabilization is output; otherwise, continue to use the training images for re-pre-training.
4. A method for classifying images of an interpretable convolutional neural network embedded with a binary tree structure according to claim 1 or 3, wherein the total loss function adopted in the training process of the convolutional neural network with an interpretable module embedded with a binary tree structure is expressed as: Among them, $L_0$ represents the classification loss function; $\lambda$ represents the balance factor, which is used to balance the classification performance and interpretability performance of the network; $J$ i1 $(\theta)$ represents the loss value of the left - hand branch neural network during the $i$-th binary tree branch, and $J$ i2 $(\theta)$ represents the loss value of the right - hand branch neural network during the $i$-th binary tree branch; $N$ represents the number of branch neural networks.
5. A method for classifying images of an interpretable convolutional neural network embedded with a binary tree structure according to claim 4, wherein the loss values of the two-side branch neural networks are respectively expressed as: Among them, f i1 f(x) represents the average activation level of the left - hand branch neural network during the i - th binary tree branch for the feature map x; f i2 f(x) represents the average activation level of the right - hand branch neural network during the i - th binary tree branch for the feature map x; δ represents a constant greater than 0.