An interpretability method and system for deep learning models of multi-source image data

By normalizing the multi-source image data and introducing noise interference, the deep learning model is retrained to evaluate the contribution of single-source image indicators, and the problem of lack of interpretability of the deep learning model of multi-source image data is solved, and the reliable and reliable interpretability analysis of the model is achieved.

CN117934450BActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410281983.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2025-05-16
Estimated Expiration
2044-03-13

AI Technical Summary

Technical Problem

Existing deep learning models lack interpretability when processing multi-source image data, making model results difficult to trust and understand.

Method used

By normalizing the multi-source image data and introducing random noise interference into the deep learning model, the model is retrained to evaluate the contribution of single-source image metrics to the multi-source image depth model.

Benefits of technology

The interpretability analysis of the deep learning model of multi-source image data is realized, and the contribution of single-source indicators to multi-source depth models is clarified, which improves the reliability and credibility of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117934450B_ABST
    Figure CN117934450B_ABST
Patent Text Reader

Abstract

The present invention proposes an interpretability method and system for a multi-source image data deep learning model. The present invention calculates the response degree of the degradation of the single-source image index to the image data deep learning model by scrambling the single-source image index, thereby obtaining the contribution of each single-source image index to the multi-source image data deep model, and then clarifying the importance of the single-source image index in the multi-source image deep model. And according to the actual meaning of the corresponding single-source image index, the interpretability result of the influence of the corresponding single-source image index on the multi-source image deep model is output. The principle of the interpretability method of the multi-source image data deep learning model of the present invention is reliable, and it can deeply analyze the "black box" problem of the multi-source image data deep model and clarify the contribution of the single-source index to the multi-source deep model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data models, and in particular to an interpretability method and system for a deep learning model of multi-source image data. Background Art

[0002] As an important model, deep neural networks have taken over the mainstream in the field of machine learning, but the inherent black box characteristics of neural networks have affected people's trust in them. Just like the human brain thinking and acting, we actually don't know how the human brain works, but we still believe in human decisions. Therefore, although the deep learning model of artificial intelligence is an end-to-end "black box" model, if its interpretability is studied in depth, the model results can also be trusted by humans. The interpretability method of deep learning strives to understand the operating mechanism of neural networks from multiple perspectives. How to uniformly define model interpretation is still an open question, but there are some recognized classification standards. In general, the interpretability method of the model can be divided from three dimensions: active interpretation or post-interpretation, local or global interpretation, and the object to be explained. Post-interpretation is to explain the trained neural network model, and active interpretation is to provide interpretability information by actively changing the network structure or training process during the training of the neural network. Local explanation only explains a single input sample at a time, while global explanation hopes to explain the decision process of all possible input samples in the feature space (“Satellite-derived bottom depth for optically shallow waters based on hydrolight simulations”, Wang YX, He XQ, Bai Y, et al., Remote Sensing, Vol. 14, 2022, p. 4590). There are four types of explanations: sample-based explanations hope to find representative samples in the data set to show the model's preference; input-based explanations hope to attribute the model's prediction results to different input features, and will mask or perturb the original input features to measure the impact of each feature on the output, and then obtain the contribution of each feature to the output ("Visualizing and understanding convolutional networks", Zeiler, MD, Fergus, R, In Computer Vision–ECCV, 2014, pp. 818-833); explaining hidden semantics to study specific hidden layers or neurons in neural networks; rule-forming explanation methods hope to extract a set of logical rules or build an approximate proxy model of a decision tree to simulate the decision-making process of the neural network. The basic principle of the perturbation-based model interpretability method is to evaluate all input features at once by constructing many perturbation samples, and finally rank the importance of the input features.Existing local interpretability methods for deep learning models, such as occlusion experiments ("Visualizing and understanding convolutional networks", Zeiler, MD, Fergus, R, In Computer Vision–ECCV, 2014, pp. 818-833), "Saliency Maps" ("Deep inside convolutional networks: Visualizing image classification models and Saliencymaps", Simonyan K., Vedaldi A, Zisserman A ICIR 2014, pp. 1-8; "Explainable deeplearning for insights in El Niño and river flows", Liu YM, Duffy K, Dy JG, Nature Communications, 2023, Vol. 14, p. 339) are calculated by calculating the contribution of input samples to the model. In simple terms, by continuously masking pixels in one area after another, such as a dog's ear. By masking such small areas one by one, the degree of influence on the accuracy after masking is determined, or disturbance is added to a certain part. If the impact is greater, the feature or area is more important. Figure 1 As shown in Figure 1, interference is added to the input single-source RGB image and its significance result is calculated. Figure 1 The black area in (a) is the dog’s face. Figure 1 (b) The wheels of the car are also black key areas. Figure 1 In (c), the man’s face is covered, making it difficult for the model to recognize him. We identify a specific feature in the image as the model’s focus area, and then explain the specific classification goal of the model.

[0003] Existing interpretability methods, such as LIME, SHAP, CAM and other methods ("Why Should I Trust You?": explaining the predictions of any classifier", Marco Tulio Ribeiro, Sameer Singh, Carlos Guestrin, https: / / doi.org / 10.48550 / arXiv.1602.04938; "Consistentindividualized feature attribution for tree ensembles", Scott M. Lundberg, Gabriel G. Erion, Su-In Lee, arXiv preprint arXiv:1802.03888, 2018; "Grad-CAM: visual explanations from deep networks via gradient-based localization"). At present, most algorithms are interpretable analysis for single-source image data, and there are no publicly published interpretability methods for deep learning models of multi-source image data. Multi-source data is also commonly used in image processing, computer vision and other fields.

[0004] Therefore, there is an urgent need for an interpretability method for multi-source data deep learning models that can explain multi-source data deep models and make the model results more reliable and credible. Summary of the invention

[0005] Due to the "black box" problem of deep learning models of multi-source image data, people have doubts about the results given by deep learning models. In this regard, the present invention proposes an interpretability method and system for deep learning models of multi-source image data. The present invention can be used to solve the computational problem of local interpretability of deep models of multi-source image data.

[0006] The present invention proposes an interpretability method and system for a multi-source image data deep learning model, comprising:

[0007] Step S1: normalize the multi-source image data to make the order of magnitude of the multi-source image data uniform, and input the multi-source image data and the corresponding target label data as input data into the initial multi-source image deep learning model;

[0008] Step S2: The initial multi-source image deep learning model is learned and trained to obtain a multi-source image deep model M0, and the accuracy A0 of the model M0 is recorded;

[0009] Step S3: For each single-source image indicator y n Add random noise interference e to get the single source image index y n ', and then use the single source image index y n 'Replace y in the initial multi-source image deep learning model n , retrain the initial multi-source image depth model to obtain the multi-source image depth model M n , record the model M n The accuracy of A n , calculate the single-source image index for the multi-source image depth model M n The degree of accuracy degradation |A0– A n |, where the subscript n is a natural number, and the value of n is equal to the number of single-source image indicators;

[0010] Step S4: Determine the contribution of the corresponding single-source image indicator to the multi-source image depth model based on the degree of influence of each single-source image indicator on the decrease in the accuracy of the multi-source image depth model, and output an interpretable result of the influence of the corresponding single-source image indicator on the multi-source image depth model according to the actual meaning of the corresponding single-source image indicator.

[0011] Furthermore, in step S4, the specific method for determining the contribution degree is:

[0012] Normalize the accuracy drop of each single-source image index on the final result, and calculate the influence of the corresponding single-source image index on the multi-source image depth model ,

[0013] ,

[0014] Among them, W0, W1, W2...W n , where i is the subscript number, i is a natural number, and n is the total number;

[0015] Will affect the value The contribution of the corresponding single-source image in the multi-source image deep model.

[0016] Furthermore, the initial multi-source image deep learning model includes a U-net deep network model.

[0017] The present invention also proposes an interpretability system for a multi-source image data deep learning model, comprising:

[0018] The preprocessing module normalizes the multi-source image data to make the multi-source image data have a uniform order of magnitude, and inputs the multi-source image data and the corresponding target label data as input data into the initial multi-source image deep learning model;

[0019] The initial accuracy calculation module learns and trains the initial multi-source image deep learning model to obtain the multi-source image deep model M0 and record the accuracy A0 of the model M0;

[0020] The impact degree calculation module is used to calculate the index y of each single source image. n Add random noise interference e to get the single source image index y n ', and then use the single source image index y n 'Replace y in the initial multi-source image deep learning model n , retrain the initial multi-source image depth model to obtain the multi-source image depth model M n , record the model M n The accuracy of A n , calculate the single-source image index for the multi-source image depth model M n The degree of accuracy degradation |A0– A n |, where the subscript n is a natural number, and the value of n is equal to the number of single-source image indicators;

[0021] The result output module determines the contribution of the corresponding single-source image indicator in the multi-source image depth model based on the degree of influence of each single-source image indicator on the decrease in the accuracy of the multi-source image depth model, and outputs the interpretable results of the influence of the corresponding single-source image indicator on the multi-source image depth model according to the actual meaning of the corresponding single-source image indicator.

[0022] Furthermore, the specific method for determining the contribution degree of the result output module is:

[0023] Normalize the accuracy drop of each single-source image index on the final result, and calculate the influence of the corresponding single-source image index on the multi-source image depth model ,

[0024] ,

[0025] Among them, W0, W1, W2...W n , where i is the subscript number, i is a natural number, and n is the total number;

[0026] Will affect the value The contribution of the corresponding single-source image in the multi-source image deep model.

[0027] Furthermore, the initial multi-source image deep learning model includes a U-net deep network model.

[0028] The present invention is based on the idea of ​​removing single-source indicators and adding disturbances to the multi-source overall model, and performs interpretable analysis on the multi-source image data deep learning model to obtain the contribution of each single-source indicator to the multi-source data deep model, thereby clarifying the importance of the single-source indicator in the multi-source deep model. The principle of the interpretability method of the multi-source image data deep learning model of the present invention is reliable, and can deeply analyze the "black box" problem of the multi-source image data deep model, clarify the contribution of the single-source indicator to the multi-source deep model, and then from the "Machine Learning" of the multi-source data deep model to the "Machine Teaching" machine to guide human work in turn, improve the interpretability of the multi-source image data deep model, and make it more reliable and credible. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram of occlusion-based saliency calculation; where: Figure 1 (a) is the animal face area occlusion map, Figure 1 (b) is the occlusion map of the key area of ​​the object, Figure 1 (c) is the occlusion map of the key areas of the face;

[0030] Figure 2 A technical flow chart of the interpretability method of the multi-source image data deep learning model of the present invention;

[0031] Figure 3 Schematic diagram of deep model environment evaluation based on multi-source image data;

[0032] Figure 4 This is the result of interpretability analysis of multi-source big data deep model; Figure 4 (a) is the deep model environmental prediction and assessment results of the study area, and Figure 4 (b) is a schematic diagram of the contribution of each feature index layer of the multi-source deep model. DETAILED DESCRIPTION

[0033] The present invention proposes an interpretability method and system for a deep learning model of multi-source image data. The specific implementation methods of the present invention are described in detail below.

[0034] The basic principle of the present invention is as follows. The present invention utilizes the data-driven results of the multi-source image data deep learning model to reversely infer the importance of the model indicator features, and combines the original attributes of the indicators to analyze the interpretability of the multi-source image data deep model. Perturbations are performed by using one or a group of input feature indicator layer (single-source image data) deviations as input to obtain the difference between the perturbed output and the original output, obtain different degrees of influence on the results, and obtain the contribution of the single-source indicator feature layer (single-source image data) to the multi-source data deep model, or importance. The method based on adding disturbances can directly estimate the importance of single-source image feature indicators, and the operation is simple and has strong versatility. The multi-source image data targeted by the present invention is slightly different from ordinary single-source image data, and is a multi-dimensional, multi-channel image data. Differences in data dimensions, structures, etc. make multi-source data deep learning more complicated. The idea of ​​obtaining the influence of an indicator feature on the entire model is as follows. Multi-source indicator feature data includes:

[0035] {y1,y2,y3,……,y n} (1)

[0036] Remove one of the indicator features or add disturbance , re-enter the trained model, calculate the accuracy of the inversion results of the indicator characteristic data, and increase the error Here y is an index layer (single-source image index data), in which the specific multidimensional data pixel contains [x1, x2, x3, ......, x n ], corresponding to different target classification labels.

[0037] Assuming the change of each single-source indicator is ∆y, and the error change of each indicator to the overall target result is ∆e, the importance of the indicator feature can be expressed as:

[0038] (2)

[0039] By calculating the S value of each indicator, we can obtain the importance of each indicator in the model, and combine it with the meaning of the indicator feature itself, so that we can use the multi-source image data intelligent model in an interpretable way, and the model is reliable and credible.

[0040] The following uses an image data model as an example to illustrate the detailed implementation of the analysis method of the present invention. Figure 2 Technical flowchart of our approach to interpretability of deep learning models for multi-source image data.

[0041] Step S1: Normalize the multi-source image data to unify the order of magnitude of the multi-source image data, and input the multi-source image data and the corresponding target label data as input data into the initial multi-source image deep learning model.

[0042] Step S2: The initial multi-source image deep learning model is learned and trained to obtain a multi-source image deep model M0, and the accuracy A0 of the model M0 is recorded.

[0043] Step S3: For each single-source image indicator y n Add random noise interference e to get the single source image index y n ', and then use the single source image index y n 'Replace y in the initial multi-source image deep learning model n , retrain the initial multi-source image depth model to obtain the multi-source image depth model M n , record the model M n The accuracy of A n , calculate the single-source image index for the multi-source image depth model M n The degree of accuracy degradation |A0– A n |, where the subscript n is a natural number, and the value of n is equal to the number of single-source image indicators.

[0044] Specifically for a single-source image indicator y1, add random noise interference e, and then replace y1 in the original model with the single-source image indicator y1' with the added random noise, retrain, obtain the multi-source image depth model M1, and record the accuracy of the model A1. At this time, the result accuracy decreases |A0- A1|.

[0045] For different single-source indicators When , the accuracy A of the record model corresponding to different single-source indicators is repeatedly calculated in the above manner n , and obtain the corresponding precision reduction results |A0- A2|, |A0- A3|, |A0- A4|, ..., |A0- A n |.

[0046] Step S4: Determine the contribution of the corresponding single-source image indicator to the multi-source image depth model based on the degree of influence of each single-source image indicator on the decrease in the accuracy of the multi-source image depth model, and output an interpretable result of the influence of the corresponding single-source image indicator on the multi-source image depth model according to the actual meaning of the corresponding single-source image indicator.

[0047] Specifically, according to the degree of decrease in the accuracy of each single-source indicator on the multi-source image depth model, the single-source indicator on the final result is normalized to obtain the single-source indicator on the multi-source depth model. , where i is the subscript number, i is a natural number, n is the total number, W0, W1, W2...W nIt reflects the importance of each single-source indicator in the multi-source image depth model or its contribution to the model, and conducts interpretability analysis in conjunction with the actual meaning of the single-source indicator characteristics themselves, so that the contribution of the single-source indicator to the multi-source depth model is consistent with the impact of the actual single-source indicator on the final target task. Finally, an interpretability result of the multi-source image depth model that is both highly accurate and credible and understandable is obtained, and the model results and the contribution of each single-source indicator and the corresponding indicator interpretation analysis results are output.

[0048] The present invention is not limited to evaluating the interpretability of deep learning models of multi-source image data, but can also be applied to various learning models with multi-source data. For example, an environmental evaluation model contains multiple environmental indicator data, that is, it constitutes multiple influencing factors as multi-source data. Taking Xi'an as an example, the environmental evaluation work of the research area is Figure 3 The multi-source image data and environmental assessment label data are input into a U-net deep network model (the present invention is not limited to this model, and can also be applied to other deep learning models, such as the Transformer model and the DEtection TRansformer (DETR) model); learning, training and interpretability analysis of the multi-source big data deep model are performed, and finally the deep model environmental prediction and assessment results of the study area and the contribution of each characteristic indicator layer of the multi-source deep model are obtained, such as Figure 4 So far, the multi-source image data deep model has been used to achieve high-precision environmental evaluation in Xi'an, and the model is interpretable, breaking through the "black box" of the multi-source deep model, clarifying the contribution of each single source indicator to the overall multi-source model, and achieving "knowing why" and interpretable high-precision environmental evaluation.

[0049] The present invention also proposes an interpretability system for a multi-source image data deep learning model, comprising:

[0050] The preprocessing module normalizes the multi-source image data to unify the order of magnitude of the multi-source image data, and inputs the multi-source image data and the corresponding target label data as input data into the initial multi-source image deep learning model.

[0051] The initial accuracy calculation module learns and trains the initial multi-source image deep learning model to obtain the multi-source image deep model M0 and record the accuracy A0 of the model M0.

[0052] The impact degree calculation module is used to calculate the index y of each single source image. n Add random noise interference e to get the single source image index y n ', and then use the single source image index y n 'Replace y in the initial multi-source image deep learning model n, retrain the initial multi-source image depth model to obtain the multi-source image depth model M n , record the model M n The accuracy of A n , calculate the single-source image index for the multi-source image depth model M n The degree of accuracy degradation |A0– A n |, where the subscript n is a natural number, and the value of n is equal to the number of single-source image indicators.

[0053] The result output module determines the contribution of the corresponding single-source image indicator in the multi-source image depth model based on the degree of influence of each single-source image indicator on the decrease in the accuracy of the multi-source image depth model, and outputs the interpretable results of the influence of the corresponding single-source image indicator on the multi-source image depth model according to the actual meaning of the corresponding single-source image indicator.

[0054] The specific method for determining the contribution degree of the result output module is as follows:

[0055] Normalize the accuracy drop of each single-source image index on the final result, and calculate the influence of the corresponding single-source image index on the multi-source image depth model ,

[0056] ,

[0057] Among them, W0, W1, W2...W n , where i is the subscript number, i is a natural number,

[0058] Will affect the value The contribution of the corresponding single-source image in the multi-source image deep model.

[0059] The principle of the interpretability method of the multi-source image data deep learning model described in the present invention is reliable, and it can deeply analyze the "black box" problem of the multi-source image data deep model, clarify the contribution of single-source indicators to the multi-source deep model, and then from the "Machine Learning" to "Machine Teaching" of the multi-source data deep model, the machine in turn guides human work, thereby improving the interpretability of the multi-source image data deep model and making it more reliable and credible.

Claims

1. A method for interpretability of deep learning models of multi-source image data, characterized in that: include: Step S1: normalize the multi-source image data, and use the multi-source image data and the corresponding target label data as input data and input them into the initial multi-source image deep learning model; Multi-source image data is multi-dimensional and multi-channel image data; Step S2: The initial multi-source image deep learning model is learned and trained to obtain a multi-source image deep model M0, and the accuracy A0 of the model M0 is recorded; Step S3: For each single-source image indicator y n Add random noise interference e to get the single source image index y n ', and then use the single source image index y n 'Replace y in the initial multi-source image deep learning model n , retrain the initial multi-source image depth model to obtain the multi-source image depth model M n , record the model M n The accuracy of A n , calculate the single-source image index for the multi-source image depth model M n The degree of accuracy degradation of |A0–A n |, where the subscript n is a natural number, and the value of n is equal to the number of single-source image indicators; Step S4: Determine the contribution of the corresponding single-source image indicator to the multi-source image depth model based on the degree of influence of each single-source image indicator on the decrease in the accuracy of the multi-source image depth model, and output an interpretable result of the influence of the corresponding single-source image indicator on the multi-source image depth model according to the actual meaning of the corresponding single-source image indicator.

2. The method according to claim 1, characterized in that: In step S4, the specific method for determining the contribution degree is: Normalize the accuracy drop of each single-source image index on the final result, and calculate the influence value W of the corresponding single-source image index on the multi-source image depth model i , Among them, W0, W1, W2...W n , where i is the subscript number, i is a natural number, and n is the total number; The impact value W i The contribution of the corresponding single-source image in the multi-source image deep model.

3. The method according to claim 1 or 2, characterized in that: The initial multi-source image deep learning model includes a U-net deep network model.

4. An interpretability system for a deep learning model of multi-source image data, characterized in that: include: The preprocessing module normalizes the multi-source image data to make the multi-source image data have a uniform order of magnitude, and inputs the multi-source image data and the corresponding target label data as input data into the initial multi-source image deep learning model; Multi-source image data is multi-dimensional and multi-channel image data; The initial accuracy calculation module learns and trains the initial multi-source image deep learning model to obtain the multi-source image deep model M0 and record the accuracy A0 of the model M0; The impact degree calculation module is used to calculate the index y of each single source image. n Add random noise interference e to get the single source image index y n ', and then use the single source image index y n 'Replace y in the initial multi-source image deep learning model n , retrain the initial multi-source image depth model to obtain the multi-source image depth model M n , record the model M n The accuracy of A n , calculate the single-source image index for the multi-source image depth model M n The degree of accuracy degradation of |A0–A n |, where the subscript n is a natural number, and the value of n is equal to the number of single-source image indicators; The result output module determines the contribution of the corresponding single-source image indicator in the multi-source image depth model based on the degree of influence of each single-source image indicator on the decrease in the accuracy of the multi-source image depth model, and outputs the interpretable results of the influence of the corresponding single-source image indicator on the multi-source image depth model according to the actual meaning of the corresponding single-source image indicator.

5. The system according to claim 4, characterized in that The specific method for the result output module to determine the contribution degree is: Normalize the accuracy drop of each single-source image index on the final result, and calculate the influence value W of the corresponding single-source image index on the multi-source image depth model i , Among them, W0, W1, W2...W n , where i is the subscript number, i is a natural number, and n is the total number; The impact value W i The contribution of the corresponding single-source image in the multi-source image deep model.

6. The system according to claim 4 or 5, characterized in that: The initial multi-source image deep learning model includes a U-net deep network model.

Citation Information

Patent Citations

  • Deep learning model interpretation method based on task cognition

    CN115017336A