Processing Method, Apparatus, Device, and Storage Medium for Deep Learning Model

By obtaining the feature parameters of the deep learning model and building an interpreted model, the problem that deep learning model is difficult to explain is solved, the quantitative interpretation and analysis of the model is realized, and the application effect of the model in specific fields is improved.

CN112580781BActive Publication Date: 2025-06-27WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011469107.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-14
Publication Date
2025-06-27
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

The existing technology is difficult to quantitatively interpret and analyze deep learning models, which makes users unable to understand the working principle of the model, limiting its application in areas such as driverless driving and medical image recognition.

Method used

By inputting the data set into the deep learning model, the information gain, sparseness parameters and complete parameters of the feature are obtained, and the tree model is built for training, the interpreted model is obtained, the classification accuracy and complete parameters of its leaf nodes are measured, and these indicators are output to provide quantitative explanation.

Benefits of technology

Quantitative interpretation and analysis of deep learning models is realized, and visual results of model performance evaluation are provided to help users understand the working principle of the model and improve their application effectiveness in specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112580781B_ABST
    Figure CN112580781B_ABST
Patent Text Reader

Abstract

The present invention discloses a processing method, device, equipment and storage medium for a deep learning model. The method includes: for a deep learning model, calculating the information gain, sparsity parameter and completeness parameter of each feature of the deep learning model according to the feature outputs of some intermediate layers and the final output extracted from the feature extractor and the classifier. And training a tree model based on the output of the intermediate layer to obtain an interpretation model, testing the leaf nodes of the interpretation model to obtain the classification accuracy of the interpretation model, and calculating the ratio of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples to obtain the tree completeness parameter. Finally, outputting the information gain, sparsity parameter, completeness parameter, tree accuracy and tree completeness of the obtained features, which are indicators for evaluating the model, so as to provide a tool for quantitative analysis and interpretation of the deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method, apparatus, device, and storage medium for processing a deep learning model. Background Art

[0002] A Convolution Neural Networks (CNNs) is a deep learning model that has excellent performance in fields such as image recognition and has been widely used. CNNs mainly consists of a convolution part and a fully connected part. The convolution part includes a convolution layer, an activation function layer, a pooling layer, etc., and its function is to extract features of data; the function of the fully connected part is to connect features and output to calculate losses, and perform operations such as recognition and classification.

[0003] However, due to the end-to-end learning strategy and extremely complex model parameter structure of the deep learning model, CNNs has always been as difficult to understand and explain its working principle as a black box. After CNNs is trained and converges, users can only obtain the final output results of the model (such as the category to which the input belongs, etc.) during use, but cannot understand how CNNs obtains the predicted output from the original input. This lack of interpretability has greatly hindered the implementation of current deep learning models such as CNNs in fields such as unmanned driving and medical image recognition.

[0004] In summary, there is currently no suitable tool for quantitatively interpreting and analyzing deep learning models. Summary of the Invention

[0005] The main object of the present invention is to provide a method, apparatus, device, and storage medium for processing a deep learning model, and to provide a tool for quantitatively interpreting and analyzing a deep learning model.

[0006] To achieve the above object, the present invention provides a method for processing a deep learning model, including:

[0007] Input a pre-acquired data set into a deep learning model to be processed, and obtain the information gain, sparsity parameter, and completeness parameter of each feature of the deep learning model; wherein, the data set includes data of multiple features, the information gain of each feature is used to represent the ability of the feature to distinguish data samples, the sparsity parameter of each feature is used to represent the degree of independence between features, and the completeness parameter of each feature is used to represent the influence degree of the feature on the deep learning model;

[0008] Extract the output of the feature extractor and the output of the classifier from the deep learning model, and use the output of the feature extractor and the output of the classifier as training data for tree model training to obtain an interpretation model;

[0009] Measure the classification accuracy of the leaf nodes of the interpretation model to obtain a tree accuracy, which is used to indicate the classification accuracy of the interpretation model;

[0010] Calculate the ratio of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples to obtain a tree completeness parameter;

[0011] Output the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter.

[0012] In a specific implementation, the method further includes:

[0013] Perform visualization processing according to the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter to obtain a visualization result for evaluating the deep learning model;

[0014] Output the visualization result.

[0015] In a specific implementation, the inputting the pre-acquired data set into the deep learning model to be processed to obtain the information gain, sparsity parameter, and completeness parameter of each feature by the deep learning model includes:

[0016] Input the data set into the deep learning model, extract the output of the feature extractor of the deep learning model, filter the output value of each feature in the output of the feature extractor, and take the mean value after filtering to obtain the information gain of the deep learning model for the feature;

[0017] Extract all filter matrices from the convolutional layer of the deep learning model, perform conversions according to all the filter matrices, calculate the K-L divergence matrix of each feature pairwise, and obtain the sparsity parameter corresponding to the feature according to the K-L divergence matrix of each feature;

[0018] Delete a feature set from the data set in order from largest to smallest according to the information gain of each feature by the deep learning model, and construct a random forest model according to all the remaining feature sets after each deletion, and calculate the test performance of the random forest model;

[0019] When the change in the test performance of a random forest model compared to the test performance of the previous model is greater than a preset value, obtain the number of deleted feature sets;

[0020] Calculate and obtain the completeness parameter according to the number of the deleted feature sets and the total number of the feature sets in the data set.

[0021] In a specific embodiment, the measurement of the classification accuracy of the leaf nodes of the interpretation model to obtain the tree accuracy includes:

[0022] Measure and obtain the total number of samples that the classification of the interpretation model finally falls on each leaf node and the number of samples correctly classified by each leaf node;

[0023] Adopt the formula: Calculate the classification accuracy Acc of each leaf node in the interpretation model i , where i is the leaf node serial number, n i is the total number of samples that finally fall on this leaf node after being classified by the interpretation model, and c i is the number of samples correctly classified by the leaf node. The tree accuracy includes the classification accuracy of each leaf node.

[0024] In another specific embodiment, the calculation of the ratio of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category in all samples to obtain the tree completeness parameter includes:

[0025] Adopt the formula: Calculate the ratio Comp of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category in all samples i , and obtain the tree completeness parameter; where i is the leaf node serial number, and c i is the number of samples correctly classified by this leaf node, and n c is the number of samples of the same category as this node in all samples.

[0026] The present invention also provides a processing device for a deep learning model, including:

[0027] A first processing module, configured to input a pre-acquired data set into a deep learning model to be processed, and obtain the information gain, sparsity parameter, and completeness parameter of each feature of the deep learning model; wherein, the data set includes data of multiple features, the information gain of each feature is used to represent the ability of the feature to distinguish data samples, the sparsity parameter of each feature is used to represent the independence degree between features, and the completeness parameter of each feature is used to represent the influence degree of the feature on the deep learning model;

[0028] A second processing module, configured to extract the output of the feature extractor and the output of the classifier from the deep learning model, and use the output of the feature extractor and the output of the classifier as training data for tree model training to obtain an interpretation model;

[0029] The second processing module is further configured to measure the classification accuracy of the leaf nodes of the interpretation model to obtain a tree accuracy, where the tree accuracy is used to indicate the classification accuracy of the interpretation model;

[0030] A third processing module, configured to calculate the ratio of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples, to obtain a tree completeness parameter;

[0031] An output module, configured to output the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter.

[0032] In a specific embodiment, the apparatus further includes:

[0033] A fourth processing module, configured to perform visualization processing according to the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter, to obtain a visualization result for evaluating the deep learning model;

[0034] The output module is further configured to output the visualization result.

[0035] In a specific embodiment, the first processing module is specifically configured to:

[0036] Input the data set into the deep learning model, extract the output of the feature extractor of the deep learning model, filter the output values of each feature in the output of the feature extractor, and take the mean value after filtering to obtain the information gain of the deep learning model for the feature;

[0037] Extract all filter matrices from the convolutional layer of the deep learning model, perform conversions according to the all filter matrices respectively, calculate the K-L divergence matrix of each feature pairwise, and obtain the sparsity parameter corresponding to the feature according to the K-L divergence matrix of each feature;

[0038] Delete a feature set from the data set in order from largest to smallest according to the information gain of each feature of the deep learning model, and construct a random forest model according to all the feature sets that have not been deleted after each deletion, and calculate the test performance of the random forest model;

[0039] When the change in the test performance of a random forest model compared to the test performance of the previous model is greater than a preset value, obtain the number of deleted feature sets;

[0040] Calculate and obtain the completeness parameter according to the number of the deleted feature sets and the total number of feature sets in the dataset.

[0041] In a specific implementation manner, the second processing module is specifically configured to:

[0042] Measure and obtain the total number of samples that the classification of the interpretation model finally falls on each leaf node and the number of samples with correct classification on each leaf node;

[0043] Adopt the formula: Calculate the classification accuracy Acc of each leaf node in the interpretation model i , where i is the leaf node serial number, and n i is the total number of samples that finally fall on this leaf node after being classified by the interpretation model, and c i is the number of samples with correct classification on the leaf node. The tree accuracy includes the classification accuracy of each leaf node.

[0044] In a specific implementation manner, the third processing module is specifically configured to:

[0045] Adopt the formula: Calculate the ratio Comp of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples, i to obtain the tree completeness parameter; where i is the leaf node serial number, and c i is the number of samples with correct classification on this leaf node, and n c is the number of samples of the same category as this node in all samples.

[0046] The present invention also provides an electronic device, which includes: a memory, a processor, and an output interface. A computer program that can run on the processor is stored on the memory. When the computer program is executed by the processor, the steps of the foregoing processing method of the deep learning model are implemented.

[0047] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing processing method of the deep learning model are implemented.

[0048] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the processing method of the deep learning model described in any one of the foregoing are implemented.

[0049] In the present invention, for a deep learning model, intermediate layer feature outputs and final outputs are extracted from a feature extractor and a classifier according to the model structure. Based on these outputs, the information gain, sparsity parameter, and completeness parameter of each feature of the deep learning model are calculated. Then, a tree model is trained based on the outputs of the intermediate layer to obtain an interpretation model, and the leaf nodes of the interpretation model are tested to obtain the classification accuracy of the interpretation model. The ratio of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples is calculated to obtain the tree completeness parameter. Finally, the information gain, sparsity parameter, completeness parameter, tree accuracy, and tree completeness of the obtained features, which are indicators for evaluating the model, are output, so as to provide a tool for quantitative analysis and interpretation of the deep learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 FIG. 6 is a schematic flowchart of the first embodiment of the processing method of the deep learning model provided by the present invention;

[0051] Figure 2 FIG. 10 is a schematic diagram of a specific information gain and the number of features provided by the present invention;

[0052] Figure 3 FIG. 14 is a schematic diagram of the corresponding relationship between the test performance of an RF model and the number of features provided by the present invention;

[0053] Figure 4 FIG. 18 is a schematic flowchart of the second embodiment of the processing method of the deep learning model provided by the present invention;

[0054] Figure 5 FIG. 22 is a radar chart provided by the present invention;

[0055] Figure 6 FIG. 26 is a schematic structural diagram of the first embodiment of the processing device of the deep learning model provided by the present invention;

[0056] Figure 7 FIG. 30 is a schematic structural diagram of the second embodiment of the processing device of the deep learning model provided by the present invention;

[0057] Figure 8 FIG. 34 is a schematic structural diagram of the first embodiment of the electronic device provided by the present invention.

[0058] The implementation, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0060] Convolution Neural Networks (CNNs) is a deep learning model that is currently widely used in various technical fields. However, there is currently no suitable technical solution for interpreting and analyzing deep learning models in the prior art, resulting in users being unable to understand their working principles and not being able to understand their specific effects and functions in fields such as autonomous driving and medical image recognition. Therefore, the present invention provides a processing method for deep learning models, analyzes the deep learning models, and better quantitatively explains the deep learning models by outputting the analysis results.

[0061] The overall idea of the technical solution of the present invention is as follows: For deep learning models of the type of the original convolutional neural network model, it can be divided into a feature extractor and a classifier according to the structure, extract the feature output of the intermediate layer of the deep learning model (i.e., the output of the feature extractor) and the final output of the model (i.e., the output of the classifier), and calculate interpretable evaluation metrics, including feature information gain, feature sparsity, and feature completeness. Then, using the aforementioned obtained feature output and final output, construct a tree model with strong interpretability such as a Decision Tree (DT) or a Random Forest (RF), and calculate the classification accuracy of the leaf nodes and the completeness of the leaf nodes. Finally, summarize and process the above results to obtain different metrics for the original convolutional neural network model, and an interpretable visualization report of the original convolutional neural network model can also be further obtained.

[0062] The processing method for deep learning models provided by the present invention can be applied to electronic devices with data processing capabilities such as servers, computers, and intelligent terminals that can perform data analysis or have data operation capabilities. This solution is not limited in this regard.

[0063] The processing method for deep learning models will be specifically described below through several specific embodiments.

[0064] Figure 1 is a schematic flowchart of the first embodiment of the processing method for deep learning models provided by the present invention. As Figure 1 shown, the processing method for deep learning models includes the following steps:

[0065] S101: Input the pre-acquired data set into the deep learning model to be processed, and obtain the information gain, sparsity parameter, and completeness parameter of each feature of the deep learning model.

[0066] In this step, in order to analyze the deep learning model in a specific scenario, different data sets need to be obtained first. The deep learning model can be used to learn and process different data sets respectively to obtain different metrics for the deep learning model to process the data set. The same data set also includes data of various features in this field.

[0067] In this solution, it should be understood that the information gain of each feature is used to represent the ability of the feature to distinguish data samples, the sparsity parameter of each feature is used to represent the degree of independence between features, and the completeness parameter of each feature is used to represent the degree of influence of the feature on the deep learning model.

[0068] Taking any data set as an example below, the process of calculating the above several metrics will be described in detail.

[0069] I. Feature Information Gain

[0070] First, input the pre-acquired data set to be learned in a certain field into the deep learning model to be analyzed and processed, extract the output of the feature extractor of the deep learning model, filter the output value of each feature in the output of the feature extractor, and take the mean value after filtering to obtain the information gain of the deep learning model for the feature.

[0071] Specifically, in this solution, the feature extractor is the middle layer of the deep learning model and outputs different features. The information gain is the difference between the information entropy of the parent and child nodes in the tree model, which can represent the ability of a feature to distinguish data samples. In this solution, the tree model is constructed based on the outputs of the feature extractor and classifier of the deep learning model, and is a model with stronger interpretability than the deep learning model. Calculate the information gain for all features obtained from the output of the feature extractor of the deep learning model (i.e., the original convolutional neural network model).

[0072] The calculation formula of the information gain can be expressed as the difference between the information entropy I(parent) of the parent node and the information entropy I(child) of the child node before and after the partitioning operation:

[0073] ΔInfoGain = I(parent) - I(child)

[0074] Among them, the information entropy I of any node can be expressed as:

[0075]

[0076] Among them, m represents the number of features output by the deep learning model at the layer where the node is located, and pk represents the number of samples corresponding to the K-th feature among all samples. The information entropy I in the above formula specifically refers to filtering all output values of a certain layer of the deep learning model with a series of thresholds (such as taking nine equal parts from 0.1 to 0.9), that is, taking the original value if it is greater than the threshold and taking 0 if it is less than the threshold, taking the average of all filtered results of each output value to obtain the information entropy of a node, and then the information gain value corresponding to a certain feature can be obtained through the difference in information entropy between the child node and the parent node. Figure 2 It is a schematic diagram of a specific information gain and the number of features provided by the present invention. After obtaining the information gain value corresponding to each feature, sort the information gain values of all features from high to low, and the result is as Figure 2 shown. Figure 2 In the figure, the horizontal axis represents the number of features, and the vertical axis represents the information gain.

[0077] Feature information gain is used to measure the influence degree of a feature on the model classification ability. The higher the information gain, the more important this feature is for the classification of the model. That is to say, in the application process of a specific deep learning model, the feature with a higher information gain value is more crucial for the classification result of the model.

[0078] Second, the sparsity parameter of the feature, also known as sparsity (Feature Sparsity)

[0079] In the above process, after inputting the data set into the deep learning model, extract all filtered matrices from the convolutional layer of the deep learning model, perform conversions according to the all filtered matrices respectively, calculate the K-L divergence matrix of each feature pairwise, and obtain the sparsity parameter corresponding to the feature according to the K-L divergence matrix of each feature.

[0080] In the previous convolutional layer of the fully connected layer of the above deep learning model (i.e., the original convolutional neural network model), extract all filter matrices (for example, the second convolutional layer of a certain convolutional neural network model has 16 10*10 feature matrices, representing a total of 1600 features, that is, m in the above information gain calculation process is 1600) and the final output result. After a series of conversion operations, calculate the K-L divergence matrix (Kullback-Leibler divergence matrix) of each feature pairwise.

[0081] Specifically, the calculation formula of the KL divergence can be expressed as:

[0082]

[0083] Among them, P(x) and Q(x) are two probability distributions on the random variable X. Here, the sparsity parameter of the feature includes the KL divergence. Feature sparsity is used to represent the mutual independence between the features extracted by the convolutional layer of the deep learning model.

[0084] III. Completeness parameter of features, also known as Feature Redundancy

[0085] Based on the above scheme, after calculating the information gain of each feature, one feature set is sequentially deleted from the dataset according to the order of the information gain of each feature from large to small by the deep learning model. And after each deletion, a random forest model is constructed according to all the undeleted feature sets, and the test performance of the random forest model is calculated. When the change in the test performance of a random forest model compared to the test performance of the previous model is greater than a preset value, obtain the number of the deleted feature sets; calculate and obtain the completeness parameter according to the number of the deleted feature sets and the total number of feature sets in the dataset.

[0086] Specifically, after calculating the information gain of each feature, all features can be sorted, and the feature sets are sequentially deleted from low to high according to the size of the information gain, that is, one feature set in the dataset is deleted. After each deletion of a feature set, several random forest models with strong interpretability are constructed using the undeleted feature sets. For example, if the total number of features is 400, 400 different RF models can be obtained. Calculate the performance of these RF models on the test dataset, and record the position where the model performance (such as prediction accuracy) undergoes a sharp drop mutation (that is, the change in the test performance of two consecutive RF models is greater than the preset value), as Figure 3 shown Figure 3 is a schematic diagram of the correspondence between the test performance of an RF model provided by the present invention and the number of features. Figure 3 In , the horizontal axis represents the number of features, and the vertical axis represents the performance of the RF model.

[0087] Taking Figure 3 as an example, it can be seen that when about 360 features are deleted, the prediction performance of the RF model begins to drop sharply. Therefore, it can be determined that the last 40 features have a significant impact on the prediction performance of the deep learning model.

[0088] Specifically, the completeness parameter of the feature can be calculated according to the number of features that have a relatively large impact on the overall deep learning model and the total number of features. Taking the above Figure 3Taking the content shown in as an example, the completeness parameter of the deep learning model can be calculated to be 40 / 400=1 / 10, that is, 0.1. Feature completeness indicates the degree of influence of the feature on the overall prediction performance of the deep learning model, and can be used to evaluate the importance of the feature to the model performance.

[0089] S102: Extract the output of the feature extractor and the output of the classifier from the deep learning model, and use the output of the feature extractor and the output of the classifier as training data to train a tree model to obtain an explanation model.

[0090] S103: Measuring the classification accuracy of the leaf nodes of the explanation model to obtain the tree accuracy, where the tree accuracy is used to indicate the classification accuracy of the explanation model.

[0091] In the above steps, the deep learning model may include a feature extractor and a classifier according to the structure. In the technology of the above scheme, after the data set is input into the deep learning model, the feature output of the middle layer of the model (i.e., the feature extractor output) and the final output of the model (i.e., the classifier output) are extracted from the deep learning model (i.e., the original convolutional neural network), and then they can be used as training data pairs to construct a tree model with strong interpretability, such as a decision tree or a random forest. This type of tree model essentially fits the behavior of the original model in an interpretable model manner, and is called an interpretation model of the original model.

[0092] Furthermore, the classification accuracy of the explanation model and the tree completeness parameter can be calculated by measuring the leaf nodes of the explanation model. Next, the tree accuracy is calculated based on the above steps.

[0093] 4. Tree Accuracy

[0094] In order to evaluate the explanation model built based on the intermediate results of the original deep learning model, the classification accuracy of the leaf nodes of the explanation model can be measured.

[0095] Specifically, the total number of samples classified by the explanation model and finally falling on each leaf node and the number of samples correctly classified at each leaf node are measured and obtained;

[0096] Using formula Calculate the classification accuracy Acc of each leaf node in the explanation model i ;

[0097] Among them, i is the leaf node number, n i is the total number of samples classified by the explanation model and finally falling into the leaf node, c i The number of samples correctly classified for the leaf node, the tree accuracy includes the classification accuracy of each leaf node.

[0098] S104: Calculate the ratio of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category among all samples, and obtain the tree completeness parameter.

[0099] On the basis of the above steps, further, after constructing a tree model with strong interpretability based on the deep learning model (i.e., the original convolutional neural network model), the completeness of the leaf nodes can also be calculated, that is, the ratio of the number of samples that each leaf node can correctly classify to the number of samples of this category among all samples, and obtain the tree completeness parameter. The specific calculation method is as follows:

[0100] V. Tree Completeness

[0101] Use the formula: Calculate the ratio Comp of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category among all samples i , and obtain the tree completeness parameter; where i is the leaf node serial number, and c i is the number of samples correctly classified by this leaf node, and n c is the number of samples of the same category as this node among all samples.

[0102] S105: Output the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter.

[0103] In this step, after calculating the information gain, sparsity parameter, completeness parameter, tree accuracy, and tree completeness parameter of each feature through the above process, in order to help the user intuitively understand the deep learning model, these parameter indicators need to be output. The specific output method can be to directly display them on the interface of the electronic device, or to interact with the user's terminal devices such as monitors, computers, and mobile phones, and present them on the terminal devices.

[0104] The processing method of the deep learning model provided in this embodiment, for the deep learning model, calculates the information gain, sparsity parameter, and completeness parameter of each feature of the deep learning model according to the feature outputs and final outputs of some intermediate layers extracted from the feature extractor and the classifier, and then constructs a tree model based on the output of the intermediate layer to obtain an interpretation model. Based on the tests of each node of the interpretation model, the tree accuracy and tree completeness parameter are calculated, and these parameter indicators are output, so as to provide a tool for quantitative analysis and interpretation of the deep learning model to the user.

[0105] Figure 4 It is a schematic flowchart of the second embodiment of the processing method of the deep learning model provided by the present invention, as Figure 4As shown, based on the foregoing embodiments, the processing method of the deep learning model further includes the following steps:

[0106] S106: Perform visualization processing according to the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter to obtain a visualization result for evaluating the deep learning model.

[0107] S107: Output the visualization result.

[0108] In the above steps, in order to help users better understand various metrics of the deep learning model, the calculated metrics for explaining the model can be visualized to obtain a relatively intuitive visualization result. The visualization result can be a visualized pattern, table, or other charts, such as a radar chart, and finally the visualization result is displayed or output through the user's terminal device.

[0109] In a specific example, the above-mentioned evaluation metric dimensions can be visualized to obtain a radar chart for evaluating the explanatory model. Figure 5 A radar chart provided by the present invention is as Figure 5 shown. In this solution, the electronic device processes three convolutional neural network models with different structures (LeNet (represented by a longer dashed line in the figure, the innermost polygon), AlexNet (represented by a solid line in the figure), VGG-16 (represented by a shorter dashed line in the figure, the outermost dashed line)) on the same dataset (CIFAR-10) to obtain a radar chart corresponding to each model, as Figure 5 shown.

[0110] The processing method of the deep learning model provided by the embodiments of the present application provides quantitative interpretable evaluation metrics for the deep learning model, which can be used to objectively compare the performance of different deep learning models. For different deep learning models, a radar chart related to the model performance can also be provided, providing an effective basis and quantitative metrics for further improving the model performance. Overall, it solves the problem that there is no tool for quantitatively interpreting and analyzing deep learning models in the prior art.

[0111] Figure 6 A schematic structural diagram of the first embodiment of the processing device for the deep learning model provided by the present invention is as Figure 6 shown. The processing device 10 for the deep learning model includes:

[0112] The first processing module 11 is configured to input a pre-acquired data set into a deep learning model to be processed, and obtain the information gain, sparsity parameter, and completeness parameter of each feature of the deep learning model; wherein, the data set includes data of multiple features, and the information gain of each feature is used to represent the ability of the feature to distinguish data samples, the sparsity parameter of each feature is used to represent the degree of independence between features, and the completeness parameter of each feature is used to represent the degree of influence of the feature on the deep learning model;

[0113] The second processing module 12 is configured to extract the output of the feature extractor and the output of the classifier from the deep learning model, and use the output of the feature extractor and the output of the classifier as training data for tree model training to obtain an interpretation model;

[0114] The second processing module 12 is further configured to measure the classification accuracy of the leaf nodes of the interpretation model to obtain a tree accuracy, and the tree accuracy is used to indicate the classification accuracy of the interpretation model;

[0115] The third processing module 13 is configured to calculate the ratio of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples to obtain a tree completeness parameter;

[0116] The output module 14 is configured to output the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter.

[0117] The processing device for the deep learning model provided in this embodiment is used to execute the technical solutions provided in any of the foregoing method embodiments, and its implementation principle and technical effects are similar, and will not be described in detail here.

[0118] Figure 7 It is a schematic structural diagram of Embodiment 2 of the processing device for the deep learning model provided by the present invention. As Figure 7 shown, the processing device 10 for the deep learning model includes:

[0119] The fourth processing module 15 is configured to perform visualization processing according to the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter to obtain a visualization result for evaluating the deep learning model;

[0120] The output module 14 is further configured to output the visualization result.

[0121] Based on any of the foregoing embodiments, the first processing module 11 is specifically configured to:

[0122] Input the dataset into the deep learning model, extract the output of the feature extractor of the deep learning model, filter the output values of each feature in the output of the feature extractor, and take the mean after filtering to obtain the information gain of the deep learning model for the feature;

[0123] Extract all filtering matrices from the convolutional layer of the deep learning model, perform conversions according to all the filtering matrices respectively, calculate the K-L divergence matrix of each feature pairwise, and obtain the sparsity parameter corresponding to the feature according to the K-L divergence matrix of each feature;

[0124] Delete one feature set from the dataset in turn according to the order of the information gain of each feature of the deep learning model from large to small, construct a random forest model according to all the feature sets that have not been deleted after each deletion, and calculate the test performance of the random forest model;

[0125] When the change in the test performance of a random forest model compared to the test performance of the previous model is greater than a preset value, obtain the number of deleted feature sets;

[0126] Calculate and obtain the completeness parameter according to the number of deleted feature sets and the total number of feature sets in the dataset.

[0127] Optionally, the second processing module 12 is specifically configured to:

[0128] Measure and obtain the total number of samples that the classification of the interpretation model finally falls on each leaf node and the number of samples correctly classified by each leaf node;

[0129] Use the formula: Calculate the classification accuracy Acc of each leaf node in the interpretation model i , where i is the leaf node number, n i is the total number of samples that finally fall on this leaf node after being classified by the interpretation model, c i is the number of samples correctly classified by the leaf node, and the tree accuracy includes the classification accuracy of each leaf node.

[0130] Optionally, the third processing module 13 is specifically configured to:

[0131] Use the formula: Calculate the ratio Comp of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category in all samples i , to obtain the tree completeness parameter; where i is the leaf node number, c i is the number of samples correctly classified by this leaf node, n c is the number of samples of the same category as this node in all samples.

[0132] The processing device of the deep learning model provided in any of the above embodiments is used to execute the technical solutions provided in any of the foregoing method embodiments, and their implementation principles and technical effects are similar, which will not be elaborated here.

[0133] Figure 8 As shown in the structural schematic diagram of the first embodiment of the electronic device provided by the present invention, Figure 8 The electronic device 20 includes: a memory 22, a processor 21, and an output interface 23. In addition, it further includes a computer program stored on the memory 22 and executable on the processor 21. When the computer program is executed by the processor 21, it implements the steps of the processing method of the deep learning model provided in any of the foregoing method embodiments.

[0134] Optionally, the above-mentioned components of the electronic device 20 can be connected through a bus 24.

[0135] The memory 22 can be a separate storage unit or an integrated storage unit in the processor 21. The number of processors 21 is one or more.

[0136] In the implementation of the above-mentioned electronic device 20, the memory and the processor are directly or indirectly electrically connected to achieve data transmission or interaction, that is, the memory and the processor can be connected through an interface or integrated together. For example, these components can be electrically connected to each other through one or more communication buses or signal lines, such as being connected through a bus. The memory can be, but is not limited to, a random access memory (Random Access Memory, abbreviated as RAM), a read-only memory (Read Only Memory, abbreviated as ROM), a programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), an erasable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), an electrically erasable read-only memory (Electric Erasable Programmable Read-Only Memory, abbreviated as EEPROM), etc. Among them, the memory is used to store programs, and the processor executes the programs after receiving execution instructions. Further, the software programs and modules in the above-mentioned memory may also include an operating system, which may include various software components and / or drivers for managing system tasks (such as memory management, storage device control, power management, etc.) and may communicate with various hardware or software components to provide a running environment for other software components.

[0137] The processor 21 may be an integrated circuit chip with the ability to process signals. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), an image processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0138] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the processing method of the deep learning model provided in any of the foregoing method embodiments.

[0139] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the processing method of the deep learning model provided in any of the foregoing embodiments.

[0140] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including that element.

[0141] The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course, they can also be implemented by hardware. However, in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing an electronic device (which may be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0142] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A processing method for a deep learning model, characterized in that, Including: Input the pre-acquired dataset into the original convolutional neural network model to be processed, and obtain the information gain, sparsity parameter, and completeness parameter of each feature of the original convolutional neural network model; wherein, the dataset includes data of multiple features, the information gain of each feature is used to represent the ability of the feature to distinguish data samples, the sparsity parameter of each feature is used to represent the degree of independence between features, the completeness parameter of each feature is used to represent the influence degree of the feature on the original convolutional neural network model, and the original convolutional neural network model is used for medical image recognition; Extract the output of the feature extractor and the output of the classifier from the original convolutional neural network model, and use the output of the feature extractor and the output of the classifier as training data for tree model training to obtain an interpretation model; Measure the classification accuracy of the leaf nodes of the interpretation model to obtain a tree accuracy, and the tree accuracy is used to indicate the classification accuracy of the interpretation model; Calculate the ratio of the number of samples that can be correctly classified by each leaf node of the interpretation model to the number of samples of the corresponding category in all samples to obtain a tree completeness parameter; Output the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter; Perform visualization processing according to the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter to obtain a radar chart for providing an interpretation related to the performance of the original convolutional neural network model; Output the radar chart through the user's terminal device to quantitatively explain the specific effects and functions of the original convolutional neural network model in medical image recognition to the user.

2. The method according to claim 1, wherein The step of inputting the pre-acquired dataset into the original convolutional neural network model to be processed and obtaining the information gain, sparsity parameter, and completeness parameter of each feature of the original convolutional neural network model includes: Input the dataset into the original convolutional neural network model, extract the output of the feature extractor of the original convolutional neural network model, filter the output value of each feature in the output of the feature extractor, and take the mean value after filtering to obtain the information gain of the feature of the original convolutional neural network model; Extract all filter matrices from the convolutional layer of the original convolutional neural network model, perform conversions according to the all filter matrices respectively, calculate the K-L divergence matrix of each feature pairwise, and obtain the sparsity parameter corresponding to the feature according to the K-L divergence matrix of each feature; Delete a feature set from the dataset in turn according to the order of the information gain of each feature of the original convolutional neural network model from large to small, construct a random forest model according to all the remaining feature sets after each deletion, and calculate the test performance of the random forest model; When the change in the test performance of a random forest model compared to the test performance of the previous model is greater than a preset value, obtain the number of deleted feature sets; Calculate and obtain the completeness parameter according to the number of the deleted feature sets and the total number of the feature sets in the data set.

3. The method according to claim 1, wherein The measuring the classification accuracy of the leaf nodes of the interpretation model to obtain the tree accuracy includes: Measuring and obtaining the total number of samples on which the classification of the interpretation model finally falls on each leaf node and the number of samples correctly classified by each leaf node; Use the formula to calculate the classification accuracy Acc of each leaf node in the interpretation model i ; where i is the leaf node serial number, and n i is the total number of samples that finally fall into this leaf node after being classified by the interpretation model, and c i is the number of samples with correct classification of the leaf node. The tree accuracy includes the classification accuracy of each leaf node.

4. The method according to claim 1, wherein The calculating the ratio of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category in all samples to obtain the tree completeness parameter includes: Using the formula: Calculate the ratio Comp of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category in all samples i , to obtain the tree completeness parameter; where i is the leaf node serial number, and c i is the number of samples correctly classified by this leaf node, and n c is the number of samples of the same category as this leaf node in all samples.

5. A processing device for a deep learning model, characterized in that, Including: A first processing module, configured to input a pre-obtained data set into an original convolutional neural network model to be processed, and obtain the information gain, sparsity parameter, and completeness parameter of each feature of the original convolutional neural network model; wherein, the data set includes data of multiple features, the information gain of each feature is used to represent the ability of the feature to distinguish data samples, the sparsity parameter of each feature is used to represent the degree of independence between features, the completeness parameter of each feature is used to represent the influence degree of the feature on the original convolutional neural network model, and the original convolutional neural network model is used for medical image recognition; A second processing module, configured to extract the output of the feature extractor and the output of the classifier from the original convolutional neural network model, and use the output of the feature extractor and the output of the classifier as training data for tree model training to obtain an interpretation model; The second processing module is further configured to measure the classification accuracy of the leaf nodes of the interpretation model to obtain the tree accuracy, and the tree accuracy is used to indicate the classification accuracy of the interpretation model; A third processing module, configured to calculate the ratio of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category in all samples to obtain the tree completeness parameter; An output module, configured to output the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter; A fourth processing module, configured to perform visualization processing according to the information gain of each feature, the sparsity parameter of each feature, the completeness parameter of each feature, the tree accuracy, and the tree completeness parameter to obtain a radar chart for providing an interpretation related to the performance of the original convolutional neural network model; The output module is further configured to output the radar chart through the user's terminal device to quantitatively interpret the specific effects and functions of the original convolutional neural network model in medical image recognition for the user.

6. The device according to claim 5, wherein The first processing module is specifically configured to: Input the data set into the original convolutional neural network model, extract the output of the feature extractor of the original convolutional neural network model, filter the output value of each feature in the output of the feature extractor, and take the mean value after filtering to obtain the information gain of the feature by the original convolutional neural network model; All filter matrices are extracted from the convolutional layers of the original convolutional neural network model, and are respectively transformed according to the all filter matrices. The K-L divergence matrix of each feature is calculated pairwise, and the sparsity parameter corresponding to the feature is obtained according to the K-L divergence matrix of each feature; According to the order of the information gain of each feature from large to small of the original convolutional neural network model, a feature set is deleted from the dataset in turn, and a random forest model is constructed according to all the remaining feature sets after each deletion, and the test performance of the random forest model is calculated; When the change in the test performance of a random forest model compared to the test performance of the previous model is greater than a preset value, the number of deleted feature sets is obtained; According to the number of the deleted feature sets and the total number of the feature sets in the dataset, the completeness parameter is calculated and obtained.

7. The device according to claim 5, characterized in that The second processing module is specifically configured to: Measure and obtain the total number of samples that the classification of the interpretation model finally falls on each leaf node and the number of samples correctly classified by each leaf node; Use the formula: Calculate the classification accuracy Acc of each leaf node in the interpretation model i , where i is the leaf node serial number, and n i is the total number of samples that finally fall on this leaf node after classification by the interpretation model, and c i is the number of samples with correct classification of the leaf node. The tree accuracy includes the classification accuracy of each leaf node.

8. The device according to claim 5, characterized in that, The third processing module is specifically configured to: Using the formula: Calculate the ratio Conp of the number of samples that each leaf node of the interpretation model can correctly classify to the number of samples of the corresponding category among all samples i , to obtain the tree completeness parameter; where i is the leaf node serial number, and c i is the number of samples correctly classified by this leaf node, and n c is the number of samples of the same category as this leaf node among all samples.

9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and an output interface. A computer program is stored on the memory and can run on the processor. When the computer program is executed by the processor, the steps of the processing method of the deep learning model according to any one of claims 1 to 4 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the processing method of the deep learning model according to any one of claims 1 to 4 are implemented.

11. A computer program product, characterized in that, It includes a computer program. When the computer program is executed, the steps of the processing method of the deep learning model according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Behavior identification method based on integrated linear classifier and analytic dictionary

    CN105938544A

  • Visual analysis system and method for boosting tree model

    CN107862342A