Training device and program

JPWO2024252857A5Active Publication Date: 2025-11-19NIKON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025526011
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-19
Estimated Expiration
2044-05-14

AI Technical Summary

Technical Problem

Deep learning models, while accurate in decision-making, often lack interpretability, making it difficult for users to understand the basis of their judgments, and existing explanation tools like Grad-CAM, LIME, and SHAP do not provide sufficient explainability.

Method used

A learning device and program that utilize Grad-CAM, LIME, SHAP, and TCAV algorithms to visualize the basis for diagnosis and differential diagnosis by calculating the contribution of image features and metadata in a skin lesion classification task, using a convolutional neural network to extract interpretable features and perform linear regression to compress feature vectors for improved explainability.

Benefits of technology

The solution enhances the explainability of deep learning models by providing a clear visualization of feature contributions, allowing users to understand the decision-making process while maintaining accuracy in classification tasks.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This training device includes: a storage unit that stores a trained model that is trained to receive input of a training image and a training feature amount obtained by digitizing a predetermined interpretable feature related to a subject of the training image, and output a result of determination for the training image and the training feature amount; a determination unit that uses the trained model stored in the storage unit to output a result of determination for an explanation target image and a first feature amount obtained by digitizing a predetermined interpretable feature related to a subject of the explanation target image; and an explanation output unit that outputs the degree of contribution of the explanation target image and the degree of contribution of the first feature amount in relation to the result of determination for the explanation target image and the first feature amount by the trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Learning devices and programs

[0001] The present invention relates to a learning device and a program.

[0002] Patent Document 1 states that "the basis calculation unit 114 can calculate an image that visualizes the basis for each diagnosis and differential diagnosis using the trained machine learning model in the inference unit 112 using algorithms such as Grad-CAM (Gradient-weighted Class Activation Mapping), LIME (LOCAL Interpretable model-agnostic Explanations), SHAP (Shapley Additive explanations), which is an advanced version of LIME, and TCAV (Testing with Concept Activation Vectors)."

[0003] Non-Patent Document 1 describes a learning model that classifies skin lesions by inputting skin images and metadata (such as the patient's age and gender). [Prior Art Documents] [Patent Documents] [Patent Document 1] International Publication No. 2022 / 176396 "Non-Patent Document 1" Nils Gessert et al., "Skin Lesion Classification Using Ensembles of Multi-Resolution EfficientNets with Meta Data," Methods X, Vol. 7, 2020 General disclosure

[0004] In a first aspect of the present invention, a learning device includes a memory unit that stores a learning model that receives as input a learning image and a learning feature that is related to the subject of the learning image and that is a quantified version of a predetermined interpretable feature, and outputs the result of a judgment on the learning image and the learning feature; a judgment unit that outputs the result of a judgment using the learning model stored in the memory unit on an image to be explained and a first feature that is related to the subject of the image to be explained and that is a quantified version of a predetermined interpretable feature; and an explanation output unit that outputs the contribution of the image to be explained and the contribution of the first feature to the result of the judgment on the image to be explained and the first feature by the learning model.

[0005] The predetermined interpretable features may include parameters that can quantitatively represent the shape or characteristics of the subject. The determination unit may use a learning model to extract second features from the image to be explained, and the explanation level output unit may calculate the contribution levels by linearly regressing neighborhoods of data consisting of the first and second features related to the image to be explained in a feature space including the second features and the first features. The determination unit may use the learning model to extract the second features from the image to be explained, and the explanation output unit may compress the dimensions of a feature vector that is the second feature and calculate the contribution level of the second feature as the contribution level of the image to be explained. The explanation output unit may compress the feature vector into one dimension.

[0006] The image processing device may further include a feature calculation unit that calculates at least one of the first feature amounts based on the explanation target image and inputs the calculated first feature amount into the learning model.

[0007] The training model may include a convolutional neural network.

[0008] The explanation output unit may further output an image showing a determination basis for the image to be explained with respect to the determination result by the learning model. The explanation output unit may display the contribution degree of the image to be explained and the contribution degree of the first feature amount side by side. The explanation output unit may display the contribution degree of the image to be explained, the contribution degree of the first feature amount, an image showing the determination basis, and the determination result side by side. The explanation output unit may display text corresponding to each of the contribution degree of the image to be explained and the contribution degree of the first feature amount.

[0009] In a second aspect of the present invention, a program is provided that causes a computer to realize the following: a memory function that receives as input a training image and training features that relate to the subject of the training image and that are quantified as predetermined interpretable features, and stores in a memory unit a learning model that has been trained using the results of judgment on the training image and training features as output; a judgment function that uses the learning model stored in the memory unit to output the results of judgment on an image to be explained and a first feature that relates to the subject of the image to be explained and that is quantified as predetermined interpretable features; and an explanation output function that outputs the contribution of the image to be explained and the contribution of the first feature to the results of judgment on the image to be explained and the first feature by the learning model.

[0010] In a third aspect of the present invention, a learning device uses a learning model that receives as input a learning image and a learning feature that relates to the subject of the learning image and that quantifies a predetermined interpretable feature, and outputs the result of judgment on the learning image and the learning feature, for an image to be explained and a first feature that relates to the subject of the image to be explained and that quantifies a predetermined interpretable feature, and includes an explanation output unit that outputs the contribution of the image to be explained and the contribution of the first feature to the result of judgment on the image to be explained and the first feature by the learning model.

[0011] The above summary of the invention does not list all of the features of the present invention, and subcombinations of these features may also be inventions.

[0012] 1 shows functional blocks of a learning device 10 according to this embodiment. 2 shows an operational flow of the learning device 10. 3 shows an example of a learning model 120. 4 shows a distribution diagram for explaining an overview of a method for linearly regressing a feature space. 5 shows a distribution diagram explaining a human-readable feature space. 6 shows a distribution diagram explaining a method for compressing the dimension of a second feature to one dimension. 7 shows an example of an image 200 displayed on a display by the explanation output unit 104. 8 shows a text template 230 displayed in a text area 220. 9 shows a list illustrating other application examples. 10 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied in whole or in part.

[0013] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0014] FIG. 1 shows functional blocks of a learning device 10 according to this embodiment. The learning device 10 outputs some kind of judgment result for a target input. The judgment may be determining which of predetermined categories the target input belongs to, or may be calculating the likelihood of the target input belonging to each category as a predicted probability. In this embodiment, the target input is a cell image and feature quantities (first feature quantities) that quantify predetermined interpretable features related to the cell depicted in the image. The output judgment is the predicted probability of the cell depicted in the image and the first feature quantities predicting which of four states of the cell cycle the cell is in. This output may also be called a decision, judgment, prediction, inference, or estimation.

[0015] For such judgments, learning models such as deep learning are used. In learning models, the accuracy of judgment and the ease of understanding the basis for the judgment are often in a contradictory relationship. For example, deep learning is a learning model that makes relatively accurate judgments, but it is difficult for users to interpret it by looking at the nodes and their weights in each layer. Therefore, tools for explaining the basis for judgments, such as Grad-CAM, LIME, SHAP, and TCAV, have been proposed, but none of them can be said to be sufficient. In this embodiment, the objective is to present the basis for judgment in an easy-to-understand manner while ensuring the accuracy of the judgment. Note that the ease of understanding the basis is sometimes referred to as high explainability.

[0016] The learning device 10 includes a feature calculation unit 100, a judgment unit 102, an explanation output unit 104, and a memory unit 106. The memory unit 106 stores a learning model 120. The learning device 10 may be an information device such as a personal computer, tablet, or smartphone, and may be realized by installing a program or application on such an information device. The learning device 10 may also be a web server that outputs some kind of judgment result for a target input using cloud computing, and may be realized by installing a program or application on the web server. The web server is connected to a microscope or the like via a network.

[0017] The feature calculation unit 100 acquires an input image from an external device such as a microscope. The feature calculation unit 100 may read out an input image pre-stored in the storage unit 160. The feature calculation unit 100 calculates feature amounts (first feature amounts) obtained by quantifying predetermined interpretable features from the input image, and inputs the first feature amounts to the learning model 120.

[0018] The determination unit 102 performs a determination on the input image using the learning model 120 stored in the storage unit 106. The determination unit 102 outputs the result of the determination to an external device such as a display.

[0019] When the learning model 120 obtains a judgment result, the explanation output unit 104 outputs the contribution of the input image (feature values ​​(second feature values ​​described later) extracted from pixel values ​​of the input image) and the contribution of feature values ​​(first feature values) related to the subject of the input image and obtained by quantifying predetermined interpretable features, in a comparable manner. The output destination is, for example, a display, similar to the judgment unit 102.

[0020] 2 shows the operational flow of the learning device 10. This operational flow is initiated, for example, by the user starting up the learning device 10. The operational flow includes a learning stage S100 in which the learning model 120 is trained, a determination stage S100 in which the learning model 120 is used to determine the explanation target, and an explanation stage S120 in which the contribution rate and the like are calculated to explain the result of the determination.

[0021] 3 schematically illustrates an example of the learning model 120. In the learning stage of step S100, the learning model 120 is first set. In this embodiment, the learning model 120 is set as a model that combines machine learning using the input image 20 itself, in other words, pixel values, with machine learning using feature quantities that are related to the subject of the input image 20 and that are predetermined interpretable features quantified. This model is sometimes called a multimodal model from the viewpoint of using different types of input.

[0022] The learning model 120 includes an image CNN 132 as a machine learning model using pixel values. The image CNN 132 is a CNN (convolutional neural network) that uses an image as input. While there are no limitations on the number of layers of the CNN or the number of nodes (also called filters, kernels, etc.) in each layer, this embodiment uses a model up to the fully connected layer in VGG16 (i.e., VGGNet16 layer). The image CNN 132 repeatedly performs convolution and pooling on an input image 20 having, for example, pixel values ​​(96 vertical × 96 horizontal × 3 colors) to calculate 512 nodes and weights for each node. The weights for the 512 nodes are used in a subsequent stage as 512-dimensional second features.

[0023] The learning model 120 includes a feature NN 134 as a machine learning model that uses quantified features. The feature NN 134 is a single-layer or multi-layer neural network, and there is no limit to the number of layers or the number of nodes in each layer. In this embodiment, the feature NN 134 receives input of a first feature including 13 features and calculates eight nodes and weights for each node. The weights for the eight nodes are used in a subsequent stage as eight-dimensional third features.

[0024] The first feature amount is calculated from the input image 20 by the feature calculation unit 100. OpenCV is used as an example of the feature calculation unit 100, but other image processing engines may also be used. Note that instead of using the feature calculation unit 100, the user may specify the first feature amount and input it directly to the feature amount NN 134.

[0025] The first feature quantity is a numerical representation of a predetermined interpretable feature related to the subject of the input image 20. The predetermined interpretable feature is a feature that can be quantitatively and intuitively understood by a user. In this embodiment, the predetermined interpretable feature includes a feature related to the shape of the cell, corresponding to determining the cycle of the cell captured in the input image 20. Examples of the feature include 13 features: cell area, convex hull, perimeter, maximum fillet size, minimum fillet size, diameter, rectangular length, rectangular width, circularity, convexity, elongation, roughness, and unevenness. The first feature quantity may be other features (e.g., cell density), and the number of features is not limited.

[0026] The learning model 120 further includes a classification NN 136. The classification NN 136 is a fully connected layer that is the final layer of the image CNN 132 and the final layer of the feature NN 134, and performs four class determinations. In other words, the classification NN 136 can be said to be a classifier that determines four classes from a 520-dimensional input that combines a 512-dimensional second feature and an 8-dimensional third feature. The four classes determined are "G1: DNA synthesis preparation" (class 1), "S: DNA synthesis and replication" (class 2), "G2: two sets of chromosomes" (class 3), and "M: cell division" (class 4), which correspond to each phase of the cell cycle.

[0027] The learning model 120 before learning is initially set based on, for example, input from a user. In this case, the learning device 10 may store an outline of the model in the storage unit 108, receive input from the user regarding the number of layers and the number of nodes in the network, and initialize the learning model 120 based on the input.

[0028] In the learning stage of step S100, for example, 100,000 pairs of input images 20 as learning images (also called teacher images), first learning features, and correct classes are prepared. The learning model 120 is trained by reducing the error between the correct answer and the judgment results when these learning images and first learning features are input. For example, backpropagation is used as a method for reducing the error. Other learning methods may be used, and further, a learning method such as dropout may be used in combination.

[0029] As described above, the pixel values ​​of the learning input image 20 and the learning first feature are input, and the results of a predetermined judgment (period) on the subject (cell) of the input image 20 are output, and the learning model 120 is trained. The trained learning model 120 is stored in the storage unit 106. This completes the operation of step S100. Note that the contents of the 512 nodes of the second feature change as learning progresses, but even if the nodes after learning are visually displayed, the user will not intuitively understand what they represent.

[0030] Next, in the judgment stage of step S110, the judgment unit 102 reads out the trained learning model 120 from the storage unit 106, inputs the input image 20 to be explained and feature amounts (first feature amounts) that relate to the subject (cell) of the input image 20 to be explained and that quantify predetermined interpretable features, into the learning model 120, and performs a predetermined judgment (cell cycle) of the subject. As a result of the judgment, the judgment unit 102 outputs a predicted probability indicating the possibility of belonging to each class for each of the four classes.

[0031] However, simply outputting the determination result may not allow the user to understand why the determination result was reached, and the determination result may not be put to good use.

[0032] Next, in the explanation stage of step S120, two processes are performed in parallel. The first process (step S122) is a process for calculating the contribution of the image itself and the contribution of a feature (first feature) that is related to the subject of the image and that quantifies a predetermined interpretable feature when the learning model 120 obtains a judgment result. The second process (step S124) is a process for visualizing the area in the image that contributed to the judgment result (visualizing the judgment basis). The first process and the second process are performed by the explanation output unit 104. First, the first process will be described. In this embodiment, the first process is a process based on the well-known method LIME (LOCAL Interpretable model-agnostic Explanations), and uses a method for estimating an explanation model in the vicinity of the data to be explained in the feature space.

[0033] The data to be explained consists of a feature (second feature) of the input image 20 to be explained, and a feature (first feature) that is related to the subject (cell) of the input image 20 to be explained and that is a quantified version of a predetermined interpretable feature.

[0034] FIG. 4 is a distribution diagram for explaining an outline of a method for linearly regressing a space of feature quantities. In the example of FIG. 4, the feature quantity x 0 and feature x 1 A two-dimensional feature space with a spatial axis of is depicted in the figure. In the figure, features of the explanation target are indicated by crosses, other input images with the highest predicted probability of belonging to the same class as the explanation target data are indicated by black circles, and other input images with the highest predicted probability of belonging to a class different from the explanation target are indicated by white circles. In the figure, the solid line represents the discrimination surface (boundary surface) learned by the learning model 120 that determines whether an input image belongs to the same class as the explanation target or a different class.

[0035] As shown by the solid line, the discrimination surface of the learning model 120 is complex. However, if we focus on the "neighborhood of the data to be explained" and perform linear regression (linear approximation) on the discrimination surface, the slope of the line will be similar to the feature value x 0 and feature x 1 In other words, by focusing on the neighborhood of the data to be explained and approximating the learning model there with a linear classifier, the weight of each feature value can be regarded as the contribution of each feature value in the learning model 120 to the judgment result.

[0036] Therefore, training data is extracted, and prediction probabilities are calculated again using the trained training model 120. The training data consists of features of the training input image 20 and features (first training features) that relate to the subject of the training input image and that are predetermined interpretable features that are quantified.

[0037] Here, in order to define the "neighborhood of the data to be explained," the feature space shown in FIG. 4 is converted into a "human-readable feature space."

[0038] The "readable feature space" is defined as follows: (1) Each feature is divided into an independent region (one region corresponds to one square in Figure 5) according to the density of the data to be explained and the training data (collectively referred to as "data"). (2) For a given feature, a readable feature value of 1 is assigned to training data that belongs to the same level as the data to be explained, and a value of 0 is assigned if it does not. In other words, the feature value of each data item is projected onto a binary space. This assignment is made for each data item and each feature, and a cost function for linear regression is set based on this.

[0039] Regarding the above definition (1) of the human-readable feature space, the division into regions is performed so that each region contains a predetermined number or range of data. As a result, as shown in Figure 5, for each feature, the width of the region where the data density is high, i.e., the numerical range, becomes narrower. Conversely, the width of the region where the data density is low becomes wider.

[0040] Regarding the above definition (2) of the human readable feature space, by using the region division, human readable feature z i The feature value of a certain data x is assigned. i If it is in the same region as the data to be explained, the readable feature z i = 1. Conversely, the feature x i If it is not in the same region as the data to be explained, the readable feature z i = 0. For each feature of the data, a readable feature is assigned according to the above. This defines a space in which 1s line up in the z component if it is close to the data to be explained, and 0s line up if it is far away.

[0041] FIG. 5 is a distribution diagram illustrating the readable feature space. In the example of FIG. 5, the feature x 0 and feature x 1 The figure depicts a two-dimensional readable feature space with the spatial axis as . In the figure, the data to be explained are indicated by crosses, and the extracted training data are represented by black circles. Furthermore, the number of data points for each feature, i.e., the data density, is shown schematically with a curve.

[0042] Next, linear regression is performed in the readable feature space. First, the i-th data is formulated as follows: Here, y i is the predicted probability that the data is predicted to be in the same class as the data to be explained, w is a vector having the dimension of the feature, and b is a constant term.

[0043] For each of the data to be explained and the training data, y i and Z i Then, under the above formulation, w and b that minimize the square error between the left and right sides of the above "Equation 1" are estimated as follows:

[0044] The i-th component w of the coefficient w obtained by this estimation i is the feature x calculated by linear regression in the readable feature space. i Furthermore, the coefficient w i The magnitude relationship between the features x of the learning model 120 is i This can be said to reflect the relationship between the magnitude of contributions of the

[0045] In addition, the coefficient w i If is a positive value, it contributes to increasing the predicted probability, and if it is negative, it contributes to decreasing it. The constant b is the model bias, which corresponds to the predicted probability when random data is input.

[0046] The above-mentioned description of the human-readable feature space shown in Fig. 5 and the linear regression in that space is a general description of the well-known LIME. In this embodiment, different types of features, namely, an image and first features (i.e., features obtained by quantifying predetermined interpretable features), are input to LIME, and the following processing is performed to calculate the contribution of each feature: (1) Convert the first features into human-readable features; (2) Reduce the dimension of the image features (second features) and convert them into human-readable features.

[0047] The above conversion (1) is performed to convert into the human-readable feature space shown in FIG.

[0048] On the other hand, for the above transformation (2), since the second feature extracted from the neural network to which the image is input has 512 dimensions in this embodiment, if these are used as the dimensions (i.e., spatial axes) of the human-readable feature space, 512 numerical values ​​are obtained as the contributions of the second feature. However, as described above, the characteristics of the second feature themselves cannot be intuitively understood, and therefore the contributions do not significantly improve interpretability.

[0049] Therefore, in this embodiment, the dimension of the second feature is reduced and the contribution is calculated by the above method. For example, the dimension of the second feature is compressed to one dimension so that the contribution after compression can be considered as the contribution of "the image itself."

[0050] Fig. 6 is a distribution diagram illustrating a method for compressing the dimension of the second feature in the feature space shown in Fig. 4 to one dimension. In Fig. 6, the data to be explained are indicated by crosses, and the training data are indicated by black circles.

[0051] The method shown in FIG. 6 compresses the 512 dimensions before compression into a feature quantity that is the distance from the explanation target data in the 512-dimensional space. That is, the distance d is calculated as follows: where x ev v , x v are the values ​​of the v-dimension out of 512 dimensions in the input image to be explained and the input image for training, respectively.

[0052] Furthermore, corresponding to the first feature, the feature after compression is also assigned a readable feature of 0 or 1. In this embodiment, if the distance d is equal to or greater than a threshold, the readable feature z=0 is assigned. Conversely, if the distance d is smaller than the threshold, the readable feature z=1 is assigned.

[0053] As a result, the second feature is converted into one dimension and projected into a binary space as a readable feature. This can be said to define a readable feature of the image itself that reflects the second feature.

[0054] A total of 14-dimensional readable feature space is set up by combining the one-dimensional readable feature of the image itself and the 13-dimensional readable feature of the first feature, and linear regression of the above "Equation 2" is performed. The resulting contribution is the contribution of the feature corresponding to the 14 dimensions, allowing the user to refer to both the contribution of the image itself and the contribution of the first feature in a form that allows comparison. Because the first feature is a numerical representation of a predetermined interpretable feature, the obtained results not only improve interpretability but also reveal the contribution of the image itself, allowing the user to evaluate the meaning of the heat map created in the second process described below.

[0055] Next, the second process (step S124) will be described. After the judgment stage of step S110, for example, a heat map is created in step S124. The heat map is a comparative output of the contribution of each region of the image to be explained when the learning model 120 obtains a judgment result for the explanation target. Note that the method is not limited to a heat map, and any method for visualizing the regions in the image to be explained that contributed to the judgment result (visualizing the basis for the judgment) may be used. For example, the contributing regions (regions that served as the basis) (if there are multiple regions, all regions) may be extracted from the image to be explained for each degree of contribution to the judgment result. In this embodiment, the second process is performed using only information from the image CNN 132 in the learning model 120.

[0056] There are various methods for generating heat maps, including CAM (Class Activation Map), a class activation mapping method for learning models, and its derivatives (which use gradients for weights), such as Grad-CAM and Grad CAM++, as well as ScoreCAM, which does not use gradients for weights but instead assigns weights through forward propagation, as well as Guide-BP and IngetrateGrad.

[0057] In a typical Grad-CAM, the gradient of the output from the final layer of a CNN is used to calculate the influence of each pixel value of the input image on the predicted probability of each class. However, instead of this, the gradient of the output from the intermediate layer, the average of the gradients of the output from all layers, etc. may be used. Furthermore, instead of using a heat map, a perturbation may be applied to the input image to divide it into several superpixels, and then the aforementioned LIME may be applied to visualize the regions in the input image that serve as the basis for judgment. Any of these methods may be used in this embodiment.

[0058] After the first process (step S122) and the second process (step S124) described above are completed, the explanation output unit 104 outputs the result of the first process (contribution degree) and the result of the second process (hereinafter referred to as the processing result) to a display or the like. Instead of or in addition to outputting the processing result to a display, the explanation output unit 104 may store the processing result in the storage unit 106. Furthermore, the processing result may be output together with the result of the determination by the determination unit 102.

[0059] 7 shows an example of a display image 200 displayed on the display by the explanation output unit 104. In Fig. 7, an input image that is input to the learning model 120 is displayed in a target image area 202. Similarly, for the first feature that is input to the learning model 120, the names of 13 features and bar graphs showing their sizes are displayed in a first feature area 204 in association with each other.

[0060] The contributions calculated in step S122 (first process) are displayed in the display image 200 in multiple ways. First, each contribution area 208 displays the name of a feature and a bar graph indicating the magnitude of its contribution. The same 13 first feature quantities as the input are displayed. Meanwhile, the second feature quantity is displayed as a single feature corresponding to the one-dimensional compression. Furthermore, these are displayed vertically. This allows the user to easily recognize the contributions of the first feature quantity and the second feature quantity, improving the interpretability of the judgment. Furthermore, since there is only one contribution quantity for the second feature quantity, this can be interpreted as the contribution of "the image itself," further improving the interpretability.

[0061] The cumulative contribution area 210 displays bar graphs for prediction probability, feature values ​​that increase prediction probability, and feature values ​​that decrease prediction probability, aligned vertically for comparison. The bar graphs for feature values ​​that increase prediction probability are arranged in series from the left in descending order of positive contribution values. The bar graphs for feature values ​​that decrease prediction probability are arranged in series from the left in descending order of negative contribution values. The right end of each bar graph is aligned with the right end of the bar graph above it. Furthermore, the name of the feature is written above each contribution value that is longer than a predetermined length. These displays further improve the ease of explanation.

[0062] Furthermore, the contribution degree is displayed using text in a text area 220. The file name of the explanation target, the predicted probability, the predicted class, and a text report are displayed in the text area 220. As the text report, text corresponding to the contribution degree of the second feature amount and the contribution degree of the first feature amount may be displayed.

[0063] 8 shows a template 230 of text to be displayed in the text area 220. The template 230 is stored in the storage unit 106.

[0064] The template 230 includes a preset sentence and variables to be inserted into the sentence. The variables are indicated by square brackets [ ], and values ​​for the symbols written within the square brackets are substituted by the determination unit 102 and the contribution output unit 104 and displayed in the text area 220.

[0065] The number of features to be displayed in the text area 220 may be predetermined, or features greater than a threshold value of the contribution degree or a threshold value of the absolute value of the contribution degree may be displayed. These rules may also be stored in the storage unit 106.

[0066] Furthermore, the heat map created in step S124 (second process) is displayed in the heat map area 206. These displays further improve the ease of explanation.

[0067] As described above, according to the present embodiment, it is possible to present the basis of the judgment in an easy-to-understand manner while ensuring the accuracy of the judgment. In particular, it is possible to improve the explainability of the judgment in a so-called multimodal model in which the judgment is made using inputs of different types of features.

[0068] A modification of the above embodiment will be described. Instead of VGGNet, the image CNN 132 may be another CNN, including AlexNet, VGGNet, ResNet, ResNeXt, etc. Furthermore, other neural networks other than CNN may also be used. Furthermore, although the 512 dimensions of the final layer of the image CNN 132 are used as the second feature, a feature from an intermediate layer before the final layer may be used instead of or in addition to this.

[0069] The reduction of the dimension of the second feature is not limited to one-dimensional reduction using distance. As another example, the dimension of the second feature may be reduced using principal component analysis or a nonlinear dimensionality reduction algorithm related thereto. As yet another example, the dimension may be reduced to one-dimensional reduction by statistical processing such as taking a simple average of the second feature or the maximum value thereof.

[0070] Explanatory models are not limited to linear regression.

[0071] In the above embodiment, an example has been described in which the learning model 120 is used to determine the cell cycle from a cell image. However, the use of the learning model 120 is not limited to this. Other examples of applicable uses are listed in FIG. 9 together with input images and first feature amounts.

[0072] The first feature amount can be calculated automatically or manually from the input image, or it can be calculated from the input image. Examples of the first feature amount that can be calculated include shape features such as the radius and length of the subject (object, living body) in the image, color features, and characteristic features. Examples of the first feature amount that cannot be calculated include attribute information such as the gender, age, and race of the owner of the object (subject) in the input image and the subject of the living body (subject) in the image (these are related to the subject and correspond to predetermined interpretable features), and for example, location information is quantified using coordinates or an index.

[0073] Various embodiments of the present invention may also be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which operations are performed or (2) sections of an apparatus responsible for performing the operations. Particular stages and sections may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable medium, and / or a processor provided with computer-readable instructions stored on a computer-readable medium. Dedicated circuitry may include digital and / or analog hardware circuitry, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuitry may include reconfigurable hardware circuitry including logical AND, OR, XOR, NAND, NOR, and other logic operations, flip-flops, registers, memory elements such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.

[0074] A computer-readable medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that the computer-readable medium having instructions stored thereon comprises an article of manufacture containing instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable media may include electronic, magnetic, optical, electromagnetic, and semiconductor storage media. More specific examples of computer-readable media may include floppy disks, diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), electrically erasable programmable read-only memories (EEPROMs), static random access memories (SRAMs), compact disc read-only memories (CD-ROMs), digital versatile discs (DVDs), Blu-ray (RTM) discs, memory sticks, integrated circuit cards, and the like.

[0075] The computer readable instructions may include either assembler instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages ​​such as the “C” programming language or similar programming languages.

[0076] The computer-readable instructions may be provided to a processor or programmable circuitry of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, either locally or over a local area network (LAN), a wide area network (WAN) such as the Internet, etc., which executes the computer-readable instructions to create means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

[0077] 10 illustrates an example of a computer 2200 in which aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 may cause the computer 2200 to function as or perform operations associated with an apparatus or one or more sections of the apparatus according to embodiments of the present invention, and / or to perform a process or steps of a process according to embodiments of the present invention. Such programs may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.

[0078] A computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.

[0079] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the graphics controller 2216 itself, and causes the image data to be displayed on the display device 2218.

[0080] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.

[0081] ROM 2230 stores therein a boot program or the like that is executed by computer 2200 upon activation, and / or programs that depend on the hardware of computer 2200. I / O chip 2240 may also connect various I / O units to I / O controller 2220 via parallel ports, serial ports, keyboard ports, mouse ports, etc.

[0082] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by implementing information manipulation or processing in accordance with the use of the computer 2200.

[0083] For example, when communication is performed between computer 2200 and an external device, CPU 2212 may execute a communication program loaded into RAM 2214 and instruct communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of CPU 2212, communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in RAM 2214, hard disk drive 2224, DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes received data received from the network to a reception buffer processing area or the like provided on the recording medium.

[0084] Furthermore, the CPU 2212 may cause all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and may perform various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.

[0085] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 2214. The CPU 2212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored on the recording medium, the CPU 2212 may search for an entry that matches a condition specified by the attribute value of the first attribute from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0086] The above-described programs or software modules may be stored in a computer-readable medium on or near the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable medium, thereby providing the programs to the computer 2200 via the network.

[0087] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.

[0088] It should be noted that the order of execution of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order.

[0089] 10 Learning device, 20 Image, 100 Feature calculation unit, 102 Judgment unit, 104 Explanation output unit, 106 Memory unit, 120 Learning model, 132 Image CNN, 134 Feature NN, 136 Classification NN, 200 Display image, 202 Target image area, 204 First feature area, 206 Heat map area, 208 Each contribution area, 210 Accumulative contribution area, 220 Text area, 230 Template, 2200 Computer, 2201 DVD-ROM, 2210 Host controller, 2212 CPU, 2214 RAM, 2216 Graphics controller, 2218 Display device, 2220 Input / output controller, 2222 Communication interface, 2224 Hard disk drive, 2226 DVD-ROM drive, 2230 ROM, 2240 input / output chip, 2242 keyboard

Claims

1. a storage unit that receives as input a training image and training features that are related to a subject of the training image and that are predetermined interpretable features that are quantified, and that stores a training model trained using as output a result of judgment on the training image and the training features; a determination unit that outputs a result of determination using the learning model stored in the storage unit for an image to be explained and a first feature amount that is related to a subject of the image to be explained and that quantifies the predetermined interpretable feature; an explanation output unit that outputs a contribution degree of the explanation target image and a contribution degree of the first feature amount in response to the determination result of the explanation target image and the first feature amount by the learning model; a feature calculation unit that calculates at least one of the first feature amounts based on the explanation target image and inputs the calculated first feature amount into the learning model; A learning device comprising:

2. The learning device according to claim 1 , wherein the predetermined interpretable features include parameters that can quantitatively represent the shape or characteristics of the subject.

3. the determination unit extracts a second feature amount from the explanation target image using the learning model; 2. The learning device according to claim 1, wherein the explanation output unit calculates the degree of contribution by performing linear regression on a neighborhood of data consisting of the first feature and the second feature related to the image to be explained in a feature space including the second feature and the first feature.

4. the determination unit extracts a second feature amount from the explanation target image using the learning model; The learning device according to claim 1 , wherein the explanation output unit compresses the dimension of a feature vector that is the second feature and calculates the contribution of the second feature as the contribution of the explanation target image.

5. The learning device according to claim 4 , wherein the explanation output unit compresses the feature vector into one dimension.

6. The learning device according to claim 1 , wherein the learning model includes a convolutional neural network.

7. The learning device according to claim 1 , wherein the explanation output unit further outputs an image showing a basis for a determination made within the explanation target image in response to the determination result made by the learning model.

8. The learning device according to claim 1 , wherein the explanation output unit displays the contribution of the explanation target image and the contribution of the first feature amount side by side.

9. The learning device according to claim 7 , wherein the explanation output unit displays the contribution of the image to be explained, the contribution of the first feature, an image showing the basis for the judgment, and the result of the judgment side by side.

10. The learning device according to claim 8 , wherein the explanation output unit displays text corresponding to each of the contribution degree of the explanation target image and the contribution degree of the first feature amount.

11. On the computer, a memory function that receives as input a learning image and learning features that are related to a subject of the learning image and that are predetermined interpretable features that are quantified, and stores in a memory unit a learning model that has been trained using a judgment result for the learning image and the learning features as an output; a determination function that uses the learning model stored in the storage unit to output a determination result for an image to be explained and a first feature amount that is related to a subject of the image to be explained and that quantifies the predetermined interpretable feature; an explanation output function that outputs a contribution degree of the explanation target image and a contribution degree of the first feature amount in response to the determination result of the explanation target image and the first feature amount by the learning model; a feature calculation function that calculates at least one of the first feature amounts based on the image to be explained and inputs the calculated feature amount into the learning model; A program to make this happen.

12. A learning device that receives as input a learning image and a learning feature that is related to a subject of the learning image and that quantifies a predetermined interpretable feature, and outputs a result of judgment on an image to be explained and a first feature that is related to a subject of the image to be explained and that quantifies the predetermined interpretable feature, using a learning model that has been trained using as output a result of judgment on the learning image and the learning feature, an explanation output unit that outputs a contribution degree of the explanation target image and a contribution degree of the first feature amount in response to the determination result of the explanation target image and the first feature amount by the learning model; a feature calculation unit that calculates at least one of the first feature amounts based on the explanation target image and inputs the calculated first feature amount into the learning model; A learning device comprising: