Method for evaluating credibility of output of convolutional neural network

By generating and analyzing the feature map covariance matrix, an extradomain classifier was constructed, which solved the problem of convolutional neural network output confidence evaluation, and improved the output reliability and accuracy of the model under extradomain input.

CN120344976APending Publication Date: 2025-07-18VALITA CELLS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380085143.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-12-12
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the credibility of convolutional neural network outputs, especially when the input data exceeds the model domain, there is a lack of simple and robust method for detecting out-of-domain inputs.

Method used

By generating multiple feature map covariance matrices (FMCMs) for each input and linearizing them into vectors to represent points in the feature map covariance space, an extradomain classifier is constructed to determine whether the new input is inside or outside the training feature map covariance space, and using principal component analysis to reduce the dimension to form the training feature map covariance space.

Benefits of technology

The credibility evaluation of the output of convolutional neural network is realized, and the ability to identify out-of-domain inputs is improved, and the reliability and accuracy of model outputs are improved, especially in image analysis such as microscope image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120344976A_ABST
    Figure CN120344976A_ABST
Patent Text Reader

Abstract

A system is disclosed that includes a CNN configured to predict an output for each input, and generate an FMCM for each input, thereby generating a plurality of training FMCMs for a plurality of training inputs that have been used to train the CNN, and forming a training feature map covariance space based on the plurality of training FMCMs. The system also includes an out-of-domain classifier constructed based on the training feature map covariance space and configured to run the new input to classify the new input as within or out of the domain of the CNN based on whether the corresponding new FMCM is within or outside the training feature map covariance space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to conventional neural networks and, more particularly, to assessing the credibility of the output of a convolutional neural network. Background Art

[0002] Convolutional neural networks (CNNs) have revolutionized image analysis, performing as well as or better than humans. This makes it possible for any image analysis currently done by humans to be replaced by artificial intelligence. However, it can be said that the main obstacle to this artificial intelligence (AI) revolution is knowing when and to what extent to trust the output of a CNN. This is all within the scope of "model trust". The prior art provides some insights into "what the model focuses on" and what information the model uses to make decisions. However, the prior art gives very little insight into when the model cannot give the correct answer for a given input. Existing methods for providing such insights are non - quantitative and rely on human interpretation. However, if human interpretation is required for all images, the meaning of using AI is completely lost.

[0003] There may be two main reasons why a model cannot process an input. First, the information may not be in the image and / or is corrupted, i.e., the image is "noisy" or "out of focus". This is particularly relevant for microscopy image analysis. Second, it may be the case that the data is "out - of - domain" for the model. For example, if a model has been designed to predict whether yeast cells are alive or dead, it is unlikely to give a reasonable answer when mammalian cells are used because this is outside the "domain" of the model. In other words, the model will extrapolate from out - of - domain data to an answer for the data used to train the model.

[0004] Some efforts have been made to quantify whether the output of a CNN model can be trusted. In fact, the simplest method is to periodically check the model against known answers, but this may not detect intermittent failures, and additionally, model failures may not be detected before performing this check.

[0005] Some work has been done on using "Bayesian" methods to model the output. In this sense, multiple samples are run through "sub - networks" within the network, and the consistency of the multiple samples is taken as a measure of confidence. However, this information is limited and may not necessarily pick up "out - of - domain" inputs.

[0006] WO 2022 / 108912 discloses an automatic image quality control system. The system includes: a first component including a first deep learning network model for estimating the quality of medical images acquired by an imaging device; and a second component for determining whether a medical image is outside the distribution of a training dataset used for training the first deep learning network model. However, the document uses a covariance matrix when creating a distance metric between feature maps. This is essentially only used to create a meaningful distance measurement. For example, if many elements of a feature map from the training data are highly correlated, their importance is reduced by the covariance metric because the results of these elements are less important.

[0007] WO 2022 / 045915 discloses processing input data through a neural network. The methods and devices of some embodiments process input data through at least one layer of a neural network and thereby obtain a feature tensor.

[0008] Therefore, in view of the above, there is a need for simple and robust methods to evaluate the credibility of the output of a convolutional neural network and also to check whether a new input is "out-of-domain" of the convolutional neural network model. Summary of the Invention

[0009] According to the present invention, as set forth in the appended claims, there is provided a system for determining whether a new input is out-of-domain of a convolutional neural network (CNN), the system including a CNN configured to: predict an output for each input; and generate a plurality of feature map covariance matrices (FMCMs) for each input. The CNN is configured to: for each CNN layer, generate the FMCM of each input; and linearize each FMCM of a given CNN layer into a vector to represent a point in the feature map covariance space of the CNN layer.

[0010] The CNN is further configured to: calculate the FMCM of a given layer for a given input based on the covariance between a plurality of feature maps of the given layer for the given input. The system further includes an out-of-domain classifier constructed based on the training feature map covariance space of a plurality of training FMCMs and configured to: run a new input to classify the new input as in-domain or out-of-domain of the CNN based on whether the corresponding new FMCM is within or outside the training feature map covariance space, and wherein the plurality of training FMCMs are generated by the CNN for a plurality of training inputs that have been used to train the CNN to predict an output for each input.

[0011] Conventionally, FMCM has been utilized in the generation network. The input of the network can be modified such that the internal FMCM of the network matches the target. For example, the input image can be modified such that the FMCM of the network matches the FMCM of a different input of the network, i.e., a smaller target image, to generate the output. However, there is no mention of extracting FMCM by the network to determine "out-of-domain" inputs.

[0012] Various embodiments of the present invention facilitate processing the training data distribution of a CNN in the training feature map covariance space to examine whether a new input / data has a "style" that is sufficiently similar to the training data. When the new input is in a different style space, then it is out-of-domain for the CNN, and thus the output of the CNN based on this new input is not trustworthy.

[0013] In one embodiment, the CNN is further configured to calculate the FMCM of a given CNN layer for a given input by computing the Gram matrix between each pair of feature maps in the corresponding plurality of feature maps.

[0014] In one embodiment, the out-of-domain classifier is further configured to: label a new input when the new input is classified as being out-of-domain for the CNN.

[0015] In one embodiment, the CNN is configured to: reduce the dimension of each FMCM based on principal component analysis (PCA) to form the training feature map covariance space.

[0016] In another embodiment, a method for determining whether a new input is out-of-domain for a convolutional neural network (CNN) is provided, the method comprising: enabling the CNN to generate an FMCM for each input in addition to predicting an output for each input. The CNN is further capable of generating, for each CNN layer, the FMCM of each input, linearizing each FMCM of a given CNN layer into a vector to represent a point in the feature map covariance space of the CNN, and calculating the FMCM of the layer for the given input based on the covariance between the plurality of feature maps of the given CNN layer for the given input. The method further comprises: constructing an out-of-domain classifier for the CNN based on the training feature map covariance space of a plurality of training FMCMs, and wherein the plurality of training FMCMs are generated by the CNN for a plurality of training inputs that have been used to train the CNN to predict an output for each input; enabling the out-of-domain classifier to run the new input to classify the new input as being in-domain or out-of-domain for the CNN based on whether the corresponding new FMCM is within or outside the training feature map covariance space.

[0017] In one embodiment, the method includes the step of enabling the CNN to calculate the FMCM of a given CNN layer for a given input by computing the Gram matrix between each pair of feature maps among the corresponding plurality of feature maps.

[0018] In one embodiment, the method includes the step of enabling an out-of-domain classifier to label a new input when the new input is determined to be out of the domain of the CNN.

[0019] In one embodiment, the method includes the step of enabling the CNN to reduce the dimension of each FMCM based on principal component analysis (PCA) to form a training feature map covariance space.

[0020] In one embodiment, each new input includes a microscopic image. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] With reference to the drawings, the present invention will be more clearly understood from the following description of embodiments of the present invention given by way of example only, in which:

[0022] Figure 1 There is shown a system for determining whether a new input is out of the domain of a CNN according to an embodiment of the present invention;

[0023] Figure 2 is a flowchart showing a method for determining whether a new input is out of the domain of a CNN according to an embodiment of the present invention;

[0024] Figure 3 There is shown an exemplary PCA reduction output corresponding to the training FMCM;

[0025] Figure 4 There is shown an exemplary feature map covariance space of the first layer of the CNN;

[0026] Figure 5 There is shown an exemplary training feature map covariance space formed by plotting the first 2 PCA components;

[0027] Figure 6 There is shown an exemplary "out-of-domain" classifier in a space with two style components;

[0028] Figure 7 There is shown the FMCM similarity for predicting the performance of the CNN on new input data;

[0029] Figure 8A There is shown an image in the style space for examining the effect of image blur on the performance of a nucleus detector model;

[0030] Figure 8B There is shown the correlation between the blur level and the model performance; and

[0031] Figure 9A shows the original image, and Figure 9B shows the original image with a blur level of 0.8. Detailed implementation mode

[0032] Figure 1 shows a system 100 for determining whether a new input is outside the domain of a CNN 102 according to an embodiment of the present invention. When the new input is outside the domain, it means that the CNN 102 will extrapolate data from outside the domain to an answer for the data used to train the CNN 102. In the context of the present invention, the system 100 can be implemented in a server or another computer system, and is implemented by a processor (e.g., a single or multiple processors) or other hardware described herein. These methods, functions, and other processes can be embodied as machine-readable instructions stored on a computer-readable medium, which can be non-transitory, such as a hardware storage device (e.g., RAM (Random Access Memory), ROM (Read-Only Memory), EPROM (Erasable, Programmable ROM), EEPROM (Electrically Erasable, Programmable ROM), hard disk drive, and flash memory).

[0033] The system 100 includes: a CNN 102 configured to predict an output for each input and generate an FMCM for each input; and an out-of-domain classifier 106 configured to classify a new input as being within or outside the domain of the CNN based on the FMCM of the new input during runtime. In an embodiment of the present invention, the CNN 102 can be an existing CNN modified to output an FMCM for each input in addition to the original output of the model, and can also be referred to as the modified CNN 102 hereinafter. The CNN 102 can be interchangeably referred to as a CNN model or a model.

[0034] In the context of the present invention, the input to the CNN 102 is typically an image, and the output can be a predicted image, a predicted binary output, or a predicted class. Additionally, multiple training inputs (also referred to as training images or training data points) can be used to train the CNN 102. The input to the CNN 102 can also be referred to as input data, data, data points, or input images.

[0035] Figure 2 is a flowchart showing a method for determining whether a new input is outside the domain of a CNN according to an embodiment of the present invention.

[0036] At step 202, the CNN 102 generates a plurality of training FMCMs for a plurality of training inputs. For each training input, the CNN 102 calculates the FMCM of the layer based on the covariance between a plurality of feature maps of a given CNN layer.

[0037] The feature map (FM) is the output of each convolutional kernel operation applied to the input from a specific layer of the CNN 102 at that specific layer. Typically, given the input from the previous layer, multiple feature maps (30+) of each layer are calculated.

[0038] FM1 = Ω(kernel1 * Output_layer -1 + B1)

[0039] FM2 = Ω(kernel2 * Output_layer -1 + B2)

[0040] FM3 = Ω(kernel3 * Output_layer -1 + B3)

[0041] Ω represents an activation function well-known to those skilled in the art, such as RELU or Sigmoid.

[0042] B is a scalar (bias)

[0043] * represents the convolution operation

[0044] The calculation of the feature map is an intermediate calculation within any convolutional CNN model architecture and is necessary for making model predictions. That is, this will occur within the model architecture anyway and does not require the CNN 102 to perform any additional calculations.

[0045] In an embodiment of the present invention, the CNN 102 calculates the FMCM of a layer for a training input by calculating the Gram matrix of the feature map at a given CNN layer, such that the Gram matrix element (1, 3) is the covariance between features Figure 1 and 3 More specifically, the FMCM for a training input is the following calculation of a given (width) × (height) × (number of feature maps) dimensional tensor:

[0046]

[0047] Here, G is the FMCM, and F is the (width) × (height) × (feature map) tensor output from the model layer

[0048] It is clear to those skilled in the art that the above function can be directly incorporated into the CNN 102, thereby allowing the simultaneous output of model predictions (e.g., class or another image) and the FMCM for that input. Thus, for each data point (e.g., image) used to train the CNN 102, the CNN 102 outputs the corresponding training FMCM.

[0049] Note that the feature maps are spatially invariant and robust to translations in the image, thus giving a compact representation of the input data. Additionally, the FMCM has a much lower dimension than the entire latent space of the CNN. For example, if there are 60 feature maps in a CNN layer, the FMCM will be only 60×60, which is much lower than the dimension of the feature maps themselves. Therefore, the training FMCMs generated based on the feature maps of the training inputs are used as a compact, model-centric representation of the training data that has already been used to train the CNN 102.

[0050] At step 204, a training feature map covariance space is formed based on a plurality of training FMCMs. In an embodiment of the present invention, each training FMCM is linearized into a vector to represent a point in an N-dimensional space, which N-dimensional space can be considered a model-specific "style space" S. This style space may also be interchangeably referred to hereinafter as the training feature map covariance space, where the FMCM vectors form the style space vectors.

[0051] For example, training data points → CNN → model_output S

[0052] where S is where d depends only on the number of feature maps. The set of S vectors forms a "style similarity space" or training feature map covariance space. The distribution of the training data within this space can be easily bounded by many techniques well-known to those skilled in the art, such as one-class SVM or fitting a parametric distribution. The covariance matrix elements are typically correlated, and thus the dimension of the training FMCMs can be further reduced by standard techniques such as principal component analysis (PCA).

[0053] Figure 3 Shows an exemplary PCA-reduced output corresponding to the training FMCMs of an exemplary single-cell CNN classifier constructed to determine whether a mammalian cell is alive or dead. The input to the CNN 102 is an image of a single cell, and the output is binary "dead" or "alive". Another example is using an image-to-image CNN 102 constructed to label cell nuclei within an image. The three axes of space 300 represent the first 3 PCA components, and the subspace 302 represents those components encapsulated by the CNN 102. The yellow dots represent those components enclosed by a one-class SVM. This is only an example classifier constructed to encapsulate the model training data.

[0054] Figure 4An exemplary feature map covariance space 400 for the first layer of the CNN 102 is shown. For viewing purposes, the dimension has been reduced via PCA. The first 3 components of the PCA are shown. The feature map covariance space 400 represents the training data 402 and the new data 404 to be predicted. This is an illustration of how it is easy to separate "out-of-domain" new data (blue) from the training data of the model in the "style space".

[0055] Figure 5 An exemplary training feature map covariance space 502 formed by plotting the first 2 PCA components of the FMCM vectors is shown. The training data is plotted on a scale from purple to blue (highest accuracy to lowest accuracy), and the FMCM 504 of the new data set is plotted from yellow to red (highest accuracy to lowest accuracy). It is obvious that as the new data points move away from the training data, the accuracy of the prediction decreases.

[0056] It is observed that by examining the distribution of the training data in the training feature map covariance space, it is possible to check whether the new input / data has a "style" similar enough to the training data. It is also empirically observed that new inputs in different style spaces are outside the domain of the CNN 102, and thus, the output of the CNN based on this new input is not trustworthy.

[0057] Returning to Figure 2 , at step 206, an out-of-domain classifier 106 is constructed for the training feature map covariance space to classify the FMCM as being within or outside the training feature map covariance space. Out-of-domain classifiers are well known to those skilled in the art. The out-of-domain classifier 106 is constructed based on the training FMCM and is used to predict whether a new FMCM is within or outside the domain relative to the training FMCM.

[0058] Figure 6 An exemplary "out-of-domain" classifier on a space with two style components is shown, i.e., the style space vector is two-dimensional. The style space vector is the FMCM vector. Relative to the training data, the first space 602 corresponds to in-domain data, and the second space 604 represents out-of-domain data.

[0059] In this example, the "out-of-domain" classifier is

[0060]

[0061] Here, 1 and 0 correspond to "out-of-domain" and "in-domain" respectively.

[0062] At step 208, the out-of-domain classifier 106 operates on the new input to classify the new input as within or outside the domain of the CNN based on whether the corresponding new FMCM is within or outside the training feature map covariance space. The out-of-domain classifier 106 receives the new FMCM for the new input from the CNN 102 and classifies the new FMCM as within or outside the training feature map covariance space. When the new FMCM is outside the training feature map covariance space, it can be determined that the new input is outside the domain of the CNN 102. When the new FMCM is within the training feature map covariance space, it can be determined that the new input is within the domain of the CNN 102. The out-of-domain classifier 106 is also configured to mark the new input and issue a warning to the user when it is found that the new input is outside the domain. When the new input data point is classified as outside the training distribution in the style space, the new input data point can be considered to be outside the domain of the CNN model, and thus the output of the CNN model 102 is not trustworthy.

[0063] Figure 7 Shows the FMCM similarity for predicting the performance 700 of the CNN 102 on new input data. For each new data set, the proportion of data outside the training data in the training feature map covariance space is calculated and compared with the accuracy of the CNN on the corresponding data set. It can be clearly seen that there is a strong relationship between the model accuracy and the similarity of the training data in the training feature map covariance space.

[0064] In an embodiment of the present invention, the CNN 102 is trained based on microscopic images and shows that if the image has a different style from the training data, it is less likely to succeed within the network.

[0065] Figure 8A Shows images in the style space for examining the effect of image blur on the performance of the nucleus detector model. The images are deliberately defocused using Gaussian blur to label dissimilar images using style similarity. The blue dots represent the 102 training data of the CNN. Orange 801 is the blur sigma level of 0.3. Green 802 is the blur sigma level of 0.4. Red 803 is the blur sigma level of 0.5. Purple 804 is the blur sigma level of 0.8.

[0066] Figure 8B Shows how the blur level is related to the model performance. It is evident that as the new data moves away from the training data in the style space, the performance of the CNN 102 degrades.

[0067] Figure 9A Shows the original image, and Figure 9BThe original image with a blur level of 0.8 is shown. The purple 901 represents the model prediction of the cell nucleus, and the pink is the actual cell nucleus. It can be seen that at the blur level of 0.8, the specificity of CNN 102 decreases (false positives). Although most microscopic images are given as examples here, this technique applies to any convolutional neural network. For example, if the training data is always acquired on sunny days, the model may have difficulty adapting to a style shift, such as changing to a rainy day.

[0068] The embodiments described with reference to the accompanying drawings in the present invention include computer devices and / or processes executed in computer devices. However, the present invention also extends to computer programs suitable for putting the present invention into practice, particularly computer programs stored on or in a carrier. The program may be in the form of source code, object code, or intermediate code between source code and object code, for example, in a partially compiled form, or in any other form suitable for use in implementing the method according to the present invention. The carrier may include a storage medium such as a ROM, for example, a memory stick or a hard disk. The carrier may be an electrical or optical signal that can be transmitted via a cable or an optical cable or by radio or other means.

[0069] In this specification, the terms "comprise, comprises, comprised, comprising" or any variation thereof and the terms "include, includes, included, including" or any variation thereof are considered to be completely interchangeable, and they should all be given the widest possible interpretation, and vice versa.

[0070] The present invention is not limited to the embodiments described previously herein, but may be different in both construction and detail.

Claims

1. A system for determining whether a new input is outside the domain of a Convolutional Neural Network (CNN), comprising: The CNN, configured to: Predict an output for each input; And For each CNN layer, generate a Feature Map Covariance Matrix (FMCM) for each input; Linearize each FMCM of a given CNN layer into a vector to represent a point in the feature map covariance space of that CNN layer; Calculate the FMCM of the layer for the given input based on the covariance between multiple feature maps of the given CNN layer for the given input; And An out-of-domain classifier, the out-of-domain classifier being constructed based on the training feature map covariance space of multiple training FMCMs, and configured to: run a new input to classify the new input as being within or outside the domain of the CNN based on whether the corresponding new FMCM is within or outside the training feature map covariance space, and wherein the multiple training FMCMs are generated by the CNN for multiple training inputs that have been used to train the CNN to predict an output for each input.

2. The system according to claim 1, wherein, The CNN is further configured to calculate the FMCM of the given CNN layer for the given input by calculating the Gram matrix between each pair of feature maps in the corresponding multiple feature maps.

3. The system according to any one of the preceding claims, wherein, The out-of-domain classifier is further configured to: label the new input when the new input is classified as being outside the domain of the CNN.

4. The system according to any one of the preceding claims, wherein, The CNN is configured to: reduce the dimension of each FMCM based on Principal Component Analysis (PCA) to form the training feature map covariance space.

5. A method for determining whether a new input is outside the domain of a Convolutional Neural Network (CNN), comprising: Enabling the CNN to: In addition to predicting an output for each input, also generate an FMCM for each input; For each CNN layer, generate an FMCM for each input and linearize each FMCM of a given CNN layer into a vector to represent a point in the feature map covariance space of that CNN layer; And Calculate the FMCM of the layer for the given input based on the covariance between multiple feature maps of the given CNN layer for the given input; Construct an out-of-domain classifier for the CNN based on the training feature map covariance space of multiple training FMCMs, and wherein the multiple training FMCMs are generated by the CNN for multiple training inputs that have been used to train the CNN to predict an output for each input; And Enabling the out-of-domain classifier to run the new input to classify the new input as being within or outside the domain of the CNN based on whether the corresponding new FMCM is within or outside the training feature map covariance space.

6. The method according to claim 5 further comprises: Enabling the CNN to calculate the FMCM of the given CNN layer for the given input by calculating the Gram matrix between each pair of feature maps in the corresponding multiple feature maps.

7. The method according to claim 6, further comprising: Enabling the out-of-domain classifier to label the new input when the new input is determined to be outside the domain of the CNN.

8. The method according to claim 6, further comprising: Enable the CNN to reduce the dimension of each FMCM based on principal component analysis (PCA) to form the training feature map covariance space.

Citation Information

Patent Citations

  • Distances between distributions for the belonging-to-the-distribution measurement of the image

    WO2022045915A1

  • Automated medical image quality control system

    WO2022108912A1