Machine Learning for Automatic Detection of Intracranial Hemorrhage Using Uncertainty Measures from CT Images
By generating probability scores and uncertainty measures related to intracranial hemorrhage detection, the problem that machine learning systems cannot explain the evaluation results is solved, reliable clinical decision support is achieved, and the reliability and efficiency of intracranial hemorrhage detection is improved.
Patent Information
- Application Number
- CN202210243754.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-12
- Filing Date
- 2022-03-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-03-11
AI Technical Summary
Existing machine learning systems cannot interpret the evaluation results when detecting intracranial hemorrhage, resulting in clinicians being unable to reliably rely on their evaluation results and being unable to incorporate them into the clinical workflow.
Through a machine learning-based network to generate probability scores associated with medical imaging analysis tasks and determine uncertainty measures associated with probability scores, uncertainty measures are calculated using calibration functions and entropy of beta distributions, combined with Dempster-Shafer uncertainty measures, providing reliable clinical decision support.
Provides reliable uncertainty measures to help clinicians understand the uncertainty of testing, typing and segmentation intracranial hemorrhage, thereby making reliable clinical decisions and improving the reliability and efficiency of testing.
Smart Images

Figure CN115082579B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to detection of intracranial hemorrhage, and in particular to machine learning for automatic detection of intracranial hemorrhage using uncertainty measures from CT (computed tomography) images. Background Art
[0002] Non-contrast CT (computed tomography) imaging is the most common form of medical imaging used to evaluate urgent and emergent neurological conditions. In current clinical practice, non-contrast CT imaging is the standard of care for the evaluation of acute stroke and head trauma and is particularly useful for detecting acute intracranial hemorrhage (ICH). Minimizing the interpretation time and intervention time of CT imaging is crucial for patient outcomes.
[0003] Recently, machine learning systems have been proposed for automatically evaluating CT images for ICH and other medical conditions. However, such machine learning systems are unable to interpret the evaluation results. While machine learning systems typically generate a score representing the evaluation result, such a score itself can be unreliable. Consequently, clinicians cannot rely on the evaluation results from such machine learning systems and, therefore, cannot reliably incorporate such machine learning systems into clinical workflows. Summary of the Invention
[0004] According to one or more embodiments, a system and method for performing a medical imaging analysis task to facilitate clinical decision making is provided. One or more input medical images of a patient are received. A medical imaging analysis task is performed based on the one or more input medical images using a machine learning-based network. The machine learning-based network generates a probability score associated with the medical imaging analysis task. An uncertainty measure associated with the probability score is determined. A clinical decision is made based on the probability score and the uncertainty measure.
[0005] In one embodiment, the uncertainty measure is determined by applying a calibration function to the probability scores and calculating the entropy of the probability scores based on the results of the applied calibration function. In another embodiment, a machine learning-based network generates parameters of a Beta distribution of the scores, and the uncertainty measure is determined by calculating the entropy of the Beta distribution based on the parameters and combining the entropy of the Beta distribution with the entropy of the probability scores.
[0006] In one embodiment, making a clinical decision includes stratifying a patient into one of a plurality of patient groups based on a probability score and an uncertainty measure. In a case where the medical imaging analysis task includes detecting intracranial hemorrhage in a patient, the plurality of patient groups may include a high confidence positive detection patient group, a high confidence negative detection patient group, and a low confidence patient group. In another embodiment, making a clinical decision includes determining whether to treat the patient based on the probability score and the uncertainty measure. In another embodiment, making a clinical decision includes determining whether to perform a clinical test on the patient based on the probability score and the uncertainty measure. In another embodiment, making a clinical decision includes prioritizing a radiologist's work list based on the probability score and the uncertainty measure.
[0007] In one embodiment, the medical imaging analysis task includes at least one of detection, classification, or segmentation of intracranial hemorrhage in a patient.
[0008] These and other advantages of the present invention will become apparent to those of ordinary skill in the art upon reference to the following detailed description and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A method for making clinical decisions based on automatically performed medical imaging analysis tasks according to one or more embodiments is shown;
[0010] Figure 2 shows a network architecture of a deep learning network trained for detecting, segmenting, and classifying intracranial hemorrhage according to one or more embodiments;
[0011] Figure 3 shows a method according to one or more embodiments Figure 2 Network architecture details of various networks in the network architecture;
[0012] Figure 4 An exemplary artificial neural network that can be used to implement one or more embodiments is shown;
[0013] Figure 5 shows a convolutional neural network that can be used to implement one or more embodiments; and
[0014] Figure 6 A high-level block diagram of a computer that can be used to implement one or more embodiments is shown. DETAILED DESCRIPTION
[0015] The present invention generally relates to methods and systems for machine learning to automatically detect intracranial hemorrhage using uncertainty metrics from CT (computed tomography) images. Embodiments of the present invention are described herein to provide an intuitive understanding of such methods and systems. Digital images typically consist of digital representations of one or more objects (or shapes). Digital representations of objects are generally described herein in terms of identifying and manipulating the objects. Such manipulations are virtual manipulations performed within the computer system's memory or other circuitry / hardware. Therefore, it should be understood that embodiments of the present invention can be executed within a computer system using data stored within the computer system.
[0016] Embodiments described herein provide for automated detection, classification, and segmentation of intracranial hemorrhage (ICH) on non-contrast head CT images using a machine learning-based network. The machine learning-based network generates a probability score (e.g., a classification score) for the detection and classification of ICH. According to embodiments described herein, an uncertainty measure associated with the probability score is determined. Advantageously, the uncertainty measure associated with the probability score reliably informs clinicians of the uncertainty associated with the detection, classification, and segmentation of ICH, thereby enabling reliable clinical decision making based on the probability score and the uncertainty measure.
[0017] Figure 1 A method 100 for making clinical decisions based on automatically performed medical imaging analysis tasks according to one or more embodiments is shown. The steps of the method 100 may be performed by one or more suitable computing devices, such as, for example, Figure 6 Computer 602.
[0018] At step 102, one or more input medical images of a patient are received. The one or more input medical images may be images of any anatomical object of interest to the patient, such as, for example, an organ, bone, lesion, etc. In one example, the one or more input medical images are images of the patient's head.
[0019] In one embodiment, the one or more input medical images include CT images, such as, for example, non-contrast CT images. However, the one or more input medical images may include any other suitable modality, such as, for example, MRI (magnetic resonance imaging), ultrasound, x-ray, or any other medical imaging modality or combination of medical imaging modalities. The one or more input medical images may include 2D (two-dimensional) images and / or 3D (three-dimensional) volumes and may include a single input medical image or multiple input medical images. In one embodiment, the one or more input medical images include 2.5D (2D plus time) images. The one or more input medical images may be received directly from an image acquisition device, such as, for example, a CT scanner, as the medical images are acquired, or may be received by loading previously acquired medical images from a storage device or memory of the computer system or receiving medical images that have been transmitted from a remote computer system.
[0020] At step 104, a medical imaging analysis task is performed based on one or more input medical images using a machine learning-based network. The machine learning-based network generates a probability score associated with the medical imaging analysis task. The medical imaging analysis task can be any one or more tasks performed on one or more medical images using any suitable machine learning-based network. For example, the medical imaging analysis task can include detection, classification, segmentation, etc. The probability score can represent the likelihood of a particular outcome of the medical imaging analysis task. In one example, if the medical imaging analysis task is classification, the probability score can be a classification score (e.g., 0.8) representing the likelihood of a particular classification.
[0021] In one embodiment, the medical imaging analysis task includes detecting, segmenting, and classifying ICH on a non-contrast CT input medical image of a patient's head. Figure 2A network architecture 200 of a deep learning network trained for detecting, segmenting, and classifying ICH according to one or more embodiments is shown. The network architecture 200 includes three blocks: 1) a feature extractor block 204 comprising an axial feature extractor 204-A and a coronal feature extractor 204-B; 2) a classifier block 206 comprising an axial classifier 206-A, a coronal classifier 206-B, and a common classifier 208; and 3) a segmentation block 214 comprising an axial segmenter 214-A and a coronal segmenter 214-B. The axial feature extractor 204-A and the coronal feature extractor 204-B extract features from an axial CT input medical image 202-A and a coronal CT input medical image 202-B, respectively. The extracted features are input to the axial classifier 206-A and the coronal classifier 206-B, and the axial segmenter 214-A and the coronal segmenter 214-B, respectively. The results of the axial classifier 206-A and the coronal classifier 206-B are combined by a common classifier 208 to generate a probability score 210. The probability score 210 includes a bleeding score, which indicates the probability of detecting an ICH, and an IPH (intraperitoneal hemorrhage) score, an SDH (subdural hemorrhage) score, an EDH (extradural hemorrhage) score, a SAH (subarachnoid hemorrhage) score, and an IVH (intraventricular hemorrhage) score, which indicate the probability of ICH classification. For example, by comparing the probability score 210 to one or more thresholds, the probability score 210 can be converted into a positive or negative (i.e., yes or no) result 212. The axial segmenter 214-A and the coronal segmenter 214-B generate an axial segmentation map 216-A and a coronal segmentation map 216-B, respectively, which are combined to generate a bleeding location map 218.
[0022] The deep learning network of the network architecture 200 is trained using training data during a previous offline or training phase. The network weights can be optimized in two phases: the feature extractor network and the segmenter network can first be optimized separately to minimize the voxel-wise binary cross entropy segmentation loss relative to the manual segmentation, and then all networks can be optimized together to minimize the linear combination of the segmentation loss and the study-level sigmoid binary cross entropy classification loss relative to the six manual labels for ICH and IPH, SDH, EDH, SAH, and IVH classification. Once trained, the trained deep learning network can be used during an online or testing phase (e.g., in Figure 4 104 ) is applied.
[0023] Figure 3 shows a method according to one or more embodiments Figure 2 The network architecture details 300 of various networks in the network architecture 200 will be continued. Figure 2 The network architecture 200 is described Figure 3Network architecture details 300. Axial feature extractor 204-A and coronal feature extractor 204-B can be implemented as DenseUNet 302. Axial classifier 206-A and coronal classifier 206-B can be implemented as axial / coronal classifier 304. Common classifier 208 can be implemented as common classifier 306. The annotated arrows in network architecture details 300 indicate the number of channels in each feature map. Strided convolution (SConv) with a stride of 2 results in the feature map size being halved within a plane (for 2D) or across all dimensions (for 3D). Deconvolution (Deconv) results in the feature map size being doubled. All convolutions use a kernel of size 3. The axial CT input medical image 202 -A and the coronal CT input medical image 202 -B are preprocessed to have a voxel spacing of 1 mm×1 mm×4 mm and 1 mm×4 mm×1 mm, and are rescaled so that the range 0-1 is mapped to a wide window with a window level WL=55 HU (Hounsfield unit) and a window width WW=200 HU.
[0024] Back to Figure 1 At step 106, an uncertainty metric associated with the probability score is determined. The uncertainty metric can be any suitable metric that represents the uncertainty associated with the probability score. For example, the uncertainty metric can represent the error associated with the probability score. In this example, the probability score is 0.8, and the uncertainty metric can be plus / mins 0.1.
[0025] In one embodiment, the uncertainty measure is a calibrated classifier uncertainty measure. The calibrated classifier uncertainty measure considers the entropy of the probability scores determined by the machine learning based network. The calibration function is selected from a family of monotonically increasing functions with three fixed points: <0, 0>, < d , 0.5 > and < 1, 1 >, where d is the decision threshold chosen by the user. All such calibration functions have the property ,in H is entropy, and p is the probability score (in Figure 1 In other words, the probability score that is exactly at the selected decision threshold is considered to be completely uncertain. In one embodiment, the calibration function is applied to the probability score to set the specified operating point threshold t Mapped to 0.5, as follows:
[0026] .
[0027] Other calibration functions may also be applied, such as, for example, the cumulative distribution function of the Kumaraswamy distribution. The entropy of the probability scores is then calculated based on the results of the applied calibration function as a calibration classifier uncertainty measure, as follows:
[0028] .
[0029] In one embodiment, the uncertainty metric is a Dempster-Shafer uncertainty metric that takes into account additional uncertainty estimated during training. In this embodiment, the machine learning based network generates distribution parameters of a beta distribution of scores, where the mean is the probability score and a narrower distribution indicates a higher confidence. The distribution parameters of the beta distribution of scores may be a belief quality representing the evidence for the classification label in the input medical image. For example, the distribution parameters of the beta distribution of scores may include a probability score representing the evidence for a positive classification label. α and evidence that negates the classification label β . The entropy of a beta distribution around an operating point threshold is calculated based on the distribution parameters as a measure of additional uncertainty. The entropy of the beta distribution is combined (e.g., added) with a calibrated classifier uncertainty metric (i.e., the entropy of the probability score) to determine a Dempster-Shafer uncertainty metric. In another embodiment, instead of the distribution parameters of the beta distribution, the machine learning-based network generates distribution parameters of a Kumaraswamy distribution, the entropy of the Kumaraswamy distribution around the operating point threshold is calculated based on the distribution parameters, and the entropy of the Kumaraswamy distribution is combined with the calibrated classifier uncertainty metric to determine the Dempster-Shafer uncertainty metric. The calculation of the Dempster-Shafer uncertainty metric is further described in U.S. Patent Publication No. 2020 / 0320354, filed on September 5, 2019, and U.S. Patent Application No. 17 / 072,424, filed on October 16, 2020, the disclosures of which are incorporated herein by reference in their entirety.
[0030] To generate the parameters of the beta distribution of scores, a machine learning based network can be trained using the typical binary cross entropy loss, and the final layer of the network can then be fine-tuned with a second loss that takes the following form:
[0031]
[0032] The first term measures the data fitting cost, and the second term is a regularization term for non-deterministic samples. y k represents the ground truth label, p k ExpressN The first of the training cases k predictions, is α and β parameterized Gamma function, and KL is the Kullback-Leibler divergence between the beta-distributed prior and the beta distribution with total uncertainty (i.e. α = β= 1). A calibration function that measures the entropy of the probability fraction exceeding the calibration operating point can be applied, which can be calculated from the cumulative distribution function of the parameterized beta distribution:
[0033]
[0034] in B is the beta function, x is the chosen operating point, p ds is the probability score, and E ds is the entropy of the probability score.
[0035] At step 108, the probability score and uncertainty measure are output. For example, the probability score and uncertainty measure can be output by displaying the probability score and uncertainty measure on a display device of the computer system, storing the probability score and uncertainty measure on a memory or storage device of the computer system, or transmitting the probability score and uncertainty measure to a remote computer system. In one embodiment, the probability score and uncertainty measure can be output to a system for use, for example, in making an automated clinical decision or further treatment of the patient.
[0036] At step 110, a clinical decision is made based on the probability score and the uncertainty measure. In one embodiment, the clinical decision includes stratifying the patient into one of a plurality of patient groups, for example, a three-way triage separating the patients into a high confidence positive classification (e.g., a positive test for ICH) patient group, a high confidence negative classification (e.g., a negative test for ICH) patient group, and a low confidence patient group. In another embodiment, the clinical decision includes determining whether to treat the patient. In another embodiment, the clinical decision includes determining whether to perform a clinical test on the patient. In another embodiment, the clinical decision includes prioritizing a radiologist's work list based on the probability score and the uncertainty measure. The clinical decision can be made based on one or more thresholds. The clinical decision can be any other suitable clinical decision.
[0037] Embodiments described herein provide an uncertainty measure associated with a probability score determined by a machine learning-based network. While the probability score determined by the machine learning-based network can represent the likelihood of a medical imaging analysis task, the probability score determined by the machine learning-based network may involve a level of uncertainty. Advantageously, the uncertainty measure calculated according to the embodiments described herein provides a level of uncertainty in the probability score, thereby enabling the results of the medical imaging analysis task to be reliably used to make clinical decisions.
[0038] The embodiments described herein were experimentally validated on 46,057 non-contrast head CT images acquired from ten centers using scanners from various manufacturers. 25,946 of those images were acquired from seven "visible" centers to develop and optimize the system, including iterative architecture selection, parameter tuning, and training. 400 of those images were acquired from three "unseen" centers from the RSNA (Radiological Society of North America) ICH Challenge training dataset to calibrate the system. The system's performance was then measured using 2,947 images from the "visible" centers without patient overlap and 16,764 images from the RSNA ICH Challenge training dataset.
[0039] Study-level labels for ICH and classification were generated in one of three ways. For the RSNA dataset, labels were provided with the dataset and assigned at the slice level by one of sixty radiologists who reviewed the images. For images from visible centers, for 6,649 images across three centers, study-level labels were provided by the contributing center based on review of the images or radiology reports by the center's radiologist. For 22,244 images across four centers, radiology reports were provided by the contributing center and manually transcribed by a trained team of annotators (with at least forty hours of training in CT hemorrhage annotation under the supervision of a radiologist) and reviewed by a technologist (with a bachelor's degree in radiology and medical imaging technology and one year of neuroimaging training) and one of three radiologists (with at least five years of experience). For a subset of 3,278 ICH-positive cases, acute / subacute ICH was manually segmented and reviewed by the same team.
[0040] Probability scores and uncertainty measures are determined according to the embodiments described herein. The uncertainty measures are evaluated by measuring ROC (receiver operating characteristic) metrics such as AUC (area under the curve), sensitivity, and specificity. The uncertainty measures are further evaluated by modeling the average RTAT (report turnaround time) for positive cases, comparing a three-way prioritized worklist (e.g., low uncertainty (i.e., high confidence) positive cases, low uncertainty negative cases, and high uncertainty cases) to a first-in, first-out worklist. Specifically, two simplifying assumptions are used to model radiologists reading a fixed worklist in either prioritized or non-prioritized sequence: 1) within each priority level, true bleeding cases are randomly placed in the queue, and 2) each case takes a fixed amount of time. T In this model, the mean RTAT t It can be written as follows:
[0041]
[0042] in, n 1. n 2. n 3 is the number of cases in the three priority levels, p 1. p 2. p 3 is the number of positive cases, and p is the prevalence in the population.
[0043] To evaluate the uncertainty measures used to identify high-confidence subsets with improved performance, the ROC metric and RTAT were compared on the full test set with the highest uncertainty metric at 80%. Differences in AUC were tested using the DeLong test of the associated ROC curves. The Youden index and RTAT were tested for differences in means across 1000 samples using bootstrap paired tests.
[0044] For bleeding detection, the ROCs from the visible and invisible centers were compared. The system yielded a ROC-AUC of 0.97 for data from the visible center and a ROC-AUC of 0.95 for data from the invisible center. At the selected operating point, the system yielded a sensitivity / specificity of 0.92 / 0.93 for data from the visible center and a sensitivity / specificity of 0.86 / 0.92 for data from the invisible center.
[0045] For bleeding segmentation, for less than and greater than 20 cm 3 The total amount of acute bleeding estimated by automatic and manual segmentation was compared. 3, the 95% agreement limit was 6 cm 3 , and for bleeding is 20-456cm 3 For patients with 3 .
[0046] For bleeding classification, the ROCs for data from the visible center and the invisible center were compared. The system yielded ROC-AUCs of 0.90, 0.92, 0.93, 0.92, and 0.96 for SAH, SDH, EDH, IPH, and IVH for data from the visible center and 0.88, 0.85, 0.77, 0.92, and 0.96 for data from the invisible center.
[0047] Performance coverage curves for area under the curve (AUC), sensitivity, and specificity were generated. Excluding the 20% lowest-confidence cases from the visible center using the calibrated classifier uncertainty metric improved the Youden index (sensitivity + specificity - 1) from 0.84 to 0.93 (p < 0.001) and from 0.78 to 0.88 (p < 0.001) from the invisible center using the calibrated classifier uncertainty metric. The Dempster-Shafer uncertainty metric improved the index from 0.84 to 0.92 (p < 0.001) from the visible center and from 0.78 to 0.89 (p < 0.001) from the invisible center using the calibrated classifier and the Dempster-Shafer uncertainty metric. The AUC for the visible center improved from 0.972 to 0.986 and 0.985, respectively. Excluding the 20% lowest-confidence cases improved the AUC from 0.945 to 0.963 and 0.962, respectively.
[0048] Performance coverage curves for three-way triage were generated. Using the calibrated classifier uncertainty metric, triage with an intermediate priority for the lowest confidence level of 20% improved the simulated RTAT from visible centers from a baseline of 20% to 15% (p < 0.001) and from unseen centers from 26% to 20% (p < 0.001). Using the Dempster-Shafer uncertainty metric, it improved the simulated RTAT from visible centers from a baseline of 20% to 15% (p < 0.001) and from unseen centers from 26% to 19% (p < 0.001).
[0049] The embodiments described herein are described with respect to the claimed systems and with respect to the claimed methods. Features, advantages, or alternative embodiments herein may be assigned to other claimed objects, and vice versa. In other words, a claim for a system may be modified using features described or claimed in the context of a method. In this case, the functional features of the method are embodied by the target unit providing the system.
[0050] Furthermore, certain embodiments described herein are described with respect to methods and systems for utilizing trained machine learning-based networks (or models), and with respect to methods and systems for training machine learning-based networks. Features, advantages, or alternative embodiments herein may be assigned to other claimed subject matter, and vice versa. In other words, claims to methods and systems for training machine learning-based networks may be improved with features described or claimed in the context of methods and systems for utilizing trained machine learning-based networks, and vice versa.
[0051] In particular, the trained machine learning-based networks used in the embodiments described herein can be adapted by methods and systems for training machine learning-based networks. Furthermore, the input data of the trained machine learning-based networks can include advantageous features and embodiments of the training input data, and vice versa. Furthermore, the output data of the trained machine learning-based networks can include advantageous features and embodiments of the output training data, and vice versa.
[0052] Generally speaking, trained machine learning-based networks mimic the cognitive functions that humans associate with other human minds. In particular, through training based on training data, trained machine learning-based networks are able to adapt to new environments and detect and infer patterns.
[0053] In general, the parameters of a machine learning-based network can be adapted using training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning, and / or active learning can be used. Furthermore, representation learning (an alternative term is "feature learning") can be used. In particular, the parameters of a trained machine learning-based network can be iteratively adapted over several training steps.
[0054] In particular, the trained machine learning-based network may include a neural network, a support vector machine, a decision tree, and / or a Bayesian network, and / or the trained machine learning-based network may be based on k-means clustering, Q-learning, a genetic algorithm, and / or association rules. In particular, the neural network may be a deep neural network, a convolutional neural network, or a convolutional deep neural network. Furthermore, the neural network may be an adversarial network, a deep adversarial network, and / or a generative adversarial network.
[0055] Figure 4 An embodiment of an artificial neural network 400 according to one or more embodiments is shown. Alternative terms for "artificial neural network" are "neural network," "artificial neural network," or "neural network." The artificial neural network 400 can be used to implement the machine learning networks described herein, such as in Figure 1 The machine learning-based network or Figure 2 Or the machine learning-based network shown in 3.
[0056] Artificial neural network 400 includes nodes 402-422 and edges 432, 434, ..., 436, wherein each edge 432, 434, ..., 436 is a directed connection from a first node 402-422 to a second node 402-422. Generally speaking, first node 402-422 and second node 402-422 are different nodes 402-422, but it is also possible that first node 402-422 and second node 402-422 are the same. For example, in Figure 4 , edge 432 is a directed connection from node 402 to node 406, and edge 434 is a directed connection from node 404 to node 406. Edges 432, 434, ..., 436 from a first node 402-422 to a second node 402-422 are also labeled as "incoming edges" of the second node 402-422 and "outgoing edges" of the first node 402-422.
[0057] In this embodiment, the nodes 402-422 of the artificial neural network 400 may be arranged in layers 424-430, wherein the layers may include an inherent order introduced by edges 432, 434, ..., 436 between the nodes 402-422. In particular, edges 432, 434, ..., 436 may only exist between adjacent node layers. Figure 4 In the embodiment shown in FIG, there is an input layer 424 that includes only nodes 402 and 404 and no incoming edges, an output layer 430 that includes only nodes 422 and no outgoing edges, and hidden layers 426, 428 between the input layer 424 and the output layer 430. In general, the number of hidden layers 426, 428 can be chosen arbitrarily. The number of nodes 402 and 404 within the input layer 424 is generally related to the number of input values of the neural network 400, and the number of nodes 422 within the output layer 430 is generally related to the number of output values of the neural network 400.
[0058] In particular, a (real) number may be assigned as a value to each node 402-422 of the neural network 400. Here, Indicates the value of the i-th node 402-422 of the n-th layer 424-430. The value of the nodes 402-422 of the input layer 424 corresponds to the input value of the neural network 400, and the value of the node 422 of the output layer 430 corresponds to the output value of the neural network 400. In addition, each edge 432, 434, ..., 436 may include a weight as a real number, in particular, the weight is a real number in the interval [-1, 1] or in the interval [0, 1]. Here, The weight of the edge between the i-th node 402-422 of the m-th layer 424-430 and the j-th node 402-422 of the n-th layer 424-430 is indicated. For weight is defined.
[0059] Specifically, to calculate the output value of the neural network 400, the input value is propagated through the neural network. Specifically, the values of the nodes 402-422 of the (n+1)th layer 424-430 can be calculated based on the values of the nodes 402-422 of the nth layer 424-430 by the following formula:
[0060] .
[0061] In this article, function f is a transfer function (another term is "activation function"). Known transfer functions are step functions, sigmoid functions (e.g., logistic function, generalized logistic function, hyperbolic tangent, inverse tangent function, error function, smoothed step function), or rectifier functions. Transfer functions are primarily used to introduce nonlinearity into a composite function and may be used for normalization purposes.
[0062] In particular, values are propagated layer by layer through the neural network, where the value of the input layer 424 is given by the input of the neural network 400, where the value of the first hidden layer 426 can be calculated based on the value of the input layer 424 of the neural network, where the value of the second hidden layer 428 can be calculated based on the value of the first hidden layer 426, and so on.
[0063] To set the edge value , training data must be used to train the neural network 400. In particular, the training data includes training input data and training output data (denoted as t i ). For the training step, the neural network 400 is applied to the training input data to generate the calculated output data. In particular, the training data and the calculated output data include a number of values, the number of which is equal to the number of nodes in the output layer.
[0064] In particular, the comparison between the calculated output data and the training data is used to recursively adapt the weights within the neural network 400 (back-propagation algorithm). In particular, the weights are changed according to
[0065]
[0066] where γ is the learning rate, and if the (n+1)th layer is not the output layer, then the number Can be based on Recursively computed as
[0067]
[0068] And if the (n+1)th layer is the output layer 430, then it is calculated as
[0069]
[0070] where f' is the first derivative of the activation function, and y (n+1) j is the comparative training value of the j-th node of the output layer 430.
[0071] Figure 5 A convolutional neural network 500 is shown in accordance with one or more embodiments. The convolutional neural network 500 may be used to implement the machine learning networks described herein, such as in Figure 1 The machine learning-based network or Figure 2 Or the machine learning-based network shown in 3.
[0072] exist Figure 5 In the embodiment shown in FIG, the convolutional neural network 500 includes an input layer 502, a convolutional layer 504, a pooling layer 506, a fully connected layer 508, and an output layer 510. Alternatively, the convolutional neural network 500 may include several convolutional layers 504, several pooling layers 506, and several fully connected layers 508, as well as other types of layers. The order of the layers can be selected arbitrarily, and the fully connected layer 508 is usually used as the last layer before the output layer 510.
[0073] In particular, within the convolutional neural network 500, the nodes 512-520 of a layer 502-510 can be considered to be arranged as a d-dimensional matrix or a d-dimensional image. In particular, in the two-dimensional case, the values of the nodes 512-520 indexed by i and j in the nth layer 502-510 can be represented as However, the arrangement of the nodes 512-520 of a layer 502-510 as such has no effect on the computations performed within the convolutional neural network 500, as these are given solely by the structure and weights of the edges.
[0074] In particular, the convolution layer 504 is characterized by the structure and weights of the incoming edges that form the convolution operation based on a specific number of kernels. In particular, the structure and weights of the incoming edges are selected so that the value of the node 514 of the convolution layer 504 is is calculated as the value x of the node 512 based on the previous layer 502 (n-1) Convolution , where convolution* is defined in two dimensions as
[0075] .
[0076] Here, the kth kernel K k is a d-dimensional matrix (a two-dimensional matrix in this embodiment) that is typically small compared to the number of nodes 512-518 (e.g., a 3×3 matrix or a 5×5 matrix). Specifically, this implies that the weights of the incoming edges are not independent, but are chosen so that they produce the convolution equation. In particular, for a kernel that is a 3×3 matrix, there are only 9 independent weights (one for each entry in the kernel matrix), regardless of the number of nodes 512-520 in the corresponding layer 502-510. In particular, for a convolutional layer 504, the number of nodes 514 in the convolutional layer is equal to the number of nodes 512 in the previous layer 502 multiplied by the number of kernels.
[0077] If the nodes 512 of the previous layer 502 are arranged as a d-dimensional matrix, using multiple kernels can be interpreted as adding another dimension (denoted as the "depth" dimension), so that the nodes 514 of the convolutional layer 504 are arranged as a (d+1)-dimensional matrix. If the nodes 512 of the previous layer 502 are already arranged as a (d+1)-dimensional matrix including the depth dimension, using multiple kernels can be interpreted as expanding along the depth dimension, so that the nodes 514 of the convolutional layer 504 are also arranged as a (d+1)-dimensional matrix, where the size of the (d+1)-dimensional matrix with respect to the depth dimension is the number of kernels times the size of the (d+1)-dimensional matrix with respect to the depth dimension in the previous layer 502.
[0078] The advantage of using convolutional layers 504 is that spatially local correlations in the input data can be exploited by enforcing local connectivity patterns between nodes in adjacent layers, in particular by having each node connected only to a small region of nodes in the previous layer.
[0079] exist Figure 5 In the embodiment shown in , the input layer 502 includes 36 nodes 512 arranged in a two-dimensional 6×6 matrix. The convolution layer 504 includes 72 nodes 514 arranged in two two-dimensional 6×6 matrices, each of which is the result of convolving the values of the input layer with the kernel. Equivalently, the nodes 514 of the convolution layer 504 can be interpreted as being arranged in a three-dimensional 6×6×2 matrix, where the last dimension is the depth dimension.
[0080] The pooling layer 506 can be characterized by the structure and weights of the incoming edges and the activation functions of its nodes 516, which form a pooling operation based on a nonlinear pooling function f. For example, in the two-dimensional case, the value x of the node 516 of the pooling layer 506 is (n) It can be based on the value x of the node 514 of the previous layer 504 (n-1) To calculate
[0081] .
[0082] In other words, by using the pooling layer 506, the number of nodes 514, 516 can be reduced by replacing the number of adjacent nodes 514 in the previous layer 504 with a single node 516 calculated as a function of the value of said number of adjacent nodes in the pooling layer. In particular, the pooling function f can be a maximum function, an average, or an L2 norm. In particular, for the pooling layer 506, the weights of the incoming edges are fixed and not modified by training.
[0083] The advantage of using the pooling layer 506 is that the number of nodes 514, 516 and the number of parameters are reduced. This results in a reduction in the amount of computation in the network and controls overfitting.
[0084] exist Figure 5 In the embodiment shown in , pooling layer 506 is max pooling, which replaces four adjacent nodes with only one node whose value is the maximum of the four adjacent nodes. Max pooling is applied to each d-dimensional matrix of the previous layer; in this embodiment, max pooling is applied to each of the two 2-dimensional matrices, reducing the number of nodes from 72 to 18.
[0085] The fully connected layer 508 is characterized by the fact that most, in particular all, edges exist between the nodes 516 of the previous layer 506 and the nodes 518 of the fully connected layer 508 , and wherein the weight of each edge can be adjusted individually.
[0086] In this embodiment, the nodes 516 of the previous layer 506 of the fully connected layer 508 are each displayed as a two-dimensional matrix, and are additionally displayed as unrelated nodes (indicated as a row of nodes, where the number of nodes is reduced for better presentation). In this embodiment, the number of nodes 518 in the fully connected layer 508 is equal to the number of nodes 516 in the previous layer 506. Alternatively, the number of nodes 516, 518 may be different.
[0087] Furthermore, in this embodiment, the values of the nodes 520 of the output layer 510 are determined by applying the Softmax function to the values of the nodes 518 of the previous layer 508. By applying the Softmax function, the sum of the values of all the nodes 520 of the output layer 510 is 1, and all the values of all the nodes 520 of the output layer are real numbers between 0 and 1.
[0088] Convolutional neural network 500 may also include a ReLU (Rectified Linear Unit) layer, which is an activation layer with a nonlinear transfer function. Specifically, the number and structure of nodes in a ReLU layer are equal to those in the previous layer. Specifically, the value of each node in a ReLU layer is calculated by applying a rectification function to the value of the corresponding node in the previous layer.
[0089] The inputs and outputs of different CNN blocks can be wired using summation (residual / dense neural networks), element-wise multiplication (attention), or other differentiable operators. Therefore, if the entire pipeline is differentiable, the CNN architecture can be nested rather than sequential.
[0090] In particular, the convolutional neural network 500 can be trained based on a backpropagation algorithm. To prevent overfitting, regularization methods can be used, such as dropout of nodes 512-520, random pooling, the use of artificial data, weight decay based on L1 or L2 norms, or maximum norm constraints. Different loss functions can be combined to train the same neural network to reflect joint training objectives. A subset of neural network parameters can be excluded from optimization to retain weights pre-trained on another dataset.
[0091] The systems, apparatus, and methods described herein can be implemented using digital circuitry, or using one or more computers using well-known computer processors, memory units, storage devices, computer software, and other components. Typically, a computer includes a processor for executing instructions and one or more memories for storing instructions and data. A computer may also include or be coupled to one or more mass storage devices, such as one or more magnetic disks, internal hard disks and removable disks, magneto-optical disks, optical disks, and the like.
[0092] The systems, devices, and methods described herein can be implemented using computers operating in a client-server relationship. Typically, in such a system, the client computer is remotely located from the server computer and interacts via a network. The client-server relationship can be defined and controlled by computer programs running on the respective client and server computers.
[0093] The systems, devices, and methods described herein can be implemented within a network-based cloud computing system. In such a network-based cloud computing system, a server or another processor connected to the network communicates with one or more client computers via the network. For example, a client computer can communicate with a server via a web browser application resident on and operating on the client computer. The client computer can store data on the server and access the data via the network. The client computer can transmit a request for data or a request for an online service to the server via the network. The server can perform the requested service and provide the data to the client computer(s). The server can also transmit data suitable for causing the client computer to perform a specific function, such as performing a calculation, displaying specified data on a screen, etc. For example, the server can transmit data suitable for causing the client computer to perform one or more steps or functions (including Figure 1 Some steps or functions of the methods and workflows described herein (including Figure 1 One or more steps or functions of the method and workflow described herein may be performed by a server or another processor in a network-based cloud computing system. Figure 1 One or more steps of the method and workflow described herein may be performed by a client computer in a network-based cloud computing system. Figure 1 One or more steps of the process may be performed by a server and / or a client computer in any combination in a network-based cloud computing system.
[0094] The systems, apparatus, and methods described herein may be implemented using a computer program product tangibly embodied in an information carrier, such as a non-transitory machine-readable storage device, for execution by a programmable processor; and the methods and workflow steps described herein (including Figure 1 One or more steps or functions of a program (e.g., a program that is executed by a processor) can be implemented using one or more computer programs that can be executed by such a processor. A computer program is a set of computer program instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0095] Figure 6A high-level block diagram of an example computer 602 that can be used to implement the systems, apparatus, and methods described herein is depicted in FIG. The computer 602 includes a processor 604 operatively coupled to a data storage device 612 and a memory 610. The processor 604 controls the overall operation of the computer 602 by executing computer program instructions that define such operations. The computer program instructions may be stored in the data storage device 612 or other computer-readable medium and loaded into the memory 610 when execution of the computer program instructions is desired. Thus, Figure 1 The methods and workflow steps or functions may be defined by computer program instructions stored in the memory 610 and / or data storage device 612 and controlled by the processor 604 executing the computer program instructions. For example, the computer program instructions may be implemented as computer executable code programmed by a person skilled in the art to perform Figure 1 Thus, by executing computer program instructions, the processor 604 performs Figure 1 The computer 602 may also include one or more network interfaces 606 for communicating with other devices via a network. The computer 602 may also include one or more input / output devices 608 that enable a user to interact with the computer 602 (e.g., a display, keyboard, mouse, speakers, buttons, etc.).
[0096] The processor 604 may include general-purpose and special-purpose microprocessors and may be the sole processor or one of multiple processors of the computer 602. For example, the processor 604 may include one or more central processing units (CPUs). The processor 604, the data storage device 612, and / or the memory 610 may include, be supplemented by, or incorporate one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs).
[0097] The data storage device 612 and the memory 610 each include a tangible, non-transitory computer-readable storage medium. The data storage device 612 and the memory 610 may each include high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate synchronous dynamic random access memory (DDR RAM), or other random access solid-state memory devices, and may include non-volatile memory, such as one or more magnetic disk storage devices, such as internal hard disks and removable disks, magneto-optical disk storage devices, optical disk storage devices, flash memory devices, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM), digital versatile disk read-only memory (DVD-ROM), or other non-volatile solid-state storage devices.
[0098] The input / output devices 608 may include peripheral devices such as a printer, a scanner, a display screen, etc. For example, the input / output devices 608 may include a display device such as a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor for displaying information to a user, a keyboard through which a user can provide input to the computer 602, and a pointing device such as a mouse or a trackball.
[0099] Image acquisition device 614 can be connected to computer 602 to input image data (e.g., medical images) into computer 602. It is possible to implement image acquisition device 614 and computer 602 as a single device. It is also possible for image acquisition device 614 and computer 602 to communicate wirelessly over a network. In a possible embodiment, computer 602 can be remotely located relative to image acquisition device 614.
[0100] Any or all of the systems and devices discussed herein may be implemented using one or more computers, such as computer 602 .
[0101] Those skilled in the art will recognize that actual computer or computer system implementations may have other structures and may include other components, and Figure 6 is a high-level representation of some components of such a computer for illustrative purposes.
[0102] The foregoing detailed description should be understood to be illustrative and exemplary in every respect and not restrictive, and the scope of the invention disclosed herein will not be determined by the detailed description, but by the claims interpreted in accordance with the full breadth permitted by patent law. It should be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention, and that various modifications may be implemented by those skilled in the art without departing from the scope and spirit of the invention. Various other feature combinations may be implemented by those skilled in the art without departing from the scope and spirit of the invention.
Claims
1. A computer-implemented method comprising: receiving one or more input medical images of a patient; performing a medical imaging analysis task based on the one or more input medical images using a machine learning-based network, the machine learning-based network generating a probability score associated with the medical imaging analysis task; Determine the uncertainty measure associated with the probability score by: applying a calibration function to the probability scores, the calibration function including fixed points defined according to a user-selected threshold at which the probability scores are completely uncertain; and calculating the entropy of the probability scores based on the results of the applied calibration function; as well as Make clinical decisions based on probability scores and uncertainty measures. 2 . The computer-implemented method of claim 1 , wherein the medical imaging analysis task comprises at least one of detection, classification, or segmentation of intracranial hemorrhage in a patient.
3. The computer-implemented method of claim 1 , wherein the machine-learning-based network generates parameters of a Beta distribution of scores, and determining an uncertainty measure associated with the probability scores further comprises: calculating the entropy of a beta distribution based on the parameters; and Combine the entropy of the beta distribution with the entropy of the probability scores.
4. The computer-implemented method of claim 1 , wherein making a clinical decision based on a probability score and an uncertainty measure comprises: Patients are stratified into one of a plurality of patient groups based on a probability score and an uncertainty measure.
5. The computer-implemented method of claim 4, wherein the medical imaging analysis task comprises detecting intracranial hemorrhage in a patient, and the plurality of patient groups comprises a high-confidence positive detection patient group, a high-confidence negative detection patient group, and a low-confidence patient group.
6. The computer-implemented method of claim 1 , wherein making a clinical decision based on a probability score and an uncertainty measure comprises: A decision is made whether to treat the patient based on the probability score and the uncertainty measure.
7. The computer-implemented method of claim 1 , wherein making a clinical decision based on a probability score and an uncertainty measure comprises: A determination is made whether to perform a clinical test on the patient based on the probability score and the uncertainty measure.
8. The computer-implemented method of claim 1 , wherein making a clinical decision based on a probability score and an uncertainty measure comprises: Prioritize radiologists' worklists based on probability scores and uncertainty measures.
9. A device comprising: means for receiving one or more input medical images of a patient; means for performing a medical imaging analysis task based on the one or more input medical images using a machine learning-based network, the machine learning-based network generating a probability score associated with the medical imaging analysis task; A component that determines the uncertainty measure associated with a probability score by: applying a calibration function to the probability scores, the calibration function including fixed points defined according to a user-selected threshold at which the probability scores are completely uncertain; and calculating the entropy of the probability scores based on the results of the applied calibration function; as well as Components for making clinical decisions based on probability scores and uncertainty measures.
10. The apparatus of claim 9, wherein the machine-learning-based network generates parameters of a Beta distribution of scores, and the means for determining an uncertainty measure associated with the probability score further comprises: means for calculating the entropy of a beta distribution based on the parameters; and Means for combining the entropy of a beta distribution with the entropy of a probability score.
11. The apparatus of claim 9, wherein the means for making a clinical decision based on a probability score and an uncertainty measure comprises: Component for stratifying a patient into one of a plurality of patient groups based on a probability score and an uncertainty measure. 12 . The apparatus according to claim 11 , wherein the medical imaging analysis task comprises detecting intracranial hemorrhage in a patient, and the plurality of patient groups comprises a high-confidence positive detection patient group, a high-confidence negative detection patient group, and a low-confidence patient group.
13. A non-transitory computer-readable medium storing computer program instructions that, when executed by a processor, cause the processor to perform operations comprising: receiving one or more input medical images of a patient; performing a medical imaging analysis task based on the one or more input medical images using a machine learning-based network, the machine learning-based network generating a probability score associated with the medical imaging analysis task; Determine the uncertainty measure associated with the probability score by: applying a calibration function to the probability scores, the calibration function including fixed points defined according to a user-selected threshold at which the probability scores are completely uncertain; and calculating the entropy of the probability scores based on the results of the applied calibration function; as well as Make clinical decisions based on probability scores and uncertainty measures. 14 . The non-transitory computer-readable medium of claim 13 , wherein the medical imaging analysis task comprises at least one of detection, classification, or segmentation of intracranial hemorrhage in a patient.
15. The non-transitory computer-readable medium of claim 13, wherein making a clinical decision based on a probability score and an uncertainty measure comprises: A decision is made whether to treat the patient based on the probability score and the uncertainty measure.
16. The non-transitory computer-readable medium of claim 13, wherein making a clinical decision based on a probability score and an uncertainty measure comprises: A determination is made whether to perform a clinical test on the patient based on the probability score and the uncertainty measure.
17. The non-transitory computer-readable medium of claim 13, wherein making a clinical decision based on a probability score and an uncertainty measure comprises: Prioritize radiologists' worklists based on probability scores and uncertainty measures.
Citation Information
Patent Citations
Medical image assessment with classification uncertainty
US20200320354A1
Machine learning from noisy labels for abnormality assessment in medical imaging
US20220028063A1
System and method for quantifying uncertainty in reasoning about 2d and 3D spatial features with a computer machine learning architecture
US20190122073A1