Automatic hemorrhage expansion detection from head ct images

By receiving and registering the patient's CT images, generating composite images, and using machine learning networks to assess bleeding expansion, the time-consuming and unreliable problems of existing technologies are solved, achieving efficient and accurate detection of bleeding expansion.

CN115205192BActive Publication Date: 2025-11-25SIEMENS HEALTHINEERS AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210299474.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-25
Filing Date
2022-03-25
Publication Date
2025-11-25
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

In the existing technology, the automatic detection of cerebral hemorrhage expansion relies on manual reading of CT images, which is time-consuming and limited by inter- and intra-rater variability, making timely diagnosis difficult. Furthermore, the automated system is unreliable in its assessment due to imaging artifacts and imperfect segmentation.

Method used

By receiving first and second input medical images from the patient, registration and segmentation are performed to generate a composite image, and features are extracted using a trained machine learning-based network to assess abnormal expansion, including hemorrhagic expansion.

Benefits of technology

It improves the accuracy of automated assessment of bleeding expansion, reduces data acquisition costs, shortens training time, and lowers development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205192B_ABST
    Figure CN115205192B_ABST
Patent Text Reader

Abstract

Systems and methods for assessing expansion of an abnormality are provided. A first input medical image of a patient and a second input medical image of the patient are received, the first input medical image depicting an abnormality at a first time, the second input medical image depicting the abnormality at a second time. The second input medical image is registered to the first input medical image. The abnormality is segmented from 1) the first input medical image to generate a first segmentation map, and from 2) the registered second input medical image to generate a second segmentation map. The first and second segmentation maps are combined to generate a combined map. Features are extracted from the first input medical image and the registered second input medical image based on the combined map. An expansion of the abnormality is assessed based on the extracted features using a trained machine learning-based network. A result of the assessment is output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates generally to hemorrhage expansion detection, and in particular to automatic hemorrhage expansion detection from head CT (computed tomography) images. BACKGROUND

[0002] Cerebral hemorrhage is typically caused by a blood vessel rupture within the brain resulting in localized bleeding in the surrounding tissue. Expansion of the hemorrhage, referred to as hemorrhage expansion, has been identified as an important biomarker indicating high risk of early neurological deterioration and poor long-term clinical outcome. Therefore, accurately detecting hemorrhage expansion within a patient is important to effectively stratify patients and tailor meticulous and timely patient care.

[0003] In current clinical practice, a pair of patient head CT (computed tomography) images acquired at different time points are qualitatively read by a radiologist to manually detect hemorrhage expansion. However, manual detection of hemorrhage expansion is time-consuming and limited by inter- and intra-rater variability due to the large amount of human interaction and judgment involved in reading the CT images, thus hindering timely diagnosis.

[0004] Recently, automated systems for localizing hemorrhage in a pair of baseline and follow-up head CT images of a patient have been proposed. In such automated systems, hemorrhage is segmented from the baseline and follow-up images, and the segmented hemorrhages are compared to assess hemorrhage expansion. However, assessing hemorrhage expansion by comparing the hemorrhage segmentations is unreliable due to imaging artifacts and imperfect segmentation. SUMMARY

[0005] According to one or more embodiments, a system and method for assessing expansion of an abnormality are provided. A first input medical image of a patient and a second input medical image of the patient are received, the first input medical image depicting an abnormality at a first time, and the second input medical image depicting the abnormality at a second time. The second input medical image is registered with the first input medical image. The abnormality is segmented from a) the first input medical image to generate a first segmentation map, and from b) the registered second input medical image to generate a second segmentation map. The first and second segmentation maps are combined to generate a combined map. Features are extracted from the first input medical image and the registered second input medical image based on the combined map. Expansion of the abnormality is assessed based on the extracted features using a trained machine learning-based network. A result of the assessment is output.

[0006] In one embodiment, the abnormality comprises a hemorrhage. The first and second input medical images can be CT (computed tomography) images of a patient's head.

[0007] In one embodiment, features are extracted from the first input medical image and the registered second input medical image based on the combined image by generating an input image based on the first input medical image, the registered second input medical image, and the combined image, extracting 2D in-plane features from slices of the generated input image, and extracting out-of-plane features from the extracted 2D in-plane features. The expansion of the abnormality can be assessed by determining an expansion score based on the extracted out-of-plane features. The expansion score can be compared to one or more thresholds. The input image can be generated by generating a 3-channel input image that includes the first input medical image, the registered second input medical image, and the combined image.

[0008] In one embodiment, the 2D in-plane features are extracted using a first trained machine learning-based feature extraction network, and the out-of-plane features are extracted using a second trained machine learning-based feature extraction network, and the trained machine learning-based network, the first trained machine learning-based feature extraction network, and the second trained machine learning-based feature extraction network are jointly trained.

[0009] In one embodiment, the first segmentation map and the second segmentation map are combined to generate the combined image by applying a voxel-wise OR operation on the first segmentation map and the second segmentation map.

[0010] These and other advantages of the application will become apparent to those of ordinary skill in the art upon reading the following detailed description and upon having reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 A method for assessing expansion of an abnormality according to one or more embodiments is shown;

[0012] Figure 2 A workflow for assessing expansion of a hemorrhage according to one or more embodiments is shown;

[0013] Figure 3 An assessment of expansion of a hemorrhage determined according to one or more embodiments is shown;

[0014] Figure 4 A table comparing a conventional segmentation-based detection system to a longitudinal detection network according to embodiments described herein is shown;

[0015] Figure 5 An exemplary artificial neural network that can be used to implement one or more embodiments is shown;

[0016] Figure 6A convolutional neural network that can be used to implement one or more embodiments is shown; and

[0017] Figure 7 A high-level block diagram of a computer that can be used to implement one or more embodiments is shown. Detailed Implementation

[0018] This invention generally relates to methods and systems for automated detection of hemorrhage expansion based on head CT (computed tomography) images. Embodiments of the invention are described herein to provide a visual understanding of the methods and systems. Digital images typically consist of digital representations of one or more objects (or shapes). The digital representations of objects are generally described herein in terms of identifying and manipulating them. Such manipulation is a virtual manipulation performed in the memory or other circuitry / hardware of a computer system. Therefore, it is to be understood that embodiments of the invention can be performed within a computer system using data stored within the computer system. Furthermore, references to image pixels herein can be equivalently used to refer to voxels of an image, and vice versa.

[0019] The embodiments described herein provide an automated assessment of the expansion of hemorrhage and other abnormalities. The expansion of hemorrhage is clinically termed hemorrhage expansion. The embodiments described herein apply a hemorrhage segmentation system to effectively distinguish pathological changes between a baseline input medical image and subsequent input medical images. The segmentation results are combined, and features are extracted from both the baseline and subsequent input medical images based on the combined segmentation results. A trained machine learning-based classifier network is applied to assess the expansion of hemorrhage based on the extracted features. Advantageously, the embodiments described herein provide an automated assessment of hemorrhage expansion with higher accuracy compared to conventional methods.

[0020] Figure 1 A method 100 for evaluating the expansion of an anomaly according to one or more embodiments is shown. Figure 2 A workflow 200 for assessing bleeding expansion is illustrated according to one or more embodiments. These will be described together. Figure 1 and Figure 2 The steps of method 100 can be performed by one or more suitable computing devices (such as, for example, Figure 7 The computer (702) is used to execute this.

[0021] exist Figure 1 At step 102, a first input medical image and a second input medical image of the patient are received. The first input medical image depicts the abnormality at a first time, and the second input medical image depicts the abnormality at a second time. In one embodiment, such as Figure 2As in the workflow 200 of FIG. 2, the abnormality is a hemorrhage. However, the abnormality can be any other abnormality of the patient, such as, for example, a lesion, a nodule, and other abnormalities in which tissue deformation and artifacts are involved. The first input medical image can be a baseline input medical image of the abnormality, and the second input medical image can be a subsequent input medical image of the abnormality. For example, as in the workflow 200 of FIG. 2, the first input medical image can be a baseline scan 202 of the patient’s head, and the second input medical image can be a subsequent scan 205 of the patient’s head. Figure 2 As shown in the workflow 200 of FIG. 2, the first input medical image can be a baseline scan 202 of the patient’s head, and the second input medical image can be a subsequent scan 205 of the patient’s head.

[0022] In one embodiment, the first input medical image and / or the second input medical image is a CT image. However, the first input medical image and / or the second input medical image can include any other suitable modality, such as, for example, MRI (magnetic resonance imaging), ultrasound, x-ray, or any other medical imaging modality or combination of medical imaging modalities. The first input medical image and / or the second input medical image can be a 2D (two-dimensional) image and / or a 3D (three-dimensional) volume, and can include a single input medical image or multiple input medical images. In one embodiment, the first input medical image and / or the second input medical image includes a 2.5D (2D plus time) image. The input medical images can be received directly from an image acquisition device (such as, for example, a CT scanner) at the time the input medical images are acquired, or can be received by loading previously acquired input medical images from a storage device or memory of the computer system, or receiving the input medical images from a remote computer system.

[0023] At step 104 of the workflow 200 of FIG. 2, the second input medical image is registered to the first input medical image. The registration spatially aligns the first input medical image and the second input medical image. In one example, at block 206, the baseline scan 202 and the subsequent scan 204 in the workflow 200 of FIG. 2 are spatially aligned to generate an aligned image 208 of the subsequent scan 204. The second input medical image can be registered to the first input medical image using any suitable method, such as, for example, known rigid registration or linear registration techniques. Figure 1 Figure 2 At step 106 of the workflow 200 of FIG. 2, the abnormality is segmented from a) the first input medical image to generate a first segmentation map, and from b) the registered second input medical image to generate a second segmentation map. In one example, at block 208, the baseline scan 202 and the subsequent scan 204 in the workflow 200 of FIG. 2 are segmented to generate a baseline segmentation map 210 and a subsequent segmentation map 212, respectively.

[0024] At step 108 of the workflow 200 of FIG. 2, the first segmentation map and the second segmentation map are compared to identify a change in the abnormality. In one example, at block 210, the baseline segmentation map 210 and the subsequent segmentation map 212 in the workflow 200 of FIG. 2 are compared to identify a change in the abnormality. Figure 1 Figure 2 ​​In workflow 200, bleed is segmented from baseline scan 202 to generate bleed map 212, and bleed is segmented from aligned image 208 of subsequent scan 204 to generate bleed map 216. Bleed maps 212 and 216 in workflow 200 can be binary segmentation maps, where, for example, a voxel (or pixel) intensity value of 1 indicates the presence of an anomaly at that voxel, and a voxel intensity value of 0 indicates the absence of an anomaly at that voxel.

[0025] In one embodiment, a trained machine learning-based segmentation network is used to perform the segmentation. The trained machine learning-based segmentation network can be implemented using U-NET, Dense U-Net, or any other suitable machine learning-based architecture. The trained machine learning-based segmentation network is trained using a benchmark ground truth annotation graph during a previous offline or training phase to segment anomalies from medical images. Once trained, the trained machine learning-based segmentation network is applied during an online or testing phase (e.g., in...). Figure 1 Step 106).

[0026] exist Figure 1 In step 108, the first segmentation image and the second segmentation image are combined to generate a combined image. For example, ... Figure 2 Bleeding maps 212 and 216 in workflow 200 are combined to generate attention map 218. In one embodiment, the first and second segmentation maps are combined by applying a voxel-wise (or pixel-wise) OR operation to them, such that a voxel value of 1 at a corresponding voxel in either the first or second segmentation map results in a voxel value of 1 at that voxel in the combined map, otherwise a voxel value of 0. Other methods for combining the first and second segmentation maps are also considered.

[0027] exist Figure 1 At step 110, features are extracted from the first input medical image and the registered second input medical image based on the composite map. Any suitable method can be used to extract features from the first input medical image and the registered second input medical image based on the composite map. The composite map identifies specific regions in which the anomaly is located in the first segmentation map or the second segmentation map, thereby enabling feature extraction from the first input medical image and the registered second input medical image while focusing on the specific regions identified by the composite map.

[0028] In one embodiment, features are extracted by first generating an input image based on a first input medical image, a registered second input medical image, and the combined image. The input image may be a 3-channel input image including the first input medical image, the registered second input medical image, and the combined image. For example, in... Figure 2In workflow 200, a 3-channel input volume is constructed at box 220.

[0029] Then, 2D in-plane features are extracted from the 3-channel input image. For example, in Figure 2 In workflow 200, 2D in-plane features are extracted from the 3-channel input volume by 2D in-plane feature extractor 222. The 2D in-plane features include latent features extracted from each 2D slice of the 3-channel input image. The 3-channel input image is used as an attention map to focus the extraction of 2D in-plane features on the regions identified in this combined map. Any suitable 2D machine learning-based segmentation network can be used to extract 2D in-plane features, such as, for example, a pre-trained 2D segmentation network, or a Res-Net32 / Res-Net50 network pre-trained using a public dataset (e.g., ImageNet).

[0030] Then, sequential out-of-plane features are extracted from the in-plane features in 2D. For example, in Figure 2 In workflow 200, out-of-plane features are extracted from 2D in-plane features by out-of-plane feature extractor 224. Out-of-plane features model the 3D context of the 3-channel input image. Out-of-plane features can be extracted using any suitable out-of-plane feature extractor trained to learn the relationship between 2D in-plane features and out-of-plane features. The out-of-plane feature extractor can be implemented using an RNN (Recurrent Neural Network) with LSTM (Long Short-Term Memory), BGRU (Bidirectional Gated Recurrent Unit), or any other suitable machine learning-based network.

[0031] exist Figure 1 At step 112, a trained machine learning-based network is used to evaluate the expansion of the anomaly based on the extracted features. The trained machine learning-based network can be any suitable trained machine learning-based classifier network. The trained machine learning-based classifier network receives the extracted out-of-plane features as input and generates an expansion score. For example, in... Figure 2 In workflow 200, the expansion of bleeding is evaluated by HE (bleed expansion) classifier 226 based on sequential out-of-plane features to determine an HE score 228. The trained machine learning-based classifier network first estimates a global latent feature vector from the extracted sequential out-of-plane features using max pooling, global average pooling, or any other suitable pooling method. Then, an expansion score is predicted based on the global latent feature vector through fully connected layers or fully convolutional blocks. The expansion score represents the probability of expansion of the anomaly between the first and second input medical images. The expansion score can be compared to one or more thresholds to provide a final result (e.g., expanded / no expansion, or expanded / no expansion / uncertain).

[0032] The trained machine learning-based classifier network is trained during a prior offline or training phase using annotated training image pairs. The training image pairs can be annotated as dilating, where the training image pairs depict, for example, at least a 33% increase in volume of the abnormality. Any other threshold increase in volume of the abnormality can be selected for annotating training images as depicting dilating. Once trained, the trained machine learning-based classifier network is applied during an online or testing phase (e.g., at step 112 of FIG. 1). Figure 1 of FIG. 1).

[0033] At step 114 of FIG. 1, the results of the assessment are output. For example, the results of the assessment can be output by displaying the assessment results on a display device of the computer system, storing the assessment results on a memory or storage device of the computer system, or by transmitting the assessment results to a remote computer system. Figure 1

[0034] Advantageously, the embodiments described herein model longitudinal image features in medical images acquired at different points in time, improving performance. Since the machine learning-based networks for extracting 2D in-plane features and sequential out-of-plane features can be trained with 2D image slices to model 3D longitudinal radiological features, fewer 3D training images are needed compared to conventional systems, reducing the cost of data acquisition. Moreover, the embodiments described herein can leverage existing pre-trained machine learning-based networks, resulting in faster training convergence and reduced overfitting, while reducing development costs.

[0035] In one embodiment, at least some of the machine learning-based networks used in the method 100 can be jointly trained. For example, a first trained machine learning-based feature extraction network can be used for 2D in-plane feature extraction (at step 110), a second trained machine learning-based feature extraction network can be used for sequential out-of-plane feature extraction (at step 110), and the first trained machine learning-based feature extraction network, the second trained machine learning-based feature extraction network, and the trained machine learning-based network (at step 112) can be jointly trained using an optimizer such as, for example, Adam.

[0036] Figure 3 An assessment result of dilatation of a hemorrhage determined in accordance with one or more embodiments is shown. A first input medical image 302 shows a baseline scan depicting a patient’s head with a hemorrhage at a first time, and a second input medical image 204 shows a follow-up scan depicting the patient’s head with the hemorrhage at a second time. As shown, the assessment result indicates that the hemorrhage has dilated between the first time and the second time. Figure 3 ​As shown in the middle, the first input medical image 302 is manually assessed to have a GT (ground truth) hemorrhage volume of 9.9 ml (milliliter), while the second input medical image 304 is manually assessed to have a GT hemorrhage volume of 17.3 ml. According to embodiments described herein, the first input medical image 302 and the second input medical image 304 are assessed to have a HE score of 0.9756.

[0037] Figure 4 A table 400 is shown that compares a conventional segmentation-based detection system and a longitudinal detection network according to embodiments described herein. The table 400 compares AUC (area under the curve), SEN (sensitivity), SPC (specificity), precision, recall, and F1 score.

[0038] Embodiments described herein are described in relation to claimed systems as well as claimed methods. Features, advantages, or alternative embodiments herein can be assigned to other claimed objects, and vice versa. In other words, features described or claimed in the context of a method can be used to improve claims for a system. In this case, the functional features of the method are embodied by providing the target units of the system.

[0039] Furthermore, certain embodiments described herein are described in relation to methods and systems for utilizing trained machine learning-based networks (or models), as well as methods and systems for training machine learning-based networks. Features, advantages, or alternative embodiments herein can be assigned to other claimed objects, and vice versa. In other words, features described or claimed in the context of methods and systems for utilizing trained machine learning-based networks can be used to improve claims for methods and systems for training machine learning-based networks, and vice versa.

[0040] In particular, trained machine learning-based networks applied in embodiments described herein can be adapted by methods and systems for training machine learning-based networks. Furthermore, input data of trained machine learning-based networks can comprise advantageous features and embodiments of training input data, and vice versa. Furthermore, output data of trained machine learning-based networks can comprise advantageous features and embodiments of output training data, and vice versa.

[0041] Generally, trained machine learning-based networks mimic cognitive functions of humans relating to other human minds. In particular, by training based on training data, trained machine learning-based networks are able to adapt to new situations and detect and infer patterns.

[0042] Generally, the parameters of a machine learning based network can be adapted with the help of training. In particular, supervised training, semi-supervised training, unsupervised training, reinforcement learning and / or active learning can be used. Furthermore, representation learning (alternative term: feature learning) can be used. In particular, the parameters of a trained machine learning based network can be adapted iteratively by several training steps.

[0043] In particular, the trained machine learning based network can comprise a neural network, a support vector machine, a decision tree and / or a Bayesian network, and / or the trained machine learning based network can be based on k-means clustering, Q-learning, a genetic algorithm and / or association rules. In particular, the neural network can be a deep neural network, a convolutional neural network or a convolutional deep neural network. Furthermore, the neural network can be an adversarial network, a deep adversarial network and / or a generative adversarial network.

[0044] Figure 5 An embodiment of an artificial neural network 500 is shown, in accordance with one or more embodiments. Alternative terms for "artificial neural network" are "neural network", "artificial neural net" or "neural net". The machine learning networks described herein, such as the method 100 and the workflow 200, and the machine learning based network used in the workflow 200, can be implemented using the artificial neural network 500. Figure 1 Figure 2

[0045] The artificial neural network 500 comprises nodes 502-522 and edges 532, 534,..., 536, wherein each edge 532, 534,..., 536 is a directed connection from a first node 502-522 to a second node 502-522. Generally, the first node 502-522 and the second node 502-522 are different nodes 502-522, it is also possible that the first node 502-522 and the second node 502-522 are the same. For example, in the embodiment shown in Fig. 5, the edge 532 is a directed connection from the node 502 to the node 506, and the edge 534 is a directed connection from the node 504 to the node 506. The edges 532, 534,..., 536 from a first node 502-522 to a second node 502-522 are also denoted as "incoming edges" for the second node 502-522, and as "outgoing edges" for the first node 502-522. Figure 5

[0046] In this embodiment, the nodes 502-522 of the artificial neural network 500 can be arranged in layers 524-530, wherein these layers can comprise an inherent order introduced by the edges 532, 534,..., 536 between the nodes 502-522. In particular, the edges 532, 534,..., 536 can only exist between adjacent layers of nodes. For example, in the embodiment shown in Fig. 5, the nodes 502-522 are arranged in three layers 524-530, wherein the edges 532, 534,..., 536 only exist between adjacent layers of nodes. In particular, the edges 532, 534,..., 536 do not exist between the nodes 502-522 of the same layer 524-530. Figure 5 ​​​In the embodiment shown in the middle, there is an input layer 524 comprising only nodes 502 and 504 without incoming edges, an output layer 530 comprising only nodes 522 without outgoing edges, and hidden layers 526, 528 between the input layer 524 and the output layer 530. In general, the number of hidden layers 526, 528 can be chosen arbitrarily. The number of nodes 502 and 504 within the input layer 524 is typically related to the number of input values of the neural network 500, and the number of nodes 522 within the output layer 530 is typically related to the number of output values of the neural network 500.

[0047] In particular, a (real) number can be assigned as a value to each node 502-522 of the neural network 500. Here, x (n) i denotes the value of the i-th node 502-522 of the n-th layer 524-530. The values of the nodes 502-522 of the input layer 524 are identical to the input values of the neural network 500, and the values of the nodes 522 of the output layer 530 are identical to the output values of the neural network 500. Furthermore, each edge 532, 534,..., 536 can comprise a weight which is a real number, in particular, the weight is a real number within the interval [-1, 1] or within the interval [0, 1]. Here, w (m,n) i,j denotes the weight of the edge between the i-th node 502-522 of the m-th layer 524-530 and the j-th node 502-522 of the n-th layer 524-530. Further, the abbreviation w (n) i,j is defined for the weight w (n,n+1) i,j .

[0048] In particular, in order to calculate the output values of the neural network 500, the input values are propagated through the neural network. In particular, the values of the nodes 502-522 of the (n+1)-th layer 524-530 can be calculated based on the values of the nodes 502-522 of the n-th layer 524-530 by the following equation:

[0049] .

[0050] In this context, the function f is a transfer function (another term is "activation function"). Known transfer functions are a step function, a sigmoid function (e.g. a logistic function, a generalized logistic function, a hyperbolic tangent function, an inverse tangent function, an error function, a smoothstep function), or a rectifier function. The transfer function is mainly used for normalization purposes.

[0051] In particular, these values are propagated through the neural network layer by layer, where the values of the input layer 524 are given by the input to the neural network 500, where the values of the first hidden layer 526 can be computed based on the values of the input layer 524 of the neural network, where the values of the second hidden layer 528 can be computed based on the values of the first hidden layer 526, and so on.

[0052] To set the values w (m,n) i,j of the edges, the neural network 500 has to be trained using training data. In particular, the training data comprises training input data and training output data (denoted as t i ). For a training step, the neural network 500 is applied to the training input data to generate computed output data. In particular, the training data and the computed output data comprise a certain number of values, which is equal to the number of nodes of the output layer.

[0053] In particular, the weights within the neural network 500 are adapted recursively using a comparison between the computed output data and the training data (backpropagation algorithm). In particular, the weights are changed according to the following formula:

[0054]

[0055] where γ is the learning rate, and if the (n+1)-th layer is not the output layer, δ (n+1) j is computed recursively as: (n) j

[0056]

[0057] and if the (n+1)-th layer is the output layer 530, it is computed as:

[0058]

[0059] where f' is the first derivative of the activation function, and y (n+1) j is the comparison training value of the j-th node of the output layer 530.

[0060] Figure 6 A convolutional neural network 600 according to one or more embodiments is shown. The convolutional neural network 600 can be used to implement the machine learning networks described herein, such as for example the machine learning based network used in the method 100 of Figure 1 and the workflow 200 of Figure 2 described herein.

[0061] In Figure 6 ​​In the embodiment shown in the middle, the convolutional neural network 600 comprises an input layer 602, a convolutional layer 604, a pooling layer 606, a fully connected layer 608, and an output layer 610. Alternatively, the convolutional neural network 600 can comprise several convolutional layers 604, several pooling layers 606, and several fully connected layers 608, as well as other types of layers. The order of these layers can be chosen arbitrarily, typically with the fully connected layer 608 being used as the last layer before the output layer 610.

[0062] In particular, within the convolutional neural network 600, the nodes 612-620 of one layer 602-610 can be considered to be arranged as a d-dimensional matrix or d-dimensional image. In particular, in the two-dimensional case, the value of a node 612-620 indexed by i and j in the n-th layer 602-610 can be denoted as x (n) [i, j] However, the arrangement of the nodes 612-620 of one layer 602-610 has no influence on the calculations performed within the convolutional neural network 600, as they are given by the structure of the edges and the weights only.

[0063] In particular, the convolutional layer 604 is characterized by the structure and the weights of the incoming edges forming a convolution operation based on a certain number of kernels. In particular, the structure and the weights of the incoming edges are chosen such that the value x (n) k of a node 614 of the convolutional layer 604 is calculated based on the values x (n-1) of the nodes 612 of the preceding layer 602 as a convolution where the convolution is defined in the two-dimensional case as:

[0064] .

[0065] Here, the k-th kernel K k is a d-dimensional matrix (in this embodiment a two-dimensional matrix) which is typically small compared to the number of nodes 612-618 (e.g. a 3x3 matrix or a 5x5 matrix). In particular, this means that the weights of the incoming edges are not independent, but are chosen such that they result in the convolution formula. In particular, for a kernel being a 3x3 matrix, there are only 9 independent weights (each entry in the kernel matrix corresponds to one independent weight), independent of the number of nodes 612-620 in the respective layer 602-610. In particular, for the convolutional layer 604, the number of nodes 614 in the convolutional layer is equal to the number of nodes 612 in the preceding layer 602 multiplied by the number of kernels.

[0066] If the nodes 612 of the preceding layer 602 are arranged as a d-dimensional matrix, using multiple kernels can be interpreted as adding an additional dimension (denoted as a "depth" dimension) so that the nodes 614 of the convolutional layer 614 are arranged as a (d+1)-dimensional matrix. If the nodes 612 of the preceding layer 602 are already arranged as a (d+1)-dimensional matrix including the depth dimension, using multiple kernels can be interpreted as a dilation along the depth dimension so that the nodes 614 of the convolutional layer 604 are also arranged as a (d+1)-dimensional matrix, where the size of the (d+1)-dimensional matrix with respect to the depth dimension is a kernel number times the size of the (d+1)-dimensional matrix with respect to the depth dimension in the preceding layer 602.

[0067] An advantage of using the convolutional layer 604 is that spatial local correlations of the input data can be exploited by enforcing a local connectivity pattern between the nodes of adjacent layers, in particular by connecting each node to only a small region of the nodes of the preceding layer.

[0068] In the embodiment shown in Figure 6 In the embodiment shown in

[0069] The pooling layer 606 is characterized by the structure and weights of the incoming edges and the activation function of its nodes 616 forming the pooling operation based on a non-linear pooling function f. For example, in the two-dimensional case, the value x (n) may be calculated as: (n-1)

[0070] .

[0071] In other words, by using the pooling layer 606, the number of nodes 614, 616 can be reduced by replacing a number di-d2of adjacent nodes 614 in the preceding layer 604 with a single node 616 in the pooling layer, which is calculated from the values of the number of adjacent nodes. In particular, the pooling function f can be a maximum function, an average, or an L2-norm. In particular, for the pooling layer 606, the weights of the incoming edges are fixed and not modified by training.

[0072] An advantage of using the pooling layer 606 is that the number of nodes 614, 616 and the number of parameters is reduced. This leads to a reduced amount of computation in the network and to a control of overfitting. ​

[0073] In the embodiment shown in Figure 6 In the embodiment shown in Fig. 6, the pooling layer 606 is max pooling, replacing four adjacent nodes with only one node, the value of which is the maximum of the values of the four adjacent nodes. Max pooling is applied to each d-dimensional matrix of the preceding layer; in this embodiment, max pooling is applied to each of two two-dimensional matrices, thereby reducing the number of nodes from 72 to 18.

[0074] The fully connected layer 608 is characterized by the fact that most, in particular all, edges between the nodes 616 of the preceding layer 606 and the nodes 618 of the fully connected layer 608 exist, and that the weight of each of these edges can be individually adjusted.

[0075] In this embodiment, the nodes 616 of the preceding layer 606 of the fully connected layer 608 are shown both as a two-dimensional matrix and additionally as uncorrelated nodes (they are indicated as a row of nodes, where the number of nodes has been reduced for better presentation). In this embodiment, the number of nodes 618 in the fully connected layer 608 is equal to the number of nodes 616 in the preceding layer 606. Alternatively, the number of nodes 616, 618 can be different.

[0076] Further, in this embodiment, the values of the nodes 620 of the output layer 610 are determined by applying a Softmax function to the values of the nodes 618 of the preceding layer 608. By applying the Softmax function, the sum of the values of all nodes 620 of the output layer 610 is 1, and all values of all nodes 620 of the output layer are real numbers between 0 and 1.

[0077] The convolutional neural network 600 can also comprise a ReLU (Rectified Linear Unit) layer or an activation layer with a non-linear transfer function. In particular, the number of nodes and the node structure contained in the ReLU layer are identical to the number of nodes and the node structure contained in the preceding layer. In particular, the value of each node in the ReLU layer is calculated by applying a rectification function to the value of the corresponding node of the preceding layer.

[0078] The inputs and outputs of different convolutional neural network blocks can be connected using summation (residual / dense neural networks), element-wise multiplication (attention), or other differentiable operators. Thus, the convolutional neural network structure can be nested rather than sequential if the entire pipeline is differentiable.

[0079] In particular, the convolutional neural network 600 can be trained based on a backpropagation algorithm. To prevent overfitting, regularized methods can be used, such as dropout of nodes 612-620, stochastic pooling, use of artificial data, weight decay based on LI or L2 norm, or maximum norm constraint. To train the same neural network, different loss functions can be combined to reflect joint training objectives. A subset of the neural network parameters can be excluded from optimization to preserve weights pre-trained on other datasets.

[0080] The systems, apparatuses, and methods described herein can be implemented using digital circuitry, or using one or more computers using well-known computer processors, memory units, storage devices, computer software, and other components. Generally, a computer includes one or more processors or processing units and a system memory, which can be composed of one or more of a number of memory devices, such as magnetic disks, internal hard disks, and removable disks. In addition, a computer can also include, in some embodiments, one or more storage devices that are external or removable to the computer, such as a magnetic hard disk, an optical disk, a flash memory, a universal serial bus (USB) drive, or a memory card. A computer can also include one or more input devices, such as a keyboard, a mouse, a trackball, a microphone, and one or more output devices, such as a display device, a printer, and speakers. A computer can also include one or more network interfaces, which can connect the computer to an outside network, such as the Internet, and can connect the computer to other computers via the Internet or other networks.

[0081] The systems, apparatuses, and methods described herein can be implemented using computers operating in a client-server relationship. Generally, in such a system, the client computer is located at a remote location from the server computer, and interacts via a network. The client-server relationship can be defined and controlled by computer programs running on the respective client and server computers.

[0082] The systems, apparatuses, and methods described herein can be implemented within a network-based cloud computing system. In such a network-based cloud computing system, a server or another processor connected to a network communicates with one or more client computers via the network. The client computers can communicate with the server via, for example, a web browser application residing and operating on the client computers. The client computers can store data on the server and access that data via the network. The client computers can transmit requests for data or requests for online services to the server via the network. The server can perform the requested services and provide the data to the client computer(s). The server can also transmit data adapted to cause the client computer to perform a specified function (e.g., perform a calculation, display specified data on a screen, etc.). For example, the server can transmit a request adapted to cause the client computer to perform one or more steps or functions of the methods and workflows described herein, including one or more steps or functions of the methods and workflows of FIGS. 1-2. Certain steps or functions of the methods and workflows described herein, including one or more steps or functions of the methods and workflows of FIGS. 1-2, can be performed by the server and transmitted to the client computer for execution. Figure 1 or functions of the methods and workflows described herein, including one or more steps or functions of the methods and workflows of FIGS. 1-2, can be performed by the server and transmitted to the client computer for execution. Figure 1One or more steps or functions of the methods and workflows described herein (including Figure 1 One or more steps of the methods and workflows described herein (including Figure 1 One or more steps of the methods and workflows described herein (including

[0083] The systems, apparatuses, and methods described herein can be implemented using a computer program product tangibly embodied in an information carrier, e.g., in a non-transitory machine- readable storage device; and the method and workflow steps described herein (including Figure 1 One or more steps or functions of the methods and workflows described herein (including A computer program, which is a set of instructions or statements developed in a programming language that can be executed directly or indirectly on a computer to perform a certain activity or bring about a certain result. The computer program can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0084] Figure 7 A high-level block diagram of an example computer 702 that can be used to implement the systems, apparatuses, and methods described herein is depicted in FIG. 7. The computer 702 includes a processor 704 that is operatively coupled to a data storage device 712 and a memory 710. The processor 704 controls the overall operation of the computer 702 by executing computer program instructions. The computer program instructions can be stored in the data storage device 712 or other computer readable medium, and loaded into the memory 710 when execution of the computer program instructions is desired. Thus, Figure 1 The methods and workflow steps or functions of the methods and workflows described herein (including Figure 1 The computer executable code of the methods and workflow steps or functions of the methods and workflows described herein (including Figure 1or the method and workflow steps or functions of 2. The computer 702 can also include one or more network interfaces 706 for communicating with other devices via a network. The computer 702 can also include one or more input / output devices 708 (e.g., a display, a keyboard, a mouse, a speaker, a button, etc.) that enable a user to interact with the computer 702.

[0085] The processor 704 can include both general and special purpose microprocessors, and can be the sole processor or one of multiple processors of the computer 702. For example, the processor 704 can include one or more central processing units (CPUs). The processor 704, the data storage device 712, and / or the memory 710 can include, be supplemented by, or be incorporated into, one or more application-specific integrated circuits (ASICs) and / or one or more field programmable gate arrays (FPGAs).

[0086] The data storage device 712 and the memory 710 each include a tangible, non-transitory computer readable storage medium. The data storage device 712 and the memory 710 can each include high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate synchronous dynamic random access memory (DDR RAM), or other random access solid state memory devices, and can include non-volatile memory, such as one or more magnetic disk storage devices (such as internal hard disks and removable disks), magneto-optical disk storage devices, optical disk storage devices, flash memory devices, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), digital versatile disc read-only memory (DVD-ROM) disks, or other non-volatile solid state storage devices.

[0087] The input / output devices 708 can include peripherals, such as a printer, scanner, display screen, etc. For example, the input / output devices 708 can include a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, for displaying information to the user and a keyboard, including alphanumeric and other keys, for communicating information to the computer 702. The input / output devices 708 can also include a pointing device, such as a mouse or trackball, for communicating direction information and command selections to the computer 702.

[0088] The image acquisition device 714 can be connected to the computer 702 for inputting image data (e.g., medical images) to the computer 702. It is possible to implement the image acquisition device 714 and the computer 702 as one device. It is also possible for the image acquisition device 714 and the computer 702 to communicate wirelessly over a network. In possible embodiments, the computer 702 can be remotely located with respect to the image acquisition device 714.

[0089] Any or all of the systems and apparatus discussed herein can be implemented using one or more computers, such as computer 702.

[0090] Those skilled in the art will realize that the implementation of the actual computer or computer system can have other structures and can contain other components than those described in the figure, Figure 7 is a high-level representation of some of the components in such a computer.

[0091] The foregoing detailed description has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the invention be limited not with this detailed description, but rather determined from the claims as interpreted in accordance with the full breadth permitted by the patent laws. It is to be understood that the embodiments shown and described herein are merely illustrative of the principles of this invention and that various modifications can be implemented by those skilled in the art without departing from the scope and spirit of the invention. Various other combinations of features can be implemented by those skilled in the art without departing from the scope and spirit of the invention.

Claims

1. A computer-implemented method comprising: receiving 1) a first input medical image of a patient and 2) a second input medical image of the patient, the first input medical image depicting an abnormality at a first time, the second input medical image depicting the abnormality at a second time; registering the second input medical image with the first input medical image; segmenting the abnormality from a) the first input medical image to generate a first segmentation map and from b) the registered second input medical image to generate a second segmentation map; combining the first and second segmentation maps to generate a combined map; extracting features from the first input medical image and the registered second input medical image based on the combined map, wherein the combined map identifies a particular region in which the abnormality is either located in the first segmentation map or in the second segmentation map, thereby enabling the features to be extracted from the first input medical image and the registered second input medical image with a focus on the particular region identified by the combined map; evaluating an expansion of the abnormality based on the extracted features using a trained machine learning-based network; and outputting a result of the evaluation.

2. The computer-implemented method of claim 1, wherein the abnormality comprises a hemorrhage.

3. The computer-implemented method of claim 1, wherein extracting features from the first input medical image and the registered second input medical image based on the combined map comprises: generating an input image based on the first input medical image, the registered second input medical image, and the combined map; extracting 2D in-plane features from slices of the generated input image; and extracting out-of-plane features from the extracted 2D in-plane features.

4. The computer-implemented method of claim 3, wherein evaluating an expansion of the abnormality based on the extracted features using a trained machine learning-based network comprises: determining an expansion score based on the extracted out-of-plane features.

5. The computer-implemented method of claim 4, wherein evaluating an expansion of the abnormality based on the extracted features using a trained machine learning-based network further comprises: comparing the expansion score to one or more thresholds.

6. The computer-implemented method of claim 3, wherein generating an input image based on the first input medical image, the registered second input medical image, and the combined map comprises: generating a 3-channel input image comprising the first input medical image, the registered second input medical image, and the combined map.

7. The computer-implemented method of claim 3, wherein extracting the 2D in-plane features is performed using a first trained machine learning-based feature extraction network and extracting the out-of-plane features is performed using a second trained machine learning-based feature extraction network, and the trained machine learning-based network, the first trained machine learning-based feature extraction network, and the second trained machine learning-based feature extraction network are jointly trained. ​ 8. The computer-implemented method of claim 1, wherein combining the first segmentation map and the second segmentation map to generate a combined map comprises: applying a voxel-wise OR operation on the first segmentation map and the second segmentation map to generate the combined map.

9. The computer-implemented method of claim 1, wherein the first input medical image and the second input medical image are computed tomography images of a patient’s head.

10. An apparatus comprising: means for receiving 1) a first input medical image of a patient and 2) a second input medical image of the patient, the first input medical image depicting an abnormality at a first time, the second input medical image depicting the abnormality at a second time; means for registering the second input medical image with the first input medical image; means for segmenting the abnormality from a) the first input medical image to generate a first segmentation map and from b) the registered second input medical image to generate a second segmentation map; means for combining the first segmentation map and the second segmentation map to generate a combined map; means for extracting features from the first input medical image and the registered second input medical image based on the combined map, wherein the combined map identifies a particular region in which the abnormality is either located in the first segmentation map or in the second segmentation map, thereby enabling the extraction of features from the first input medical image and the registered second input medical image with a focus on the particular region identified by the combined map; means for assessing an expansion of the abnormality based on the extracted features using a trained machine learning-based network; and means for outputting a result of the assessment.

11. The apparatus of claim 10, wherein the abnormality comprises a hemorrhage.

12. The apparatus of claim 10, wherein the means for extracting features from the first input medical image and the registered second input medical image based on the combined map comprises: means for generating an input image based on the first input medical image, the registered second input medical image, and the combined map; means for extracting 2D in-plane features from slices of the generated input image; and means for extracting out-of-plane features from the extracted 2D in-plane features.

13. The apparatus of claim 12, wherein the means for assessing an expansion of the abnormality based on the extracted features using a trained machine learning-based network comprises: means for determining an expansion score based on the extracted out-of-plane features.

14. The apparatus of claim 13, wherein the means for assessing an expansion of the abnormality based on the extracted features using a trained machine learning-based network further comprises: means for comparing the expansion score to one or more thresholds.

15. A non-transitory computer-readable medium storing computer program instructions that, when executed by a processor, cause the processor to perform operations comprising: ​ ​ receiving 1) a first input medical image of a patient and 2) a second input medical image of the patient, the first input medical image depicting an abnormality at a first time, the second input medical image depicting the abnormality at a second time; registering the second input medical image with the first input medical image; segmenting the abnormality from a) the first input medical image to generate a first segmentation map and from b) the registered second input medical image to generate a second segmentation map; combining the first and second segmentation maps to generate a combined map; extracting features from the first input medical image and the registered second input medical image based on the combined map, wherein the combined map identifies a particular region in which the abnormality is either in the first segmentation map or in the second segmentation map, thereby enabling the extraction of features from the first input medical image and the registered second input medical image with a focus on the particular region identified by the combined map; evaluating an expansion of the abnormality using a trained machine learning based network based on the extracted features; and outputting a result of the evaluation.

16. The non-transitory computer readable medium of claim 15, wherein extracting features from the first input medical image and the registered second input medical image based on the combined map comprises: generating an input image based on the first input medical image, the registered second input medical image, and the combined map; extracting 2D in-plane features from slices of the generated input image; and extracting out-of-plane features from the extracted 2D in-plane features.

17. The non-transitory computer readable medium of claim 16, wherein generating an input image based on the first input medical image, the registered second input medical image, and the combined map comprises: generating a 3-channel input image comprising the first input medical image, the registered second input medical image, and the combined map.

18. The non-transitory computer readable medium of claim 17, wherein extracting the 2D in-plane features is performed using a first trained machine learning based feature extraction network and extracting the out-of-plane features is performed using a second trained machine learning based feature extraction network, and the trained machine learning based network, the first trained machine learning based feature extraction network, and the second trained machine learning based feature extraction network are jointly trained.

19. The non-transitory computer readable medium of claim 15, wherein combining the first and second segmentation maps to generate a combined map comprises: applying a voxel-wise OR operation on the first and second segmentation maps to generate the combined map.

20. The non-transitory computer readable medium of claim 15, wherein the first input medical image and the second input medical image are computed tomography images of a patient’s head. ​

Citation Information

Patent Citations

  • Image processing method, model training method and device and storage medium

    CN109978037A

  • System and method for detecting abnormal tissue using vascular features

    CN111801704A