Systems and methods for inferring thickness of a class of objects of interest in two-dimensional medical images using deep neural networks

By extracting features from 2D medical images using deep neural networks and generating segmentation and thickness masks, the usability of 3D imaging systems in resource-limited areas is addressed. This enables accurate assessment of the volume of objects of interest and enhances the diagnostic and evaluation capabilities of 2D imaging systems.

CN116964629BActive Publication Date: 2026-04-10GE PRECISION HEALTHCARE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GE PRECISION HEALTHCARE LLC
Filing Date
2022-01-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In the present technology, the availability of three-dimensional imaging systems such as CT systems is limited in rural areas or developing countries, and they require highly trained technicians, which makes it difficult to assess the volume of object of interest classes, especially during pandemics when the demand from a large number of patients increases, and 2D imaging systems cannot effectively assess the volume of object of interest classes.

Method used

By using deep neural networks, especially convolutional neural networks (CNNs), features are extracted from 2D medical images to generate segmentation and thickness masks for object classes of interest, and then their volumes are calculated. This process includes feature extraction, segmentation mask and thickness mask mapping, and spatial regularization constraints are used to reduce noise and roughness.

Benefits of technology

It enables accurate assessment of the volume and thickness of object classes of interest in 2D medical images, expands the diagnostic and assessment capabilities of 2D imaging systems, reduces reliance on 3D imaging systems, and improves diagnostic capabilities in resource-limited areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116964629B_ABST
    Figure CN116964629B_ABST
Patent Text Reader

Abstract

Methods and systems for inferring thickness and volume of one or more object classes of interest in two-dimensional (2D) medical images using deep neural networks are provided. In an exemplary embodiment, thickness of an object class of interest can be inferred by obtaining a 2D medical image; extracting features from the 2D medical image; mapping the features to a segmentation mask of the object class of interest using a first convolutional neural network (CNN); mapping the features to a thickness mask of the object class of interest using a second CNN, wherein the thickness mask indicates a thickness of the object class of interest at each pixel of a plurality of pixels of the 2D medical image; and determining a volume of the object class of interest based on the thickness mask and the segmentation mask.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the subject matter disclosed herein relate to medical imaging, including x-ray and computed tomography (CT) imaging. In particular, the present disclosure provides systems and methods for inferring a three-dimensional (3D) depth or thickness of one or more materials of interest in a two-dimensional (2D) image. BACKGROUND

[0002] Determining the volume of a class of objects of interest, such as a tissue type, an organ, or a disease-affected region (e.g., the volume of inflamed tissue, necrotic tissue, a tumor, etc.) can be useful in diagnosing or assessing the condition of a patient. As an example, diagnosing a disease affecting lung tissue, such as pneumonia or severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), can be based on the volume of inflamed lung tissue compared to non-inflamed lung tissue. However, determining the volume of a class of objects of interest typically relies on a three-dimensional (3D) imaging system, such as a computed tomography (CT) system, a magnetic resonance imaging (MRI) system, or a positron emission tomography (PET) system. In comparison to two-dimensional (2D) imaging modalities, conventional 3D imaging systems are expensive, large / fixed, and typically require more highly trained technicians. As a result, 3D imaging systems can be less readily available than 2D imaging systems, e.g., 3D imaging systems can have limited availability in rural areas or developing countries, and even in large hospitals, there can be more 2D imaging systems available than 3D imaging systems. This reduced availability of 3D imaging systems can be exacerbated in situations where a large number of patients can benefit from volume assessment of a class of objects of interest, such as during a pandemic where a large number of patients can seek diagnosis / assessment via volume analysis of a class of objects of interest. Thus, it is generally desirable to explore systems and methods for determining the volume of a class of objects of interest from a 2D medical image. SUMMARY

[0003] The inventors herein have developed systems and methods that can enable determination of volume information for at least a first object class of interest from a 2D medical image, thereby extending the functionality of 2D imaging modalities for diagnosis and assessment of medical conditions. In one embodiment, the present disclosure provides a method comprising: obtaining a 2D medical image; extracting features from the 2D medical image; mapping the features to a segmentation mask for an object class of interest using a first convolutional neural network (CNN); mapping the features to a thickness mask for the object class of interest using a second CNN, wherein the thickness mask indicates a thickness of the object class of interest at each pixel of a plurality of pixels of the 2D medical image; and determining a volume of the object class of interest based on the thickness mask and the segmentation mask. In this way, deep / thickness information can be extracted from a 2D medical image by utilizing a CNN to produce a thickness mask for an object class of interest. The volume of the object class of interest can then be determined using the thickness mask, which can be beneficial for patient assessment / diagnosis.

[0004] The above advantages of the present specification, and other advantages and features, will be apparent from the following detailed description, when read in conjunction with the appended drawings. It is to be understood that the above description is a summary of aspects of the disclosure and not intended to be used to limit the claimed subject matter, which is defined by the claims below. The summary is not intended to be used to determine the critical or essential features of the claimed subject matter used in order to determine the scope of claims, which is limited solely by the claims below. Furthermore, the claimed subject matter is not limited to implementations that solve any or all of the disadvantages mentioned above or any of the problems in the background art described herein. BRIEF DESCRIPTION OF DRAWINGS

[0005] Various aspects of the disclosure can be better understood with reference to the following detailed description when considered in connection with the accompanying drawings, in which:

[0006] Figure 1 is a block diagram of a system for determining a volume of an object class of interest from a 2D medical image according to an example embodiment;

[0007] Figure 2 is a block diagram of an example embodiment of a medical imaging system;

[0008] Figure 3 is a flowchart illustrating a method for determining a thickness mask for at least a first object class of interest according to an example embodiment;

[0009] Figure 4 is a flowchart illustrating a method for generating training data that can be used to train a deep neural network to map a 2D medical image to a thickness mask for one or more object classes of interest according to an example embodiment;

[0010] Figure 5is a flowchart illustrating a method for determining projection parameters for projecting a 3D image onto a 2D plane to produce a synthetic 2D image based on separately acquired and corresponding 2D medical images, according to an example embodiment;

[0011] Figure 6A a process of projecting a 3D image onto a 2D plane to produce a synthetic 2D image is shown, according to an example embodiment;

[0012] Figure 6B an example synthetic 2D image generated by the process shown; Figure 6A

[0013] Figure 7 is a flowchart illustrating a method for training a deep neural network to map a 2D medical image to a thickness mask of one or more object classes of interest, according to an example embodiment;

[0014] Figure 8A an example embodiment of a thickness heat map that can be generated from a thickness mask of an object class of interest is shown;

[0015] Figure 8B an example embodiment of a pseudo-3D image that can be generated from a thickness mask of an object class of interest is shown;

[0016] Figure 9 a pathology prediction overlaid on a 2D medical image is shown, according to an example embodiment, where the pathology prediction can be based on an inferred volume of a disease region;

[0017] Figure 10 an example embodiment of a spatial regularization constraint that can be imposed on a deep neural network trained to infer a thickness of an object class of interest based on a 2D medical image is shown; and

[0018] Figure 11 generation of a depth information encoding vector is shown, according to an example embodiment.

[0019] The accompanying drawings illustrate the described systems and methods for using a deep neural network to infer a thickness of an object class of interest from a 2D medical image. The drawings, together with the description, illustrate and explain the structures, methods, and principles described herein. In the drawings, the size of components can be exaggerated or otherwise modified to illustrate the various embodiments. Well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the described components, systems, and methods. DETAILED DESCRIPTION

[0020] ​The following description relates to using a deep neural network to infer the depth of one or more object classes of interest in a 2D medical image. This disclosure includes aspects relating to generating training data for a deep neural network, training said deep neural network, and implementing the deep neural network to infer a depth mask for one or more object classes of interest. This disclosure also includes systems and methods for determining the volume and / or pathological prediction of one or more object classes of interest based on the inferred thickness mask.

[0021] In one implementation, the thickness prediction system such as Figure 1 The thickness prediction system 100 shown can use one or more trained convolutional neural networks (CNNs) to determine one or more of the thickness mask and volume prediction for one or more object classes of interest in a 2D medical image. The thickness prediction system 100 can be generated by an imaging system (such as...) Figure 2 The imaging system 200 shown is implemented to process the acquired 2D images. Figure 3 The illustrated method 300 provides an exemplary embodiment of a method by which a 2D image acquired by a 2D imaging system (such as imaging system 200) can be mapped to a thickness mask and optionally used to determine pathological predictions and / or volumes of an object class of interest. Training data including the 2D image and the corresponding ground-based thickness mask can be used to train a deep neural network to infer the thickness of the object class of interest (such as a first CNN 106 of thickness prediction system 100). The training data can be generated according to exemplary method 400, such as... Figure 4 As shown. In some implementations, the generation of training data may include projecting annotated three-dimensional (3D) images onto a 2D plane to produce synthetic 2D images. Figure 5 The method 500 shown illustrates an exemplary method for determining projection parameters, thereby projecting a 3D image onto a 2D plane to produce a synthetic 2D image that matches a previously acquired 2D medical image. Adjusting the projection parameters until the difference between the 2D synthetic image and the 2D medical image is below a threshold allows the same projection parameters to be applied to 3D annotations of the 3D image to produce a ground-based thickness mask corresponding to the 2D medical image. Figure 6A This demonstrates the process by which 3D images can be projected onto a 2D plane, while Figure 6B It shows that it can be based on Figure 6A The projection process shown generates a ground-based thickness mask (and a synthesized 2D image). This can be viewed... Figure 7 In the method 700 shown, a deep neural network is trained using training data pairs generated by method 400 to infer a thickness mask from an input 2D medical image.

[0022] Thickness masks can be used to generate a visual display of the thickness of an object class of interest, such as...Figure 8A the thickness heat map 802 or Figure 8B the pseudo 3D image 804. Again, the thickness mask can also be used to determine a pathology prediction, where in Figure 9 An example pathology prediction 902 is shown in FIG. 8B. Further, the inventors herein determined that applying a spatial regularization constraint to the filters of the CNN can advantageously reduce noise and roughness in the generated thickness mask. In Figure 10 A graphical illustration of an example regularization method is shown in FIG. 8C. Figure 11 An example process by which a depth information encoding vector can be generated from a 3D image (or cross-sectional 2D image) to generate a ground truth thickness mask is shown in FIG. 9, where each point of the ground truth thickness mask includes a depth information encoding vector indicating a depth-wise density of an object class of interest or a depth-wise object class label of a plurality of object classes of interest. In addition to thickness, a deep neural network trained using the depth information encoding vector can also infer a depth-dependent density and / or a depth-dependent location of one or more object classes of interest.

[0023] As used herein, the term object class of interest can refer to one or more of a biological tissue, an organ, a disease-affected region, a surgically implanted device, a tumor, a cavity or space within a bioimaging subject, a plaque or fat accumulation, and a biological fluid. As used herein, the terms thickness or depth, which can be used interchangeably, refer to the extent of an object class of interest in a direction perpendicular to the surface or plane of a 2D medical image when used to describe the object class of interest in a 2D medical image. As an example, if the width of a 2D medical image extends parallel to an x-axis, and the height of the 2D medical image extends parallel to a y-axis, then the thickness or depth of an object imaged by the 2D medical image can be considered to extend parallel to a z-axis that extends into (and out of) the plane of the 2D medical image, where the x-axis, y-axis, and z-axis are each perpendicular to one another. When applied to a description of an object class of interest captured by a 2D image, the term “area” refers herein to the area of the plane of the 2D medical image occupied by the object class of interest. In other words, in a 2D medical image comprising a plurality of pixels, the area of an object class of interest can refer to the plurality of pixels in the plurality of pixels that depict the object class of interest. The area of an imaged object class of interest can be converted into physical units, such as cm 2 Similarly, the term volume, when used herein to describe an object class of interest, refers to the three-dimensional (3D) volume occupied by the object class of interest in 3D space. In one example, the volume of an object class of interest captured by a 3D image can be proportional to the number of voxels of the 3D image occupied by the object class of interest. The volume of an object class of interest can be converted into physical units, such as cm 3The volume of the object of interest captured in the 3D image can be determined by multiplying the number of voxels occupied by the object class of interest by a conversion factor. Alternatively, the physical unit of the volume of the object class of interest captured in the 3D image can be determined by multiplying the fraction of the total 3D image occupied by the object class of interest by the total physical volume of the 3D image.

[0024] Go to Figure 1 An exemplary embodiment of a thickness prediction system 100 is illustrated. The thickness prediction system 100 is configured to determine the thickness mask and volume of an object class of interest in a 2D medical image, and optionally to determine pathological predictions of the 2D medical image. The thickness prediction system 100 may be provided by an image processing system such as... Figure 2 The image processing device 202 of the imaging system 200 shown is implemented. The thickness prediction system 100 includes a first feature extractor 104 configured to extract features from an input 2D medical image 102. A first CNN 106 is configured to receive the features extracted from the 2D medical image 102 and segment one or more object classes of interest to generate a segmentation mask 110. Similarly, a second CNN 108 is configured to map the features extracted from the 2D medical image 102 to a thickness mask 112, thereby indicating the thickness of at least a first object class of interest at each point / pixel in the 2D medical image 102. The segmentation mask 110 and the thickness mask 112 can be used to generate a segmented thickness mask 114 from which a volume 116 for at least the first object class of interest can be determined. Optionally, in addition to the segmentation mask 110 and the thickness mask 112, a classifier 130 may determine a pathology prediction 132 based on the features extracted by the feature extractor 104. Furthermore, an optional second thickness mask 174 may be generated by a second feature extractor 170 and a third CNN 172, wherein, in addition to the features extracted by the first feature extractor 104, the second thickness mask 174 may also be fed as input to the second CNN 108.

[0025] The thickness prediction system 100 can receive 2D medical images 102 from one or more external devices, such as image libraries or imaging devices. The 2D medical images 102 can include 2D medical images acquired through virtually any 2D imaging modality known in the field of medical imaging, including but not limited to X-ray imaging, MRI, CT imaging, PET imaging, ultrasound imaging, optical imaging, etc. In some embodiments, the 2D medical image 102 includes a matrix of intensity values ​​in one or more color channels, wherein each intensity value in each of the one or more color channels uniquely corresponds to the intensity value of the associated pixel. The 2D medical image 102 can include an image of an anatomical region of the object being imaged. Figure 1 In the example shown, 2D medical image 102 is a chest X-ray of the patient.

[0026] 2D medical image 102 is fed to a first feature extractor 104. The first feature extractor 104 is configured to extract features from the 2D medical image 102 to produce a feature map. The features can include pixel intensity values, patterns of pixel intensity values, or patterns of previously extracted features. In some embodiments, the feature map indicates, for each of a plurality of sub-regions of the 2D medical image 102, a degree of matching between the sub-region and a filter, wherein the relative positions of each of the plurality of sub-regions of the 2D medical image 102 are maintained in the relative positions of the features in the feature map. In some embodiments, the first feature extractor 104 is a deep neural network configured to map a matrix of pixel intensity values of the 2D medical image 102 to the feature map using one or more convolutional layers, fully connected layers, activation functions, regularization layers, and stochastic dropout layers. In some embodiments, the first feature extractor 104 can not include learnable parameters, but can include an expert system configured to extract one or more predetermined features from the 2D medical image 102 based on hard-coded domain knowledge.

[0027] The features of the 2D medical image 102 extracted by the first feature extractor 104 are fed to a first CNN 106. The first CNN 106 includes one or more convolutional layers, where each convolutional layer of the one or more convolutional layers includes one or more filters comprising a plurality of learnable weights having a predetermined receptive field and stride. The first CNN 106 can receive the features extracted by the first feature extractor 104 as a feature map, where the spatial relationships of each of the extracted features are preserved within the feature map and encoded in the relative positions of each feature within the feature map. The first CNN 106 is configured to map the features of the 2D medical image 102 to a segmentation mask of at least a first object class of interest. In one embodiment, the segmentation mask 110 includes a plurality of values or a matrix of values corresponding to a plurality of pixel intensity values of the 2D medical image 102, where each value of the segmentation mask 110 indicates a classification of a corresponding pixel of the 2D medical image 102. In some embodiments, the segmentation mask 110 can be a binary segmentation mask including a matrix of 1s and 0s, where a 1 indicates that a pixel belongs to the object class of interest and a 0 indicates that a pixel does not belong to the object class of interest. In some embodiments, the segmentation mask 110 can include a multi-class segmentation mask including a matrix of N different integers (e.g., 0, 1,... N), where each different integer uniquely corresponds to an object class, thus enabling the multi-class segmentation mask to encode location and area information for a plurality of object classes of interest. The values of the segmentation mask 110 spatially correspond to the pixels / intensity values of the 2D medical image such that if the segmentation mask 110 is overlaid onto the 2D medical image 102, each value of the segmentation mask will align with (i.e., be overlaid on) a corresponding pixel of the 2D medical image 102, the object classification of each pixel of the 2D medical image will be indicated by the corresponding value of the segmentation mask.

[0028] The features of the 2D medical image 102 extracted by the first feature extractor 104 are also fed to the second CNN 108. The second CNN 108 comprises one or more convolutional layers, where each convolutional layer of the one or more convolutional layers comprises one or more filters comprising a plurality of learnable weights having a predetermined receptive field size and stride. The second CNN 108 can receive the features extracted by the first feature extractor 104 as a feature map, where the spatial relationship of each of the extracted features is preserved within the feature map and encoded in the relative position of each feature within the feature map. The second CNN 108 is configured to map the features of the 2D medical image 102 to a thickness mask of at least a first object class of interest. The thickness mask 112 output by the second CNN 108 can comprise a matrix of thickness values of at least the first object class of interest, where each value of the matrix of thickness values indicates a thickness of at least the first object class of interest at a corresponding pixel / position of the 2D medical image 102.

[0029] The segmentation mask 110 can be applied to the thickness mask 112 to suppress the thickness values of regions corresponding to non-object classes of interest, thereby resulting in a segmented thickness mask 114. In some embodiments, the segmented thickness mask 114 can be generated by performing a pairwise multiplication between the plurality of values of the segmentation mask 110 and the plurality of thickness values of the thickness mask 112, where each value of the segmented thickness mask 114 is a product of the multiplication between a spatially corresponding value of the segmentation mask 110 and a spatially corresponding thickness value of the thickness mask 112. By applying the segmentation mask 110 to the thickness mask 112 to suppress the thickness values that are not classified as belonging to the object class of interest, a more accurate volume can be determined for the object class of interest.

[0030] Since both the first CNN 106 and the second CNN 108 receive the features extracted by the feature extractor 104, the first CNN 106 can be referred to as a first CNN branch and the second CNN 108 can be referred to as a second CNN branch, where the feature extractor 104, the first CNN 106, and the second CNN 108 can constitute a single deep neural network. In some embodiments, each of the feature extractor 104, the first CNN 106, and the second CNN 108 can be trained during a single training process, where a first loss can be determined based on the output of the first CNN 106 and a second loss can be determined based on the output of the second CNN 108, where the first loss is used to update the parameters of the first CNN 106 and the second loss is used to update the parameters of the second CNN 108, and both the first loss and the second loss are used to update the parameters of the feature extractor 104. In some embodiments, the feature extractor 104, the first CNN 106, and the second CNN 108 can each be trained separately.

[0031] The thickness prediction system 100 can determine a volume of at least a first object class of interest based on the segmented thickness mask 114. In some embodiments, each thickness value of the segmented thickness mask can be summed to produce a thickness sum, and then the thickness sum can be multiplied by a conversion factor to produce a volume of the object class of interest. In some embodiments, the thickness values of the segmented thickness mask 114 can be considered points in a 3D space, where the z-coordinate of a point in the 3D space is given by the thickness value, and the x and y coordinates of each point in the 3D space correspond to the row and column, respectively, of the corresponding pixel in the 2D medical image. The volume of the object class of interest can then be obtained as an integral or an approximation of an integral of a 3D surface formed by the plurality of points in the 3D space.

[0032] The thickness prediction system 100 can optionally include a second feature extractor 170 and a third CNN 172 configured to determine a second thickness mask 174 from the 2D medical image 102. When present, the second thickness mask 174 can be concatenated, merged, or otherwise combined with the feature maps produced by the first feature extractor 104, and the combined feature maps and second thickness mask 174 can be fed to the second CNN and mapped to a first thickness mask of a first object class of interest. In some embodiments, the second thickness mask indicates thicknesses of a second object class of interest that is different from the first object class of interest indicated by the first thickness mask 112. In some embodiments, the first object class of interest is a disease-affected region, and the second object class of interest is thickness of a non-disease-affected region. In some embodiments, the first object class of interest is a first tissue type, and the second object class of interest is a second tissue type. In some embodiments, the second object class of interest is a total object depth (e.g., a total depth of object tissue at each pixel of the 2D medical image 102). By first determining the second thickness mask 174, and using this thickness mask as input into the second CNN 108 to determine the first thickness mask, the inventors herein have found that accuracy of the thickness values of the first thickness mask can be increased.

[0033] The second feature extractor 170 can receive the 2D medical image 102 as input and extract features from the 2D medical image 102 to a feature map. In some embodiments, the second feature extractor 170 is a deep neural network configured to map a matrix of pixel intensity values of the 2D image 102 to a feature map using one or more convolutional layers, fully connected layers, activation functions, regularization layers, and stochastic dropout layers. In some embodiments, the second feature extractor 170 can not include learnable parameters, but can include an expert system configured to extract one or more predetermined features from the 2D medical image 102 based on hard-coded domain knowledge.

[0034] The features of the 2D medical image 102 extracted by the second feature extractor 170 are fed into a third CNN 172. The third CNN 172 comprises one or more convolutional layers, where each convolutional layer of the one or more convolutional layers comprises one or more filters comprising a plurality of learnable weights with a predetermined receptive field size and stride. The third CNN 172 can receive the features extracted by the second feature extractor 170 as a feature map, where the spatial relationship of each of the extracted features is preserved within the feature map and encoded in the relative position of each feature within the feature map. The third CNN 172 is configured to map the features of the 2D medical image 102 to a second thickness mask 174 of at least a second object class of interest. The second thickness mask 174 output by the third CNN 172 can comprise a matrix of thickness values of at least the second object class of interest, where each value of the matrix of thickness values indicates a thickness of at least the second object class of interest at a corresponding pixel / position of the 2D medical image 102. As described above, the second thickness mask 174 is then fed as input into the second CNN 108 along with the features extracted by the first feature extractor 104.

[0035] Further, the thickness prediction system 100 can optionally comprise a classifier 130 configured to receive as input the feature maps produced by the first feature extractor 104, the segmentation mask 110, and the thickness mask 112, and map these inputs to a pathology prediction 132. In some embodiments, the classifier 130 comprises one or more fully connected neural network layers, and can thus be referred to as a fully connected neural network. The pathology prediction 132 is a probability of one or more pathologies. One example of a pathology prediction is shown by the case prediction 902 in Figure 9 In some embodiments, the classifier 130 comprises a pre-trained deep neural network trained to map the segmentation map and the thickness map of at least the first object class of interest to a pathology prediction. The inventors herein have determined that the prediction of certain pathologies, such as pneumonia, SARS-CoV-19, etc., can benefit from an accurate estimate of the volume of the disease-affected region. In some embodiments, the object class of interest can be the disease-affected region, in which case the thickness mask 112 and the segmentation mask 110 implicitly contain information about the volume of the disease-affected region, and by feeding this information directly into the classifier 130, a more accurate pathology prediction can be determined. In the above examples of pneumonia and SARS-CoV-19, the disease-affected region can comprise inflamed lung tissue and / or fluid buildup in the lungs.

[0036] Referring to Figure 2FIG. 1 illustrates an imaging system 100, in accordance with example embodiments. In some embodiments, at least a portion of the imaging system 100 is disposed at a remote device (e.g., an edge device, a server, etc.) that is communicatively coupled to the imaging system 100 via a wired connection and / or a wireless connection. In some embodiments, at least a portion of the imaging system 100 is disposed at a separate device (e.g., a workstation) that can receive images from the imaging system 100 or from a storage device that stores images generated by one or more additional imaging systems. The imaging system 100 includes an image processing device 102, a display device 130, a user input device 140, and an imaging device 150.

[0037] The image processing device 102 includes a processor 104 configured to execute machine-readable instructions stored in a non-transitory memory 106. The processor 104 can be a single-core or multi-core processor, and programs executing thereon can be configured for parallel or distributed processing. In some embodiments, the processor 104 can optionally include separate components spread across two or more devices, which can be remotely located and / or configured for cooperative processing. In some embodiments, one or more aspects of the processor 104 can be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration.

[0038] The non-transitory memory 106 can store a deep neural network module 108, a training module 110, and image data 112. The deep neural network module 108 can include one or more deep neural networks including a plurality of weights and biases, activation functions, and instructions for implementing the one or more deep neural networks to receive a 2D medical image and map the 2D medical image into one or more of a thickness mask or a segmentation mask. For example, the deep neural network module 108 can store instructions for implementing the neural networks of the thickness prediction system 100, such as the first feature extractor 104, the second feature extractor 170, the first CNN 106, the second CNN 108, the third CNN 172, and / or the classifier 130. The deep neural network module 108 can include trained and / or untrained neural networks, and can further include various metadata for storing one or more trained or untrained deep neural networks therein.

[0039] The non-transitory memory 106 can also include a training module 110 including instructions for training one or more of the deep neural networks stored in the deep neural network module 108. The training module 110 can include instructions that, when executed by the processor 104, cause the image processing device 102 to perform the processes described below with reference to FIGS. 2-5. Figure 7One or more of the steps of the method 700 are discussed in detail. In one example, the training module 210 includes instructions to receive training data pairs from the image data 212, where the training data pairs include 2D medical images and corresponding ground truth thickness masks for training one or more of the deep neural networks stored in the deep neural network module 208. In another example, the training module 210 can include instructions to generate training data by performing one or more of the operations of the method 400 discussed in more detail below. In some embodiments, the training module 210 is not disposed at the imaging device 200, but is located remotely and communicatively coupled to the imaging system 200.

[0040] The non-transitory memory 206 can also store image data 212. The image data 212 can include medical images, such as 2D or 3D images of an anatomical region of one or more imaged individuals. In some embodiments, the images stored in the image data 212 can have been obtained by the imaging device 250. In some embodiments, the images stored in the image data 212 can have been obtained by an imaging system located remotely and communicatively coupled to the imaging system 200. The images stored in the image data 212 can include metadata related to the images stored therein. In some embodiments, the metadata of the medical images stored in the image data 212 can indicate one or more of image acquisition parameters used to acquire the images, conversion factors for converting pixels / voxels to physical dimensions (e.g., converting pixels or voxels to regions, lengths, or volumes corresponding to the length or volume of the region represented by the pixel / voxel), image acquisition date, anatomical region of interest included in the image, and the like.

[0041] In some embodiments, the non-transitory memory 206 can include components disposed on two or more devices that can be located remotely and / or configured for coordinated processing. In some embodiments, one or more aspects of the non-transitory memory 206 can include a remotely accessible networked storage device configured in a cloud computing configuration.

[0042] The imaging system 200 can further include a user input device 240. The user input device 240 can include one or more of a touchscreen, a keyboard, a mouse, a trackpad, a motion sensing camera, or other devices configured to enable a user to interact with and manipulate data within the image processing device 202. In one example, the user input device 240 can enable a user to annotate object classes of interest in a 3D medical image.

[0043] Display device 230 can include one or more display devices utilizing almost any type of technology. Display device 230 can be combined in a shared package with processor 204, non-transitory memory 206, and / or user input device 240, or can be a peripheral display device, and can include a monitor, a touchscreen, a projector, or other display devices known in the art, which can enable a user to view 2D medical images, 3D medical images, pseudo-3D medical images, and thickness heat maps, and / or interact with various data stored in non-transitory memory 206.

[0044] Imaging system 200 further includes imaging device 250. Imaging device 250 can include a 2D or 3D medical imaging device, including but not limited to an x-ray imaging device, a CT imaging device, an MRI system, an ultrasound, and a PET imaging device. Images acquired by imaging device 250 can be stored at image data 212 in non-transitory memory 206, or can be stored remotely at an external storage device communicatively coupled to imaging system 200.

[0045] It should be appreciated that Figure 2 The image processing system 200 shown in FIG. 1 is for purposes of illustration and not limitation. Another suitable image processing system can include more, fewer, or different components.

[0046] It should be appreciated that different systems can be used during the training phase and the implementation phase of one or more of the deep neural networks described herein. In some embodiments, a first system can be used to train a deep neural network by performing one or more steps of a training method, such as method 700 described below, and a second, separate system can be used to implement the deep neural network to infer a thickness mask for a 2D medical image, such as by performing one or more steps of method 300 described below. Further, in some embodiments, training data generation can be performed by a third system, different from the first and second systems, by performing one or more steps of methods 400 and / or 500 described below. Thus, the first, second, and third systems can each include different components. In some embodiments, the second system can not include a training module, such as training module 210, as the deep neural network stored on the non-transitory memory of the second system can be pre-trained by the first system. In some embodiments, the first system can not include an imaging device, and can receive images acquired by an external system that is communicably coupled to an imaging device. In some embodiments, the second system can not include or be communicably coupled to a 3D imaging device, but can use one or more trained deep neural networks to infer 3D information, such as depths / thicknesses of one or more object classes of interest, from a 2D medical image. However, in some embodiments, a single system can perform one or more or each of the training data generation, deep neural network training, and implementation of a trained deep neural network disclosed herein.

[0047] Referring to Figure 3 , a flowchart of a method 300 for inferring a thickness mask and a volume of at least a first object class of interest in a 2D medical image is shown. In some embodiments, method 300 can be implemented by an imaging system, such as imaging system 200 shown in Figure 2 . In some embodiments, a system performing method 300 can not include or be communicably coupled to a 3D imaging device, and thus can perform one or more steps of method 300 to infer depth / thickness information from a 2D medical image, such as depth / thickness information that can be acquired from a 3D imaging system.

[0048] At operation 302, the imaging system receives a 2D medical image of an anatomical region of an imaging subject. The 2D medical image can include, but is not limited to, a 2D x-ray image, a mammogram, or other 2D image. The 2D medical image received at operation 302 can include a plurality of intensity values in one or more color channels corresponding to a plurality of pixels. The plurality of intensity values can be arranged in a determined order. In some embodiments, the plurality of intensity values of the 2D medical image can comprise a 2D array or matrix, where each of the plurality of intensity values in a particular color channel can be uniquely identified by a first index and a second index, such as by a row number and a column number. In embodiments where the 2D medical image includes a plurality of color channels, the color channel to which an intensity value corresponds can also be indicated by a third index. The 2D image can include a grayscale image or a color image. In some embodiments, at operation 302, the imaging system acquires the 2D medical image using an imaging device, such as imaging device 250. In some embodiments, the imaging system receives the 2D medical image from an external device communicatively coupled to the imaging system, such as an image repository.

[0049] At operation 304, the imaging system can extract features from the 2D medical image to produce a feature map. In some embodiments, operation 304 includes the imaging system passing the 2D medical image into an input layer of a feature extractor, where the feature extractor can apply one or more filters to the 2D medical image to extract one or more features that match the one or more filters. In some embodiments, the filters can include learned filters of a convolutional layer. In some embodiments, the filters can be hard-coded based on domain knowledge. In some embodiments, the feature extractor can include both learned and hard-coded filters / parameters. In some embodiments, the feature extractor includes a deep neural network, such as an encoder, where the input image is mapped to a compressed or encoded representation by passing through one or more layers of learned weights / filters. In some embodiments, the feature extractor can output a feature map, where the feature map includes a spatially meaningful arrangement of identified / extracted features present in the 2D medical image. In some embodiments, operation 302 can also include the feature extractor concatenating one or more pieces of metadata related to the 2D medical image with the feature map. As an example, the one or more pieces of information related to the 2D medical image can be included in a DICOM header, and the information can be vectorized and concatenated with the feature map output by the feature extractor. Alternatively, the feature extractor can be configured to receive metadata related to the 2D medical image in addition to the pixel intensity data of the 2D medical image, and map the metadata and the pixel intensity data to the feature map.

[0050] At operation 306, the imaging system maps the features to a segmentation mask using a first CNN. The first CNN includes one or more convolutional layers, where each convolutional layer of the one or more convolutional layers includes one or more filters that include a plurality of learned weights having a predetermined receptive field size and stride. The first CNN is configured to map the features of the 2D medical image to a segmentation mask of at least a first object class of interest. In one embodiment, the segmentation mask includes a plurality of values or a matrix of values corresponding to a plurality of pixels of the 2D medical image, where each value of the segmentation mask indicates a classification of a corresponding pixel of the 2D medical image. In some embodiments, the segmentation mask can be a binary segmentation mask including a matrix of 1s and Os, where a 1 indicates that a corresponding pixel belongs to the object class of interest and a 0 indicates that a corresponding pixel does not belong to the object class of interest. A binary segmentation mask can be applied to a matrix of the same size, such as a matrix including pixel intensity values of the 2D medical image, by multiplying each pixel intensity value by a corresponding mask value (this process can also be referred to herein as a pixel-wise multiplication or pair-wise multiplication). The effect of the pixel-wise multiplication between the 2D medical image and the binary segmentation mask is an inhibition of pixel intensity values that are not classified by the first CNN as belonging to the first object class of interest.

[0051] At operation 308, the imaging system maps the features to a thickness mask of at least the first object class of interest. The second CNN includes one or more convolutional layers, where each convolutional layer of the one or more convolutional layers includes one or more filters that include a plurality of learnable weights having a predetermined receptive field size and stride. The second CNN can receive the features extracted by the feature extractor as a feature map, where the spatial relationship of each of the extracted features is preserved within the feature map and encoded in the relative position of each feature within the feature map. The second CNN is configured to map the features of the 2D medical image acquired at operation 302 to a thickness mask of at least the first object class of interest. The thickness mask can include a matrix of thickness values of at least the first object class of interest, where each value of the matrix of thickness values indicates a thickness of at least the first object class of interest at a corresponding pixel / position of the 2D medical image. In some embodiments, the thickness mask output by the second CNN can include a plurality of depth information encoding vectors that indicate depth-dependent locations and / or densities of at least the first object class of interest.

[0052] At operation 310, the imaging system applies the segmentation mask generated at operation 306 to the thickness mask generated at operation 308 to generate a segmented thickness mask. Applying the segmentation mask to the thickness mask suppresses the thickness values for regions corresponding to non-objects classes of interest, as the thickness values are associated with a segmentation value of 0, and thus will be cancelled (i.e., will become zero) upon the pairwise multiplication between the segmentation value and the corresponding thickness value. In this way, the imaging system reduces the noise in the thickness mask and enables more accurate volume determination for the object class of interest.

[0053] At operation 312, the imaging system determines the volume of the object class of interest based on the segmented thickness mask. In some embodiments, each thickness value of the segmented thickness mask can be summed to produce a thickness sum, and then the thickness sum can be multiplied by a conversion factor to produce the volume of the object class of interest. In some embodiments, the conversion factor can be included as metadata associated with the 2D medical image. In some embodiments, the thickness values of the segmented thickness mask can be plotted as points in a 3D space, where the z-coordinate of a point in the 3D space is given by the thickness value, and the x and y coordinates of each point in the 3D space correspond to the row and column, respectively, of the corresponding pixel in the 2D medical image. The volume of the object class of interest can then be obtained as the integral or an approximation of the integral of the 3D surface formed by the plurality of points in the 3D space.

[0054] At operation 314, the imaging system can optionally feed the features extracted by the feature extractor at operation 304, the segmentation mask generated at operation 306, and the thickness mask generated at operation 308 to the trained classifier. The trained classifier can then determine a pathology prediction, indicating a probability score for one or more diseases. In one embodiment, the trained classifier comprises a fully connected neural network comprising one or more fully connected layers. The output layer of the trained classifier can comprise one or more regression nodes, where each of the one or more regression nodes corresponds to a different pathology, and the output of the regression node is the predicted probability of the pathology. Through the pathology prediction 902, the imaging system can determine the presence or absence of one or more diseases in the 2D medical image. Figure 9 An example of a pathology prediction is shown in FIG. 9B. Turning temporarily to FIG. 9B, the imaging system can determine a pathology prediction 902 for the 2D medical image 900. The pathology prediction 902 can indicate a probability score for one or more diseases. In some embodiments, the pathology prediction 902 can be determined by a trained classifier, such as the trained classifier 110 of FIG. 1. In some embodiments, the pathology prediction 902 can be determined by a trained classifier, such as the trained classifier 110 of FIG. 1. Figure 9As can be seen, the pathology prediction 902 includes probability scores for multiple pathologies as well as a separate probability score for the non-pathology state. The pathology prediction 902 also includes the associated 2D medical image for which the pathology prediction 902 was generated. In the particular case of the pathology prediction 902, as can be seen, a probability of 99.9994% has been determined for COVID (SARS-CoV-19), a probability of 0.0006% has been determined for pneumonia, and a probability of 0.0% has been determined for the non-disease state. Each of the three probabilities of the pathology prediction 902 can be produced by a separate regression node of the output layer of the trained classifier. The pathology prediction determined at operation 314 can be displayed to a user via a display device communicatively coupled to the imaging system.

[0055] At operation 316, the imaging system displays the segmented thickness mask to the user via the display device. In some embodiments, the imaging system can generate a thickness heat map from the segmented thickness mask, overlay the thickness heat map onto the 2D medical image, and display the thickness heat map overlaid on the 2D medical image, as shown in the example embodiment thickness heat map 802 shown in Figure 8A In some embodiments, the imaging system can generate a pseudo-3D image from the segmented thickness mask, where the thickness value of each pixel of the 2D medical image is plotted as a z-coordinate in 3D space, where the x- and y-coordinates of each point plotted in 3D space correspond to the location of the associated pixel in the 2D medical image. The pseudo-3D image 804 shown in Figure 8B illustrates an example embodiment pseudo-3D image generated from a segmented thickness mask.

[0056] In this way, the method 300 enables the inference of depth / thickness information from a 2D medical image of at least a first object class of interest, thereby providing greater insight to the patient and clinician. Moreover, by inferring the depth information of the object class of interest, the volume of the object class of interest can be estimated, which can aid in the diagnosis or assessment of the patient.

[0057] Turning to Figure 4 , a method 400 for generating training data pairs for training a deep neural network to map 2D medical images to corresponding thickness masks of at least a first object class of interest is shown. The method 400 can be performed by one or more of the systems disclosed herein, such as the imaging system 200 of Figure 2 The training data pairs generated by the method 400 can be employed in a training method, such as the method 700, to train a deep neural network to map from 2D medical images to corresponding thickness masks.

[0058] The method 400 begins at operation 402, where the imaging system receives a 2D medical image of a first anatomical region of an imaged subject. In some embodiments, the 2D medical image is an x-ray image. The 2D medical image can include metadata related to the acquisition of the 2D medical image, where the metadata can indicate the anatomical region imaged, one or more imaging parameters used during the acquisition of the 2D medical image, the date of the image acquisition, etc. The 2D medical image can be stored on a non-transitory memory of the imaging system and / or transmitted to a remote device communicatively coupled to the imaging system, such as a remote image repository. The imaging system can acquire the 2D medical image via a 2D imaging device communicatively coupled thereto or from the image repository.

[0059] At operation 404, the imaging system receives a 3D medical image of the first anatomical region of the imaged subject. In some embodiments, the 2D medical image and the 3D medical image are acquired within a threshold time window, thereby reducing differences that can occur in the first anatomical region between the acquisition of the 2D medical image and the 3D medical image. In some embodiments, the threshold time window is based on a rate of change / growth of one or more anatomical structures of the first anatomical region and / or a disease affecting one or more anatomical structures of the first anatomical region. In one example, for rapidly developing diseases, such as pneumonia, the threshold time window can be less than 48 hours. In another example, for more slowly developing / changing diseases, such as a slowly growing tumor, the threshold time window can be 3 months. The threshold time window for a non-disease-affected anatomical region can be greater than the threshold time window for a disease-affected anatomical region when the disease does affect the first anatomical region compared to when the disease does not affect the first anatomical region, as the rate of change of the tissue / organ affected by the disease can be greater than the rate of potential growth / change in the tissue / organ. In another example, the threshold time window for a child can be shorter than the threshold time window for an adult, as the rate of change of the anatomical structures of the first anatomical region can be greater in children than in adults.

[0060] The 3D medical image can be received from a 3D imaging device using one or more known 3D imaging modalities, including but not limited to CT, MRI, PET, ultrasound, mammography, etc. The imaging system used to acquire the 2D medical image at operation 402 can be the same or different from the imaging modality used to acquire the 3D medical image at operation 404. In some embodiments, the 3D medical image is a CT image including a plurality of voxels representing a first anatomical region of an imaged subject in 3D. The 3D medical image can include metadata related to the acquisition of the 3D medical image, where the metadata can indicate the anatomical region imaged, one or more imaging parameters used during the acquisition of the 3D medical image, the date of image acquisition, etc. The 3D medical image can be stored on a non-transitory memory of the imaging system and / or transmitted to a remote device communicatively coupled to the imaging system, such as a remote image repository. In some embodiments, the 2D medical image acquired at operation 402 and the 3D medical image acquired at operation 404 are both associated with a unique identification number, thereby linking the 2D medical image and the 3D medical image.

[0061] At operation 406, the imaging system annotates voxels of the 3D medical image with one or more object class labels. In some embodiments, the imaging system annotates the voxels of the 3D medical image in response to input received from a user via a user input device. In some embodiments, the imaging system automatically annotates the voxels of the 3D medical image based on a 3D segmentation mask determined by a trained deep neural network. In some embodiments, the voxels of the 3D medical image are automatically annotated based on an unsupervised learning algorithm. The annotations can include a label, a flag, or a value associated with one or more voxels of the 3D medical image. In one example, the object class annotations can include a 3D array of values, where each value can indicate an object class label, and the location of a point within the 3D array can correspond to a spatial location of a voxel in the 3D medical image.

[0062] At operation 408, the imaging system projects the 3D medical image onto a 2D plane to produce a synthetic 2D image. Figure 5 A method for generating a synthetic 2D image from a 3D image is described in the following detailed description. Briefly, an imaging system can select one or more projection parameters, such as a radiation source location, an angle of incidence of a plurality of rays emitted by the radiation source, and a location and orientation of a 2D projection plane. The rays emitted from the radiation source can pass through voxels of the 3D medical image and intersect the 2D projection plane, where for each ray passing through the 3D medical image and onto the 2D projection plane, a synthetic pixel having an associated synthetic intensity value can be determined based on the voxels of the 3D image through which the ray passes. In some embodiments, the intensity value of the synthetic pixel for a ray can be based on an average and / or a sum of the intensity values of the voxels of the 3D image through which the ray passes before intersecting the 2D projection plane.

[0063] At operation 410, the annotations of at least a first object class of interest of the 3D medical image are projected onto the 2D plane using the same projection parameters applied at operation 408 to produce a ground truth thickness mask. In some embodiments, the ground truth thickness mask is produced by emitting rays from a radiation source, through the 3D medical image, and onto the 2D plane, where for each ray incident on the 2D plane, a thickness value is determined based on the number of voxels (annotated as belonging to the first object class of interest) that the ray passes through. The plurality of rays emitted from the radiation source can thus be converted into a plurality of thickness values for the first object class of interest, and the plurality of thickness values, along with their spatial relationship as indicated by their locations on the 2D plane, constitute the ground truth thickness mask. Although the above describes a process for generating a ground truth thickness mask for the first object class of interest, it should be understood that the same process can be used to generate a plurality of ground truth thickness masks for a plurality of object classes of interest.

[0064] At operation 412, the imaging system registers the ground truth thickness mask produced at operation 410 with the 2D medical image acquired at operation 402. Registration includes aligning the two images such that the sum of the pixel-wise differences between the two images is minimized, or conversely, such that the alignment between the anatomical regions captured by the two images is maximized. By registering the ground truth thickness mask with the 2D medical image, the alignment between the regions of the first object class of interest depicted in the 2D medical image and in the ground truth thickness mask can be maximized. In some embodiments, operation 412 can include registering the synthetic 2D image generated at operation 408 with the 2D medical image acquired at operation 402 to obtain registration parameters (e.g., the extent of x and y translations for producing a minimization of the pixel-wise differences between the 2D medical image and the synthetic 2D image), and applying these registration parameters to the ground truth thickness mask to align the thickness values of the ground truth thickness mask with their corresponding pixels in the 2D medical image.

[0065] At operation 414, the aligned ground truth thickness mask and the 2D medical image are stored together as a training data pair. In some embodiments, metadata pertaining to the training data pair can be stored along with the 2D medical image and the ground truth thickness mask. As an example, the metadata can include an indication of the object class of interest, an indication of the anatomical region captured by the 2D medical image, a date of acquisition, a type of disease associated with the training data pair, and the like. After operation 414, method 400 can end.

[0066] In this way, training data pairs comprising 2D medical images and corresponding ground truth thickness masks can be generated. The inventors herein found that, in order for a deep neural network to learn an accurate mapping from 2D medical images to thickness masks, synthetic 2D images are not sufficient for use in training data pairs directly, as synthetic 2D medical images and real 2D medical images are sufficiently different in appearance to reduce the accuracy of thickness inference during implementation on real 2D medical images. Accordingly, the inventors developed the methods disclosed herein, such as the method 400 described above, such that real 2D medical images can be paired with accurate thickness information (in the form of a thickness mask) for at least a first class of objects of interest captured by the real 2D images.

[0067] Turning to Figure 5 , an exemplary method 500 for determining projection parameters for generating synthetic 2D images from 3D images is shown. The method 500 can be performed by one or more of the systems described herein, such as the imaging system 200 shown in Figure 2 . The method 500 can be performed as part of a method of generating training data pairs for training a deep neural network to infer depth information from 2D medical images, such as at operation 408 of the method 400.

[0068] The method 500 begins at operation 502, where the imaging system selects an initial set of projection parameters. Projection parameters include, but are not limited to, the position of a radiation source relative to an imaging subject, the position and orientation of a 2D projection plane relative to the radiation source and the imaging subject, and the angle / projected direction of a plurality of rays emitted by the radiation source. Turning briefly to Figure 6A , an exemplary schematic of a projection process is shown. Figure 6A A radiation source 602 is shown positioned at a distance 604 away from an imaging subject 608, with the imaging subject 608 positioned between the radiation source 602 and a projection plane 610. As can be seen in Figure 6A , changing any of the position of the radiation source 602, the position or orientation of the imaging subject 608, the position or orientation of the projection plane 610, and the trajectory of the plurality of rays 606 can change the projection of the imaging subject 608 formed on the projection plane 610.

[0069] At operation 504, the imaging system projects a 3D medical image onto a 2D plane using the currently selected projection parameters to generate a synthetic 2D image. Again turning to Figure 6AThe radiation source 602 emits a plurality of rays 606, and a subset of the plurality of rays 606 intersects the imaging object 608. In some embodiments, the imaging object 608 can include a plurality of voxels of a 3D medical image acquired via a 3D imaging device, and as one or more of the plurality of rays pass through a voxel of the imaging object 608, a history of the travel path of the ray can be determined and / or recorded. Upon passing through the imaging object 608 and intersecting the projection plane 610, a projection of the imaging object 608 can be produced on the projection plane 610 by drawing a value of each incident ray at the intersection location between the ray and the projection plane 610, where the value of the incident ray can be determined based on the travel history of the ray. Turning to Figure 6B An example synthetic 2D image 640 is shown. The synthetic 2D image 640 can be generated according to the process shown. Figure 6A Each of the synthetic 2D images 640 includes a different synthetic image generated from a single imaging object, but applying a different set of projection parameters. The synthetic 2D images 640 provide an example embodiment of synthetic 2D images that can be produced at operation 504 of the method 500.

[0070] At operation 506, the imaging system determines a difference between the synthetic 2D image produced at operation 504 using the currently selected projection parameters and a corresponding 2D medical image (e.g., the medical image acquired at operation 402 of the method 400). In some embodiments, the difference between the synthetic 2D image and the corresponding 2D medical image can be determined using one or more of a weighted average of DICE scores, a pixel-wise mean squared error, and a degree of x and / or y translation determined by registering the synthetic 2D image with the 2D medical image.

[0071] At operation 508, the imaging system assesses whether the difference determined at operation 506 is less than a threshold difference. If it is determined at operation 508 that the difference is not less than the difference threshold, the method 500 proceeds to determining new projection parameters at operation 510, and returns to operation 504 to produce a new synthetic 2D image using the updated projection parameters. However, if at operation 508 the imaging system determines that the difference is less than the threshold difference, the method 500 proceeds to operation 512.

[0072] At operation 512, the synthetic 2D image and the projection parameters used to obtain the synthetic 2D image are stored in a non-transitory memory of the imaging system. After operation 512, the method 500 can end.

[0073] The degree of correspondence / matching between the ground truth thickness mask and the 2D medical image can be increased by iteratively adjusting the projection parameters until a synthetic 2D image is produced that has sufficient similarity to the 2D medical image (e.g., the difference between the synthetic 2D image and the 2D medical image is below a threshold), where the ground truth thickness mask is produced by projecting the annotation of the 3D medical image onto the 2D plane using the projection parameters determined by the method 500.

[0074] Referring to Figure 7 , a flowchart of an exemplary method 700 for training a deep neural network, such as the second CNN 108, to infer a thickness mask of a class of objects of interest from a 2D medical image is shown. The method 700 can be implemented by the imaging system 200 shown in Figure 2 based on instructions stored in the non-transitory memory 206.

[0075] The method 700 begins at operation 702, where a training data pair from a plurality of training data pairs is fed to the deep neural network, where the training data pair includes a 2D medical image of an anatomical region of an imaged subject, and a corresponding ground truth thickness mask that indicates a thickness of at least a first class of objects of interest at each pixel of a plurality of pixels of the 2D medical image. In some embodiments, the training data pair and the plurality of training data pairs can be stored in the imaging system, such as in the imaging data 212 of the imaging system 200. In other embodiments, the training data pair can be acquired via a communicative coupling between the imaging system and an external storage device (e.g., via an internet connection with a remote server). In some embodiments, the ground truth thickness mask includes a deep encoding vector for each pixel of the plurality of pixels of the 2D medical image, thereby enabling the deep neural network to learn a depth variable density or distribution of the class of objects of interest.

[0076] At operation 704, the imaging system uses the feature extractor to extract features from the 2D medical image, similar to operation 304 of the method 300 described above. In some embodiments, the feature extractor includes one or more learnable / adjustable parameters, and in such embodiments, the parameters can be learned by performing the method 700. In some embodiments, the feature extractor includes hard-coded parameters and does not include learnable / adjustable parameters, and in such embodiments, the feature extractor is not trained during performance of the method 700.

[0077] At operation 706, the imaging system uses a deep neural network to map the features to a predicted thickness mask of at least a first object class of interest. In some embodiments, the deep neural network comprises a CNN comprising one or more convolutional layers comprising one or more convolutional filters. The deep neural network maps the features to the predicted thickness mask by propagating the features from an input layer through one or more hidden layers until reaching an output layer of the deep neural network.

[0078] At operation 708, the imaging system computes a loss for the predicted thickness mask based on a difference between the predicted thickness mask and a ground truth thickness mask. In some embodiments, the loss comprises a mean squared error given by the following equation:

[0079]

[0080] where MSE represents the mean squared error, N is a total number of training data pairs, i is an index indicating a currently selected training data pair, x i is the predicted thickness mask for the training data pair i, and X i is the ground truth thickness mask for the training data pair i. The expression x i - X i will be understood to represent a pair-wise subtraction of each pair of corresponding thickness values in the predicted thickness mask and the ground truth thickness mask for the currently selected training data pair i. It will be appreciated that other loss functions known in the art of machine learning can be employed at operation 708.

[0081] At operation 710, the weights and biases of the deep neural network are adjusted based on the loss determined at operation 708. In some embodiments, the parameters of the feature extractor and the CNN can be adjusted to reduce the loss on the training dataset. In some embodiments, the feature extractor can not include learnable parameters, and thus operation 710 can not include adjusting the parameters of the feature extractor. In some embodiments, the backpropagation of the loss can occur according to a gradient descent algorithm, in which the gradient (first derivative or approximation of the first derivative) of the loss function is determined for each weight and bias of the deep neural network. Each weight (and bias) of the deep neural network is then updated by adding the negative of the product of the gradient determined (or approximation) for the weight (or bias) and a predetermined step size to the weight (and bias). The method 700 can then end. It should be noted that the method 700 can be repeated for each training data pair in the training dataset, and the process can be repeated until a stopping condition is met. In some embodiments, the stopping condition includes one or more of the loss reducing below a threshold loss, the rate of change of the loss reducing below a threshold rate of change of the loss, a validation loss determined on a validation dataset reaching a minimum, etc. In this way, the feature extractor can learn to extract features related to the thickness of the object class of interest, and the CNN can learn to map the features to a thickness mask of the 2D medical image.

[0082] Turning to Figure 10 , an exemplary embodiment of a spatial regularization method 1000 that can be applied to the output of a CNN layer, such as the first CNN 106 and / or the second CNN 108, is shown. The inventors herein determined that by applying a spatial regularization constraint, the noise of the output parameters determined by a trained neural network, such as the thickness values of the class of interest, can be reduced, where the output values are modified based on other output values in spatially local regions of values (e.g., neighboring pixels / voxels). Figure 10A 2D medical image 1002 is shown that includes a first region 1004, a second region 1006, a third region 1008, and a fourth region 1010, with a first filter 1014 applied to the first region 1004, a second filter 1016 applied to the second region 1006, a third filter 1018 applied to the third region 1008, and a fourth filter 1020 applied to the fourth region 1010, in order to produce a first feature f0, a second feature f1, a third feature f2, and a fourth feature f3, respectively. Spatial regularization factors W0, W1, W2, and W3 are applied to the corresponding features to produce spatially regularized outputs. More specifically, in the example shown in the spatial regularization method 1000, the first feature f0 is multiplied by a first spatial regularization factor W0, the second feature f1 is multiplied by a second spatial regularization factor W1, the third feature f2 is multiplied by a third spatial regularization factor W2, and the fourth feature f3 is multiplied by a fourth spatial regularization factor W3, to produce a corresponding plurality of spatially regularized features, which can be used as feature maps for a subsequent layer or can include output values such as thickness values for a class of interest. The spatial regularization factors can be determined from features extracted in neighboring regions of the input feature maps or images. In some embodiments, the spatial regularization factors are determined such that the absolute value of the difference between any two approximate feature values is less than a threshold difference.

[0083] Turning to Figure 11 , a process 1100 is shown by which a depth information encoding vector 1130 can be produced. The process 1100 includes obtaining an intensity distribution 1106 for a line 1104 taken across a depthwise direction image 1160, where the position of the line 1104 corresponds to a point 1140 of a 2D medical image 1102. The intensity distribution 1106 encodes depth information for an object structure extending into the plane of the 2D medical image 1160 at the point 1140. The intensity distribution 1106 can be quantized into a finite number of discrete intensity bands, such as a first intensity band 1108, a second intensity band 1110, and a third intensity band 1112. Each different intensity band can represent a different object class of interest, and / or a different density of a single object class of interest. Each of the first intensity band 1108, the second intensity band 1110, and the third intensity band 1112 can be used to generate a depth information encoding vector 1130 corresponding to the point 1140 of the 2D medical image 1160. Similar processing can be performed for each point of the 2D medical image 1160 to produce a plurality of depth information encoding vectors. The depth information encoding vectors can be used in place of or in addition to the thickness values of the thickness masks described herein to enable a deep neural network to not only infer thicknesses of object classes of interest, but also to infer depth positions and distributions of object classes of interest at each point of a 2D medical image.

[0084] When introducing elements of various embodiments of the present disclosure, the articles “a,” “an,” and “the” are intended to mean that there are one or more of the elements. The terms “first,” “second,” and the like do not denote any order, quantity, or importance, but are used to distinguish one element from another. The terms “including” and “comprising” are intended to be inclusive and mean that there can be additional elements other than the listed elements. As used herein, the terms “connected to,” “coupled to,” and the like are intended to mean that one object (e.g., material, element, structure, member, etc.) can be directly or indirectly connected or coupled to another object, whether it is connected or coupled directly, or is connected or coupled through one or more intervening objects. Further, it will be understood that references to “one embodiment” or “an embodiment” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the referenced features.

[0085] In addition to any prior indicated modifications, numerous variations and alternative arrangements can be designed by those of ordinary skill in the art without departing from the spirit and scope of the description, and it is intended that the appended claims cover such modifications and arrangements. Therefore, although the information has been described above in detail and with reference to specific aspects thereof, it will be apparent to one of ordinary skill in the art that many modifications, including but not limited to, form, function, manner of operation and use, can be made without departing from the principles and concepts set forth herein. Also, as used herein, in all aspects, the examples and embodiments are intended to be illustrative and not restrictive, in any way.

Claims

1. A method for determining a volume of a class of objects of interest from a 2D medical image, the method comprising: receiving a 2D medical image; extracting features from the 2D medical image; mapping the features to a segmentation mask of a class of objects of interest using a first convolutional neural network (CNN); mapping the features to a thickness mask of the class of objects of interest using a second CNN, wherein the thickness mask indicates a thickness of the class of objects of interest at each pixel of a plurality of pixels of the 2D medical image; and determining a volume of the class of objects of interest based on the thickness mask and the segmentation mask, wherein estimating the volume of the class of objects of interest based on the thickness mask and the segmentation mask comprises: multiplying each value of the thickness mask by a spatially corresponding value of the segmentation mask to produce a plurality of segmented thickness values; and summing the plurality of segmented thickness values to produce the volume of the class of objects of interest.

2. The method of claim 1, further comprising: generating a pseudo 3D medical image from the 2D medical image and the thickness mask by: rendering thickness values of the thickness mask as a surface in 3D space; overlaying the surface on the 2D medical image to produce the pseudo 3D medical image; and displaying the pseudo 3D medical image via a display device.

3. The method of claim 1, further comprising: generating a thickness heat map of the class of objects of interest from the thickness mask; and displaying the 2D medical image with the thickness heat map overlaid thereon.

4. The method of claim 1, further comprising: mapping the segmentation mask, the thickness mask, and the features to a pathology prediction using a trained classifier.

5. A method of training a deep neural network to learn a mapping between a 2D medical image and a thickness mask of a first class of objects of interest, the method comprising: receiving a 2D medical image of a first region of an imaged object; receiving a 3D medical image of the first region of the imaged object; annotating voxels of the 3D medical image by a class label of a first class of objects of interest to produce a first plurality of annotated voxels; projecting the 3D medical image along a plurality of rays onto a plane to produce a synthetic 2D medical image matching the 2D medical image; projecting the first plurality of annotated voxels along the plurality of rays onto the plane to produce a first plurality of thickness values of the first class of objects of interest; producing a first ground-truth thickness mask of the first class of objects of interest from the first plurality of thickness values; and training a deep neural network to learn a mapping between a 2D medical image and a thickness mask of the first class of objects of interest by: mapping the 2D medical image to a first predicted thickness mask of the first class of objects of interest; determining a loss of the first predicted thickness mask based on a difference between the first predicted thickness mask and the first ground-truth thickness mask; and updating parameters of the deep neural network based on the loss. ​ ​ ​ ​ ​ ​ 6. The method of claim 5, wherein projecting the 3D medical image along the plurality of rays onto the plane to produce the synthetic 2D medical image that matches the 2D medical image comprises: selecting a first position of a simulated radiation source relative to the 3D medical image; selecting a second position and a first orientation of the plane relative to the simulated radiation source and the 3D medical image; and projecting the plurality of rays from the simulated radiation source through the 3D medical image onto the plane to produce the synthetic 2D medical image.

7. The method of claim 6, further comprising: determining a difference between the synthetic 2D medical image and the 2D medical image; and in response to the difference between the synthetic 2D medical image and the 2D medical image being less than a threshold value: setting the simulated radiation source to the first position; setting the plane to the second position and the first orientation; and projecting the plurality of rays from the simulated radiation source through the first plurality of annotated voxels onto the plane to generate the first plurality of thickness values.

8. The method of claim 7, wherein the first plurality of thickness values are arranged in a matrix, wherein each thickness value of the first plurality of thickness values indicates a length of the first object class of interest that is traversed by a corresponding ray of the plurality of rays projected from the simulated radiation source through the first plurality of annotated voxels onto the plane.

9. The method of claim 8, wherein the length of the first object class of interest that is traversed by the corresponding ray of the plurality of rays projected from the simulated radiation source through the first plurality of annotated voxels onto the plane is proportional to a number of voxels of the first plurality of annotated voxels that the ray traverses when traveling from the simulated radiation source to the plane.

10. The method of claim 7, wherein the first ground truth thickness mask comprises a plurality of vectors, wherein each of the plurality of vectors encodes a length of one or more object class labels that is traversed by a ray projected from the simulated radiation source through the object class labels onto the plane.

11. The method of claim 7, wherein the first ground truth thickness mask comprises a plurality of vectors, wherein each of the plurality of vectors encodes a depth-dependent density of the first object class of interest that is traversed by a ray projected from the simulated radiation source through the object class labels onto the plane.

12. The method of claim 5, further comprising: annotating voxels of the 3D medical image by object class labels of a second object class of interest to produce a second plurality of annotated voxels; projecting the second plurality of annotated voxels along the plurality of rays onto the plane to produce a second plurality of thickness values of the second object class of interest; producing a second ground truth thickness mask of the second object class of interest from the second plurality of thickness values; and ​ ​ ​ ​ training the deep neural network to learn a mapping between 2D medical images and thickness masks of the second object class of interest by: mapping the 2D medical images to second predicted thickness masks of the second object class of interest; determining a loss for the second predicted thickness masks based on a difference between the second predicted thickness masks and the second ground truth thickness masks; and updating parameters of the deep neural network based on the loss.

13. The method of claim 12, wherein the first object class of interest is disease-affected tissue, and wherein the second object class of interest is non-disease-affected tissue.

14. The method of claim 5, wherein the deep neural network comprises a plurality of convolutional filters, wherein a sensitivity of each convolutional filter of the plurality of convolutional filters is modulated by a corresponding spatial regularization factor.

15. The method of claim 5, wherein generating the first ground truth thickness mask of the first object class of interest from the first plurality of thickness values comprises: registering the synthetic 2D medical image with the 2D medical image to determine a translation; and applying the translation to the first plurality of thickness values to generate the first ground truth thickness mask.

16. A medical imaging system, the medical imaging system comprising: an imaging device; a display device; a memory storing: a feature extractor; a first trained convolutional neural network (CNN); a second trained CNN; and instructions; a processor communicably coupled to the imaging device, the display device, and the memory, and, when executing the instructions, the processor is configured to: acquire, via the imaging device, a 2D medical image of an anatomical region of an imaged subject; extract, using the feature extractor, features from the 2D medical image; map, using the first trained CNN, the features to a segmentation mask of an object class of interest; map, using the second trained CNN, the features to a thickness mask of the object class of interest, wherein the thickness mask indicates a thickness of the object class of interest at each pixel of a plurality of pixels of the 2D medical image; apply the segmentation mask to the thickness mask to produce a segmented thickness mask; and display, via the display device, the segmented thickness mask.

17. The medical imaging system of claim 16, wherein the features comprise a total object thickness at each pixel of the plurality of pixels of the 2D medical image.

18. The medical imaging system of claim 16, wherein, when executing the instructions, the processor is further configured to: determine a volume of the object class of interest by approximating an integral of the segmented thickness mask. ​ ​ ​ ​ 19. The medical imaging system of claim 16, wherein the segmented thickness mask comprises a matrix of thickness values of the object class of interest, and wherein the processor is configured to display the segmented thickness mask as a pseudo-3D image by plotting each thickness value of the matrix of thickness values at a z-position corresponding to the thickness value.

Citation Information

Patent Citations

  • Segmentation-based corneal mapping

    US20190209006A1

  • Method for measuring volume of organ by using artificial neural network, and apparatus therefor

    WO2020122606A1