Deep convolutional neural networks for tumor segmentation via positron emission tomography
Through deep convolutional neural network combined with PET and CT or MRI scan, the challenge of tumor segmentation in whole-body PET-CT images is solved, and rapid and accurate tumor segmentation and treatment evaluation is achieved, improving the efficiency of tumor diagnosis and treatment.
Patent Information
- Application Number
- CN202080021150.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-31
- Filing Date
- 2020-03-14
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2040-03-14
AI Technical Summary
The prior art is difficult to accurately and automatically segment tumors in positron emission tomography images, especially in whole-body PET-CT images. Due to the unclear boundaries between tumors and normal tissues, low image resolution, and the diversity of tumors in size, shape and location, the definition of tumor segmentation is inaccurate, affecting the average SUV calculation and treatment evaluation.
The deep convolutional neural network architecture is adopted, combined with PET and CT or MRI scans, and tumor masks are generated through two-dimensional and three-dimensional segmentation models. The residual block and jump-connected convolutional neural network are used for tumor segmentation. The multimodal imaging technology is combined to improve segmentation accuracy, adapting to the size of the whole body scan and the imbalance between tumor and healthy tissues.
Fast and accurate tumor segmentation is achieved, enabling the evaluation of treatment efficacy, predict progression-free survival, staged treatment and select clinical trial subjects, and providing automatic treatment response evaluation, improving the efficiency of tumor diagnosis and treatment planning.
Smart Images

Figure CN113711271B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 880,898, filed on July 31, 2019, entitled “DEEP CONVOLUTIONAL NEURAL NETWORKS FOR TUMOR SEGMENTATION WITH POSITRON EMISSION TOMOGRAPHY,” and U.S. Provisional Application No. 62 / 819,275, filed on March 25, 2019, entitled “AUTOMATED TUMOR SEGMENTATION WITH FLUORODEOXYGLUCOSE POSITRON EMISSION TOMOGRAPHY,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] The present disclosure relates to automatic tumor segmentation, and more particularly to systems and methods for segmenting tumors in positron emission tomography images using deep convolutional neural networks for image and lesion metabolic analysis. Background Art
[0004] Positron emission tomography (PET), also known as PET imaging or PET scans, is a nuclear medicine imaging test that helps reveal the function of tissues and organs. PET scans use radioactive drugs (tracers) to visualize this activity. A tracer is a molecule linked to or labeled with a radioactive tag that can be detected on a PET scan. The tracer can be injected, swallowed, or inhaled, depending on the organ or tissue being studied. The tracer accumulates in areas of the body with higher metabolic activity or binds to specific proteins in the body (for example, cancerous tumors or areas of inflammation), which often correspond to areas of disease. These areas appear as bright spots on a PET scan. The most commonly used radioactive tracer is fluorodeoxyglucose (FDG), a molecule similar to glucose. In FDG-PET, tissues or areas with higher metabolic activity than their surroundings appear as bright spots. For example, cancer cells can absorb glucose at a higher rate, resulting in higher metabolic activity. This higher rate is visible on a PET scan and allows healthcare providers to identify tumors when they may not be visible on other imaging tests. PET scans can help diagnose and determine the severity of a variety of diseases, including several types of cancer, heart disease, gastrointestinal disorders, endocrine disorders, neurological disorders, and other abnormalities in the body. Summary of the Invention
[0005] In various embodiments, a computer-implemented method is provided that includes: obtaining a plurality of positron emission tomography (PET) scans and a plurality of computed tomography (CT) or magnetic resonance imaging (MRI) scans of a subject; preprocessing the PET scans and the CT or MRI scans to generate a first subset of standardized images of a first plane or region of the subject and a second subset of standardized images of a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scans and the CT or MRI scans; and generating a second two-dimensional segmentation model using a first two-dimensional segmentation model implemented as part of a convolutional neural network architecture that takes as input the first subset of standardized images. A two-dimensional segmentation mask, wherein the first two-dimensional segmentation model uses a first residual block comprising a first layer, wherein the first layer: (i) directly feeds into a subsequent layer, and (ii) directly feeds into a layer multiple layers away from the first layer using a skip connection; generating a second two-dimensional segmentation mask using a second two-dimensional segmentation model implemented as part of the convolutional neural network architecture that takes as input a second subset of the normalized image, wherein the second two-dimensional segmentation model uses a second residual block comprising a second layer, wherein the second layer: (i) directly feeds into a subsequent layer, and (ii) directly feeds into a layer multiple layers away from the second layer using a skip connection; and generating a final imaging mask by combining information from the first two-dimensional segmentation mask and the second two-dimensional segmentation mask.
[0006] In some embodiments, the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
[0007] In some embodiments, the method further comprises: determining a total metabolic tumor burden (TMTV) using the final imaging mask; and providing the TMTV.
[0008] In some embodiments, the method further includes: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, wherein the three-dimensional organ segmentation model takes the PET scan and the CT or MRI scan as input; using the final imaging mask and the three-dimensional organ mask to determine the metabolic tumor burden (MTV) and the number of lesions of one or more organs in the three-dimensional organ segmentation; and providing the MTV and the number of lesions of the one or more organs.
[0009] In some embodiments, the method further comprises: using a classifier having one or more of the TMTV, the MTV, and the number of lesions as input, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: the likelihood of progression-free survival (PFS) of the subject; the disease stage of the subject; and a selection decision to include the subject in a clinical trial.
[0010] In some embodiments, the method further includes: inputting, by a user, multiple PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; and receiving, by the user, one or more of the final imaging mask, the TMTV, the MTV, and the number of lesions on a display of a computing device.
[0011] In some embodiments, the method further comprises administering, by the user, a treatment to the subject based on one or more of the final imaging mask, the TMTV, the MTV, and the number of lesions.
[0012] In some embodiments, the method further comprises providing, by the user, a diagnosis of the subject based on one or more of the final imaging mask, the TMTV, the MTV, and the number of lesions.
[0013] In various embodiments, a computer-implemented method is provided that includes obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of a subject; preprocessing the PET scans and the CT or MRI scans to generate standardized images that incorporate information from the PET scans and the CT or MRI scans; generating one or more two-dimensional segmentation masks using one or more two-dimensional segmentation models implemented as part of a convolutional neural network architecture that take as input the standardized images; generating one or more three-dimensional segmentation masks using one or more three-dimensional segmentation models implemented as part of the convolutional neural network architecture that take as input image data blocks associated with segments from the two-dimensional segmentation masks; and generating a final imaging mask by combining information from the one or more two-dimensional segmentation masks and the one or more three-dimensional segmentation masks.
[0014] In some embodiments, the one or more three-dimensional segmentation models include a first three-dimensional segmentation model and a second three-dimensional segmentation model; the image data block includes: a first image data block associated with the first segment; and, a second image data block associated with the second segment; and generating the one or more three-dimensional segmentation masks includes: generating a first three-dimensional segmentation mask using the first three-dimensional segmentation model with the first image data block as input, and generating a second three-dimensional segmentation mask using the second three-dimensional segmentation model with the second image data block as input.
[0015] In some embodiments, the method further includes: evaluating the position of the region or body part captured in the standardized image as a reference point; dividing the region or body into a plurality of anatomical regions based on the reference point; generating position labels for the plurality of anatomical regions; incorporating the position labels into the two-dimensional segmentation mask; determining that the first segment is located in a first anatomical region among the plurality of anatomical regions based on the position labels; determining that the second segment is located in a second anatomical region among the plurality of anatomical regions based on the position labels; based on the determination that the first segment is located in the first anatomical region, inputting the first image data block associated with the first segment into the first three-dimensional segmentation mask; and based on the determination that the second segment is located in the second anatomical region, inputting the second image data block associated with the second segment into the second three-dimensional segmentation mask.
[0016] In some embodiments, the standardized images include a first subset of standardized images of a first plane or region of the subject and a second subset of standardized images of a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; the one or more two-dimensional segmentation models include a first two-dimensional segmentation model and a second two-dimensional segmentation model; and generating the one or more two-dimensional segmentation masks includes: generating a first two-dimensional segmentation mask using the first two-dimensional segmentation model implemented to take the first subset of the standardized images as input, wherein the first two-dimensional segmentation model uses a first residual block comprising a first layer, the first layer: (i) directly feeds into a subsequent layer, and (ii) directly feeds into a layer multiple layers away from the first layer using a skip connection; and generating a second two-dimensional segmentation mask using the second two-dimensional segmentation model taking the second subset of the standardized images as input, wherein the second two-dimensional segmentation model uses a second residual block comprising a second layer, the second layer: (i) directly feeds into a subsequent layer, and (ii) directly feeds into a layer multiple layers away from the second layer using a skip connection.
[0017] In some embodiments, the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
[0018] In some embodiments, the method further comprises: determining a total metabolic tumor burden (TMTV) using the final imaging mask; and providing the TMTV.
[0019] In some embodiments, the method further includes: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, wherein the three-dimensional organ segmentation model takes the PET scan and the CT or MRI scan as input; using the final imaging mask and the three-dimensional organ mask to determine the metabolic tumor burden (MTV) and the number of lesions of one or more organs in the three-dimensional organ segmentation; and providing the MTV and the number of lesions of the one or more organs.
[0020] In some embodiments, the method further comprises: using a classifier having one or more of the TMTV, the MTV, and the number of lesions as input, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: the likelihood of progression-free survival (PFS) of the subject; the disease stage of the subject; and a selection decision to include the subject in a clinical trial.
[0021] In some embodiments, the method further includes: inputting, by a user, multiple PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; and receiving, by the user, one or more of the final imaging mask, the TMTV, the MTV, and the number of lesions on a display of a computing device.
[0022] In some embodiments, the method further comprises administering, by the user, a treatment to the subject based on one or more of the final imaging mask, the TMTV, the MTV, and the number of lesions.
[0023] In some embodiments, the method further comprises providing, by the user, a diagnosis of the subject based on one or more of the final imaging mask, the TMTV, the MTV, and the number of lesions.
[0024] In some embodiments, a system is provided that includes: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.
[0025] In some embodiments, a computer program product is provided, which is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein.
[0026] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein.
[0027] The terms and expressions used are used as terms of description and not of limitation, and in the use of such terms and expressions, there is no intention to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications can be made within the scope of the invention claimed. Therefore, it should be understood that although the invention claimed has been specifically disclosed by way of embodiments and optional features, modifications and variations of the concepts disclosed herein may be made by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The present disclosure is described in conjunction with the following drawings:
[0029] Figure 1 shows an example computing environment for automatic tumor segmentation using a cascaded convolutional neural network in accordance with various embodiments;
[0030] Figure 2A shows an exemplary schematic diagram representing a convolutional neural network (CNN) architecture according to various embodiments;
[0031] Figure 2B shows an exemplary schematic diagram representing a CNN architecture according to various embodiments;
[0032] Figure 2C shows an exemplary detailed schematic diagram representing a CNN architecture according to various embodiments;
[0033] Figure 3 An exemplary U-Net according to various embodiments is shown;
[0034] Figure 4A shows a residual block according to various embodiments;
[0035] Figure 4B shows pyramidal layers according to various embodiments;
[0036] Figure 5 shows a plurality of V-Nets according to various embodiments;
[0037] Figure 6 shows a process for determining an extracted total metabolic tumor volume (TMTV) according to various embodiments;
[0038] Figure 7 A process for predicting the likelihood of progression-free survival (PFS) of a subject, the clinical efficacy of a treatment, the disease stage of a subject, and / or a selection decision for inclusion of a subject in a clinical trial is shown according to various embodiments;
[0039] Figure 8 A process for providing automated end-of-treatment response assessment, which may include the Lugano classification for lymphoma staging, is shown according to various embodiments;
[0040] 9A to 9F shows segmentation results according to various embodiments;
[0041] Figure 10A and Figure 10B Figure 5 shows the tumor volume and standardized uptake value (SUV) from the predicted mask according to various embodiments, respectively. max ;
[0042] Figure 11A and Figure 11B shows clustered Kaplan-Meier estimators showing the correlation of TMTV with prognosis according to various embodiments;
[0043] Figure 12A and Figure 12B It is shown that automated TMTV provides prognostic metrics at baseline consistent with manual TMTV assessment according to various embodiments;
[0044] Figure 13A and Figure 13B shows that baseline TMTV is prognostic in non-small cell lung cancer (NSCLC) and melanoma according to various embodiments; and
[0045] Figures 14A to 14F Shown are Kaplan-Meier analyses of the association of extranodal involvement with the probability of progression-free survival in the GOYA (AC) and GALLIUM (DE) studies, according to various embodiments.
[0046] In the drawings, similar parts and / or features may have the same reference numerals. Furthermore, parts of the same type may be distinguished by following the reference numeral with a dash and a second reference numeral that distinguishes the similar part. If only the first reference numeral is used in the specification, the description applies to any similar part having the same first reference numeral, regardless of the second reference numeral. DETAILED DESCRIPTION
[0047] I. Overview
[0048] The present disclosure describes techniques for automatic tumor segmentation. More specifically, some embodiments of the present disclosure provide systems and methods for segmenting tumors in positron emission tomography images using deep convolutional neural networks for image and lesion metabolic analysis.
[0049] The use of standardized uptake values (SUVs) is now commonplace in clinical PET / CT oncology imaging and has a specific role in assessing a subject's response to treatment. Oncology imaging using fluorodeoxyglucose (FDG) accounts for the majority of all positron emission tomography (PET) / computed tomography (CT) (PET / CT) imaging procedures, as increased FDG accumulation relative to normal tissue is a useful marker for a variety of cancers. Furthermore, PET / CT imaging is becoming increasingly important as a tool for quantitatively monitoring individual responses to treatment and for evaluating new drug therapies. For example, changes in FDG accumulation have been shown to be useful as an imaging biomarker for assessing treatment response. Several methods exist to measure the rate and / or total amount of FDG accumulation in tumors. PET scanners are designed to measure the in vivo radioactivity concentration [kBq / ml], which is directly related to the FDG concentration. However, the relative tissue uptake of FDG is generally of interest. The two most important sources of variation that occur in practice are the amount of FDG injected and the subject's body size. To compensate for these variations, the SUV is used as a relative measure of FDG uptake. Ideally, the use of SUVs reduces the variability of the signal, which depends on the injected dose of radiotracer and its depletion, and is defined in equation (1).
[0050]
[0051] Where activity is the concentration of radioactivity [kBq / ml] measured by the PET scanner within the region of interest (ROI), dose is the decay-corrected amount of injected radiolabeled FDG [Bq], and weight is the subject's weight [g], which serves as a surrogate for the volume of distribution of the tracer. Using SUV as a measure of relative tissue / organ uptake facilitates comparisons between subjects and has been suggested as a basis for diagnosis.
[0052] However, there are numerous potential sources of bias and variance in determining SUV. One source of bias and variance arises from the analytical methods used to analyze tracer uptake in PET images. Ideally, without loss of resolution or uncertainty in boundary definition, calculating the mean SUV within a ROI on a PET scan would yield a reliable estimate of the mean SUV for the corresponding tissue. Nevertheless, regional and whole-body PET-CT scans are challenging due to their large size and the low proportion of tumor voxels in each image. Regional and whole-body PET-CT also present challenges due to the low resolution of FDG-PET and the low contrast of CT. These effects lead to problems in tumor segmentation, which defines the boundaries of the ROI for which the mean SUV is to be calculated, and ultimately introduce bias and variance in the calculation of the mean SUV.
[0053] Automatic segmentation of tumors and substructures from PET scans has the potential to accurately and reproducibly delineate tumors, which could lead to more efficient and improved tumor diagnosis, surgical planning, and treatment assessment. Most automatic tumor segmentation methods use hand-crafted features. These methods implement a classic machine learning pipeline, where features are first extracted and then fed to a classifier whose training procedure does not affect the properties of these features. Another approach to designing task-adaptive feature representations is to learn a hierarchy of increasingly complex features directly from in-domain data. However, accurate automatic segmentation of tumors from PET scans is challenging for several reasons. First, the boundary between tumor and normal tissue is often unclear due to specific (e.g., brain) and nonspecific (e.g., blood pool) regions of high metabolic activity, the heterogeneity of low-resolution images (e.g., variable density and metabolism of organs), sparse signal (e.g., tumor tissue typically occupies less than approximately 1% of the image), and the sheer number of body structures to be distinguished from the tumor. Second, tumors vary greatly from subject to subject in size, shape, and location. This prevents the use of strong priors on shape and localization that are commonly used for robust image analysis in many other applications, such as facial recognition or navigation.
[0054] To address these limitations and issues, embodiments of the present invention provide a technique for automatic tumor segmentation using a convolutional neural network architecture that is fast and allows the model to handle the size and fuzzy nature of regional and whole-body PET-CT or PET-MRI images. An illustrative embodiment of the present disclosure relates to a method comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of a subject; and preprocessing the PET scans and the CT or MRI scans to generate a first subset of standardized images of a first plane or region of the subject and a second subset of standardized images of a second plane or region of the subject. The first subset of standardized images and the second subset of standardized images incorporate information from the PET scans and the CT or MRI scans. The method may also include generating a first two-dimensional segmentation mask using a first two-dimensional segmentation model implemented as part of a convolutional neural network architecture that takes the first subset of standardized images as input. The first two-dimensional segmentation model uses a first residual block comprising a first layer that: (i) directly feeds into subsequent layers and (ii) directly feeds into layers multiple layers away from the first layer using skip connections. The method may also include generating a second two-dimensional segmentation mask using a second two-dimensional segmentation model implemented as part of the convolutional neural network architecture that takes as input a second subset of the normalized image. The second two-dimensional segmentation model uses a second residual block comprising a second layer that: (i) directly feeds into a subsequent layer, and (ii) directly feeds into a layer multiple layers away from the second layer using a skip connection. The method may also include generating a final imaging mask by combining information from the first two-dimensional segmentation mask and the second two-dimensional segmentation mask.
[0055] Another illustrative embodiment of the present disclosure is directed to a method comprising: obtaining a positron emission tomography (PET) scan and a computed tomography (CT) or magnetic resonance imaging (MRI) scan of a subject; preprocessing the PET scan and the CT or MRI scan to generate standardized images; generating a two-dimensional segmentation mask using a two-dimensional segmentation model implemented as part of a convolutional neural network architecture that takes as input the standardized images; generating a three-dimensional segmentation mask using a three-dimensional segmentation model implemented as part of the convolutional neural network architecture that takes as input image data patches associated with segments from the two-dimensional segmentation mask; and generating a final imaging mask by combining information from the two-dimensional segmentation mask and the three-dimensional segmentation mask. In some cases, the one or more three-dimensional segmentation models include a first three-dimensional segmentation model and a second three-dimensional segmentation model; the image data block includes: a first image data block associated with the first segment; and, a second image data block associated with the second segment; and generating the one or more three-dimensional segmentation masks includes: (i) generating a first three-dimensional segmentation mask using the first three-dimensional segmentation model with the first image data block as input, and (ii) generating a second three-dimensional segmentation mask using the second three-dimensional segmentation model with the second image data block as input.
[0056] Advantageously, these methods include a convolutional neural network architecture that utilizes a two-dimensional segmentation model (modified U-Net) that allows for an optimal mixture of local and general features in the image, and optionally a three-dimensional segmentation model (modified V-Net) that performs volumetric segmentation of the subject in a fast and accurate manner. These methods are also applicable to multi-rate residual learning and multi-scale pyramid learning. In addition, channel filters followed by point-by-point convolution enable deep learning to overcome the limitations of shallow networks. The solution is scalable to whole-body PET-CT or PET-MRI and allows for the assessment of clinical efficacy of treatment, assessment of TMTV in multiple types of cancer, rapid assessment of whole-body FDG-PET tumor burden, prediction of progression-free survival (PFS), staging of treatment subjects, selection of clinical trial subjects and automated end-of-treatment response assessment (e.g., Lugano).
[0057] II. Definitions
[0058] As used herein, when an action is "based on" something, it means that the action is based at least in part on at least a portion of the something.
[0059] As used herein, the terms "substantially," "approximately," and "about" are defined as being largely, but not necessarily entirely, as understood by one of ordinary skill in the art (and including all that is specified). In any disclosed embodiment, the terms "substantially," "approximately," or "about" may be replaced by the specified "within [percentage]," where the percentage includes 0.1%, 1%, 5%, and 10%.
[0060] III. Techniques for Automatic Tumor Segmentation
[0061] Image segmentation is the process of dividing an image into components that display similarities in different features (e.g., shape, size, color, etc.). Tumor segmentation allows visualization of the size and location of tumors within a region of the body (e.g., the brain or lungs) and also provides a basis for analyzing tracer uptake in PET or single-photon emission computed tomography (SPECT) images. For a long time, the gold standard for tumor segmentation has been manual segmentation, which is time-consuming and labor-intensive, making it unsuitable for large-scale studies. Numerous studies have attempted to fully or partially automate the tumor segmentation process. For example, image segmentation techniques such as thresholding, region growing, fuzzy clustering, and the use of watershed algorithms have been used to separate abnormal tissue (e.g., tumor mass) from normal tissue (e.g., white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF) of the brain). Despite this, the segmentation process remains challenging due to the diversity of tumor shapes, locations, and sizes.
[0062] Multimodal imaging techniques that combine information from multiple imaging technologies, such as PET / CT, PET / magnetic resonance imaging (MRI), and SPECT / CT, can help improve the accuracy of tumor segmentation. Combined with PET, the additional imaging modalities can provide information about metabolic or biochemical activity, physiological processes, and other detailed information about the image.
[0063] Described herein is an end-to-end method that incorporates a model that uses two-dimensional and three-dimensional convolutional neural networks (CNNs) to segment tumors and extract metabolic information about lesions from PET (or SPECT) and CT (or MRI) images (scans). As used herein, a "scan" is a graphical representation of a signal on a single plane through the subject's body. The developed model has a low computational load (e.g., it can be run on a common desktop computing device and return predictions on demand, such as, for example, within a few minutes) and is designed to adapt to the size of whole-body scans, the extreme imbalance between tumors and healthy tissues, and the heterogeneity of input images (e.g., variable density and metabolism of organs). The performance of the model in tumor segmentation is comparable to traditional algorithms (such as thresholding, edge-based segmentation, or region-based segmentation) that rely on manual intervention (e.g., manually selecting seeds or manually identifying bounding boxes).
[0064] Tumor metabolic information obtained through PET, SPECT, PET / CT, PET / MRI, and SPECT / CT can be used to assess the clinical efficacy of treatment. Tumor metabolic information can be used to assess TMTV in various types of cancer, with examples including but not limited to subjects with lymphoma (e.g., non-Hodgkin's lymphoma (NHL)) and lung cancer. This method may be a tool for radiologists to quickly assess whole-body PET or SPECT tumor burden. This method can also be used to predict progression-free survival (PFS). This method can also be used to stage subjects for treatment and select subjects for clinical trials. The output of the model can also be used to perform automated end-of-treatment response assessment (e.g., Lugano).
[0065] III.A. Example Computing Environment
[0066] Figure 1 An example computing environment 100 (i.e., a data processing system) for performing tumor segmentation using a deep convolutional neural network is shown, according to various embodiments. The computing environment 100 may include a deep convolutional neural network (CNN) system 105 for training and executing a two-dimensional CNN model, a three-dimensional CNN model, or a combination thereof. More specifically, the CNN system 105 may include classifier subsystems 110a-n that may train their respective CNN models. In some embodiments, each CNN model corresponding to a classifier subsystem 110a-n is individually trained based on one or more PET scans, CT scans, MRI scans, or any combination thereof within a set of input image elements 115a-n. In some cases, the PET scans, CT scans, or MRI scans of the set of input image elements 115a-n may be segmented by an image detector into one or more portions (e.g., body slices: transverse scans, coronal scans, and sagittal scans and / or anatomical portions or regions: head and neck, thorax, and abdomen and pelvis), such that each of the classifier subsystems 110a-n may process a corresponding portion of the PET scan, SPECT scan, CT scan, or MRI scan for training and deployment.
[0067] In some embodiments, each of the input image elements 115a-n may include one or more digital images depicting a portion or region of the body, or the entire body (e.g., a whole-body scan). Each of the input image elements 115a-n may correspond to a single subject at one or more time points (e.g., baseline or before treatment, during treatment, after treatment, etc.), at which underlying image data corresponding to the digital image was obtained. The underlying image data may include one or more PET scans, SPECT scans, CT scans, MRI scans, or any combination thereof. Thus, in some cases, a single input image element 115a-n may include an image corresponding to multiple PET scans, SPECT scans, CT scans, MRI scans, or any combination thereof, each of which depicts a different (e.g., overlapping or non-overlapping) portion or region of the body, or slice of the entire body. In some embodiments, multiple images corresponding to multiple PET scans, SPECT scans, CT scans, MRI scans, or any combination thereof, associated with different portions or regions of the body, or slices of the entire body, are stitched together to form a montage of images to capture several portions or regions of the entire body. Thus, in some cases, a single input image element 115a-n may include a single stitched image.
[0068] The parts or regions of the body represented in a single input image element 115a-n may include one or more central regions of the body (including the head and neck, chest, and abdomen and pelvis) and, optionally, one or more peripheral regions of the body (such as, for example, the legs, feet, arms, or hands). Furthermore, virtual slices, fields, planes, or projections of the body may be represented in a single input image element 115a-n. For example, with PET-CT imaging techniques, sequential images from two devices (a PET scanner and a CT scanner) may be acquired in the same session and combined into a single or multiple superimposed (co-registered) images. Thus, the functional imaging acquired by the PET scanner, which describes the spatial distribution of metabolic or biochemical activity in the body, may be more accurately aligned or correlated with the anatomical imaging acquired by the CT scanner. In some embodiments, a single input image element 115a-n includes both a PET scan and a CT or MRI scan. With respect to the foregoing, it should also be understood that the examples and embodiments described herein with respect to PET scanning and CT scanning are for illustrative purposes only, and that other imaging methods and alternatives thereto will be suggested to those skilled in the art. For example, a PET scan or CT scan with different tracers or angle configurations may capture different structures or regions of the body, and one or more of these types of imaging techniques may be combined with other imaging techniques (such as SPECT and MRI) to gain a deeper understanding of the body's pathology and / or anatomical structures.
[0069] The input image elements 115a-n may include one or more training input image elements 115a-d, validation input image elements 115e-g, and unlabeled input image elements 115h-n. It should be understood that the input image elements corresponding to the training, validation, and unlabeled groups need not be accessed simultaneously. For example, the initial training and validation input image elements may be accessed and used to train the model first, and the unlabeled input image elements may be accessed or received later (e.g., at a single or multiple subsequent times). Furthermore, depending on the specific circumstances of the model training process (such as, for example, when performing k-fold cross-validation), each of the input image elements 115a-g may be accessed and used for training or validation.
[0070] In some cases, supervised training can be used to train the CNN model, and each of the training input image elements 115a-d and the validation input image elements 115e-g can be associated with one or more labels that identify the "correct" interpretation of the presence and / or severity of the tumor. The labels can alternatively or additionally be used to classify the corresponding input image elements, or pixels or voxels therein, according to the presence and / or severity of the tumor at a time point corresponding to the time when the base scan was performed or at a subsequent time point (e.g., after a predetermined duration when the scan was performed). In some cases, unsupervised training can be used to train the CNN model, and each of the training input image elements 115a-d and the validation input image elements 115e-g need not be associated with one or more labels. Each of the unlabeled image elements 115h-n need not be associated with one or more labels.
[0071] The CNN model can be trained using training input image elements 115a-d (and validation input image elements 115e-h to monitor training progress), a loss function, and / or a gradient descent method. In the case where the input image data elements correspond to multiple underlying scans, each scan corresponding to a different part, field, plane, or slice of the body, each of a set of CNN models can be trained to process image data corresponding to a specific part, field, plane, or slice of the body.
[0072] In some embodiments, the classifier subsystems 110a-n include a feature extractor 120, a parameter data store 125, a classifier 130, and a trainer 135, which are collectively used to train a CNN model (e.g., learn parameters of the CNN model during a supervised or unsupervised training process) using training data (e.g., training input image elements 115a-d). In some embodiments, the classifier subsystems 110a-n access the training data from the training input image elements 115a-d at the input layer. The feature extractor 120 can pre-process the training data to extract relevant features (e.g., edges) detected in specific portions of the training input image elements 115a-d. The classifier 130 can receive the extracted features and, based on weights associated with a set of hidden layers in one or more CNN models, convert the features into one or more output metrics that segment one or more tumors and optionally indicate clinical efficacy, assess TMTV, assess whole-body PET tumor burden, predict PFS, stage subjects for treatment and select subjects for clinical trials, automatically perform end-of-treatment response assessments (e.g., Lugano), or a combination thereof. The trainer 135 can use the training data corresponding to the training input image elements 115a-d to train the feature extractor 120 and / or the classifier 130 by facilitating the learning of one or more parameters. For example, the trainer 135 can use backpropagation techniques to facilitate the learning of weights associated with a set of hidden layers of the CNN model used by the classifier 130. Backpropagation can use, for example, a stochastic gradient descent (SGD) algorithm to cumulatively update the parameters of the hidden layers. The learned parameters may include, for example, weights, biases, linear regression, and / or other hidden layer related parameters, which can be stored in the parameter data memory 125.
[0073] An ensemble of trained CNN models (e.g., a plurality of trained CNN models individually trained to identify different features, objects, or metrics in an input image that are then combined into an image mask and / or output metric) can be deployed to process unlabeled input image elements 115h-n to segment one or more tumors and, optionally, predict one or more output metrics indicative of clinical efficacy of treatment, estimate TMTV, estimate whole-body FDG-PET tumor burden, predict PFS, stage subjects for treatment and select subjects for clinical trials, automatically perform end-of-treatment response assessment (e.g., Lugano), or a combination thereof. More specifically, a trained version of the feature extractor 120 can generate feature representations of the unlabeled input image elements, which can then be processed by a trained version of the classifier 130. In some embodiments, image features can be extracted from the unlabeled input image elements 115h-n based on one or more convolutional blocks, convolutional layers, residual blocks, or pyramidal layers that utilize dilation of the CNN models in the classifier subsystems 110a-n. These features can be organized into a feature representation, such as a feature vector for the image. The CNN model can be trained to learn feature types based on classification and subsequent adjustment of parameters in hidden layers (including fully connected layers of the CNN model). In some embodiments, the image features extracted by the convolutional blocks, convolutional layers, residual blocks, or pyramidal layers include feature maps that are matrices of values representing one or more portions of a scan on which one or more image processing operations have been performed (e.g., edge detection, sharpening image resolution). These feature maps can be flattened for processing by the fully connected layers of the CNN model, which output a tumor mask or one or more metrics corresponding to current or future predictions about the tumor.
[0074] For example, the input image elements can be fed to the input layer of the CNN model. The input layer may include nodes corresponding to specific pixels or voxels. The first hidden layer may include a set of hidden nodes, each hidden node connected to multiple input layer nodes. The nodes in the subsequent hidden layers can be similarly configured to receive information corresponding to multiple pixels or voxels. Thus, the hidden layers can be configured to learn to detect features that extend across multiple pixels or voxels. Each of the one or more hidden layers may include a convolutional block, a convolutional layer, a residual block, or a pyramidal layer. The CNN model may also include one or more fully connected layers (e.g., a softmax layer).
[0075] At least a portion of the training input image elements 115a-d, validation input image elements 115e-g, and / or unlabeled input image elements 115h-n may include data collected using one or more imaging systems 160 or data received from the system or may have been derived from that data. The imaging system 160 may include a system configured to collect image data (e.g., a PET scan, a CT scan, an MRI scan, or any combination thereof). The imaging system 160 may include a PET scanner and optionally a CT scanner and / or an MRI scanner. The PET scanner may be configured to detect photons (subatomic particles) emitted by radionuclides in the organ or tissue being examined. The radionuclides used in PET scans may be made by attaching radioactive atoms to chemicals that a particular organ or tissue naturally uses in its metabolic processes. For example, in a PET scan of the brain, a radioactive atom (e.g., radiofluorine, such as 18 F) is attached to glucose (blood sugar) to produce FDG, as the brain uses glucose for metabolism. Other radioactive tracers and / or substances can be used for image scanning, depending on the purpose of the scan. For example, if blood flow and perfusion of an organ or tissue are of interest, the radionuclide can be a radioactive oxygen, carbon, nitrogen, or gallium; if infectious diseases are of interest, a radioactive atom can be attached to sorbitol (e.g., fluorodeoxysorbitol (FDS)); and if oncology is of interest, a radioactive atom can be attached to misonidazole (e.g., flumisonidazole (FMISO)). The raw data collected by a PET scanner is a list of "coincidence events," representing the near-simultaneous detection (usually within a window of 6 to 12 nanoseconds) of annihilation photons by a pair of detectors. Each coincidence event represents a spatial line connecting the two detectors along which positron emission occurs (i.e., a line of response (LOR)). Coincidence events can be grouped into projection images, called sinograms. Sinograms are used for computer analysis to reconstruct two-dimensional and three-dimensional images of metabolic processes in the organ or tissue being examined. Two-dimensional PET images and / or three-dimensional PET images may be included in the set of input image elements 115a-n. However, the sinusoids collected in PET scans are of much lower quality in terms of anatomical structures than CT or MRI scans, which can result in noisier images.
[0076] To overcome the limitations of PET scans in terms of anatomical structures, PET scans are increasingly being read together with CT or MRI scans (called "co-registration") to provide detailed anatomical information and metabolic information (i.e., what the structure is and what it does biochemically). Because PET imaging is most useful in conjunction with separate anatomical imaging (such as CT or MRI), PET scanners can be used in conjunction with integrated high-end multi-detector row CT or MRI scanners. These two scans can be performed immediately after each other during the same session without a particular subject changing position between the two types of scans. This allows for more precise registration of the two sets of images, allowing abnormal areas on the PET imaging to be more perfectly correlated with anatomical structures on the CT or MRI images.
[0077] A CT scanner can be configured to aim a motorized X-ray source, which generates a narrow beam of X-rays, at a specific subject and rapidly rotate the source and beam around the body. A digital X-ray detector, located directly opposite the source, detects the X-rays leaving the body and generates signals that are processed by the scanner's computer to produce two-dimensional cross-sectional images, or "slices," of the body. These slices, also called tomographic images, contain more detailed information than conventional X-rays. The thickness of tissue represented in each image slice can vary depending on the CT scanner used, but typically ranges from 1 to 10 mm. Once a complete slice is taken, the two-dimensional image is stored, and the motorized table holding the subject is gradually moved forward into the gantry. The X-ray scanning process is then repeated multiple times to generate a series of two-dimensional images captured around the axis of rotation. Once the scanner's computer has collected the series of two-dimensional images, they can be digitally "stacked" together through computer analysis to reconstruct a three-dimensional image of the subject. The two-dimensional images and / or reconstructed three-dimensional images make it easier to identify and locate essential structures and possible tumors or abnormalities. When a CT scan is used to image a particular subject as part of a PET-CT scanner or as a standalone CT scanner, both sets of two-dimensional images and / or reconstructed three-dimensional images from the PET and CT scans (or a set of registered two-dimensional images and / or reconstructed three-dimensional images) are included in the set of input image elements 115a-n used to train the CNN model.
[0078] An MRI scanner can be configured to produce detailed, three-dimensional images of anatomy using a strong magnetic field and radio waves. More specifically, the magnetic field forces protons in the tissue or body to align with the magnetic field. When a radiofrequency current is then pulsed through the tissue or body, the protons are stimulated and pulled out of equilibrium, resisting the pull of the magnetic field. When the radiofrequency field is turned off, the MRI sensors are able to detect the energy released as the protons realign with the magnetic field. The time it takes for the protons to realign with the magnetic field, as well as the energy released, varies depending on the environment and the chemical properties of the molecule. A computing device is able to distinguish between various types of tissue based on these magnetic properties and generate a series of two-dimensional images. Once the scanner's computer has collected a series of two-dimensional images, these two-dimensional images can be digitally "stacked" together through computer analysis to reconstruct a three-dimensional image of the subject. The two-dimensional images and / or the reconstructed three-dimensional images allow for easier identification and location of essential structures, as well as possible tumors or abnormalities. When an MRI scan is used to image a particular subject as part of a PET-MRI scanner or as a standalone MRI scanner, the two sets of two-dimensional images and / or reconstructed three-dimensional images from the PET and CT scans (or a set of registered two-dimensional images and / or reconstructed three-dimensional images) are included in the set of input image elements 115a-n used to train the CNN model.
[0079] In some cases, the labels associated with training input image elements 115a-d and / or validation input image elements 115e-g may have already been received or may be derived from data received from one or more provider systems 170, each of which may be associated with, for example, a doctor, nurse, hospital, pharmacist, etc., associated with a particular subject. The received data may include, for example, one or more medical records corresponding to the particular subject. The medical records may indicate, for example, a diagnosis or characterization by a professional indicating whether the subject has a tumor and / or the progression stage of the subject's tumor (e.g., along a standard scale and / or by identifying a metric such as TMTV) for a time period corresponding to the time when the one or more input image elements associated with the subject were collected, or for a subsequently defined time period. The received data may also include pixels or voxels indicating the location of the tumor within the one or more input image elements associated with the subject. Thus, the medical records may include or be used to identify one or more labels for each training / validation input image element. The medical records may also indicate each of one or more therapeutic agents (e.g., medications) that the subject has taken and the time period for which the subject received the treatment. In some cases, the images or scans that are input to one or more classifier subsystems are received from provider system 170. For example, provider system 170 may receive an image or scan from imaging system 160 and may then transmit the image or scan (e.g., along with a subject identifier and one or more labels) to CNN system 105.
[0080] In some embodiments, data received or collected at one or more imaging systems 160 may be aggregated with data received or collected at one or more provider systems 170. For example, the CNN system 105 may identify corresponding or identical identifiers for subjects and / or time periods in order to associate image data received from the imaging system 160 with label data received from the provider system 170. The CNN system 105 may also process the data using metadata or automated image analysis to determine which classifier subsystem to feed a particular data component. For example, the image data received from the imaging system 160 may correspond to the entire body or multiple regions of the body. For each image, metadata, automated alignment, and / or image processing may indicate to which region the image corresponds. For example, automated alignment and / or image processing may include detecting whether the image has image properties corresponding to blood vessels and / or shapes associated with specific organs, such as the lungs or liver. The label-related data received from the provider system 170 may be region-specific or subject-specific. When the label-related data is region-specific, metadata or automated analysis (e.g., using natural language processing or text analysis) may be used to identify to which region the particular label-related data corresponds. When the label-related data is subject-specific, the same label data (for a given subject) can be fed to each classifier subsystem during training.
[0081] In some embodiments, computing environment 100 may also include user device 180, which may be associated with a user requesting and / or coordinating the performance of CNN system 105 to analyze input images associated with subjects for training, testing, validation, or use case purposes. A user may be a physician, an investigator (e.g., associated with a clinical trial), a subject, a medical professional, or the like. Therefore, it should be understood that in some cases, provider system 170 may include and / or act as user device 180. The performance of CNN system 105 to analyze input images may be associated with a specific subject (e.g., a human), who may (but need not) be different from the user. This performance may be achieved by user device 180 transmitting a performance request to the CNN system. The request may include and / or be accompanied by information about the specific subject (e.g., the subject's name or other identifier, such as a de-identified subject identifier). The request may also include identifiers of one or more other systems from which data (e.g., input image data corresponding to the subject) is collected. In some cases, the communication from user device 180 includes an identifier for each subject in a set of specific subjects, corresponding to a request for the performance of CNN system 105 to analyze input images associated with each subject represented in the set of specific subjects.
[0082] Upon receiving the request, the CNN system 105 may send a request for unlabeled input image elements (e.g., the request includes an identifier of the subject) to one or more corresponding imaging systems 160 and / or provider systems 170. The trained CNN ensemble may then process the unlabeled input image elements to segment one or more tumors and generate metrics associated with PFS, such as TMTV (quantitative tumor burden parameter). The results for each identified subject may include or be based on tumor segmentation and / or one or more output metrics from one or more CNN models of the trained CNN ensemble deployed by the classifier subsystems 110a-n. For example, tumor segmentation and / or metrics may include or be based on outputs generated by one or more fully connected layers of the CNN. In some cases, such outputs may be further processed using, for example, a softmax function. Furthermore, aggregation techniques (e.g., random forest aggregation) may then be used to aggregate the outputs and / or further processed outputs to generate one or more subject-specific metrics. One or more results (e.g., including plane-specific outputs and / or one or more subject-specific outputs and / or processed versions thereof) may be transmitted to the user device 180 and / or may be utilized by the user device. In some cases, some or all communications between the CNN system 105 and the user device 180 occur via a website. It should be understood that the CNN system 105 can gate access to results, data, and / or processing resources based on authorization analysis.
[0083] Although not explicitly shown, it should be understood that the computing environment 100 may also include a developer device associated with a developer. Communications from the developer device may indicate what type of input image elements to use for each CNN model in the CNN system 105, the number and type of neural networks to use, and hyperparameters of each neural network, such as the learning rate and the number of hidden layers, as well as how to format data requests and / or what training data to use (e.g., and how to access the training data).
[0084] III.B. Model Overview
[0085] Figure 2A A representative CNN architecture for tumor segmentation according to various embodiments is shown (e.g., Figure 1 In some embodiments, the input image elements 205 are obtained from one or more image sources (e.g., as described with respect to Figure 1The input image elements 205 may be obtained by the imaging system 160 or the provider system 170 described herein. The input image elements 205 may include one or more digital images depicting a portion or region of the body or the entire body (e.g., a whole body scan) of a subject (e.g., a patient) obtained at one or more time points (e.g., baseline or before treatment, during treatment, after treatment, etc.). The underlying image data may include acquired two-dimensional and / or reconstructed three-dimensional PET images, CT images, MRI images, or any combination thereof. Thus, in some cases, a single input image element 205 may include images corresponding to multiple PET scans, SPECT scans, CT scans, MRI scans, or any combination thereof, each of which depicts a different (e.g., overlapping or non-overlapping) portion or region of the body or slice of the entire body. In some embodiments, the multiple images corresponding to the multiple PET scans, SPECT scans, CT scans, MRI scans, or any combination thereof associated with different portions or regions of the body or slices of the entire body are stitched together to form a montage of images so as to capture several portions or regions of the entire body. Thus, in some cases, a single input image element 205 may include a single stitched image.
[0086] The input image elements 205 are structured as one or more arrays or matrices of pixel or voxel values. A given pixel or voxel location is associated with, for example, a general intensity value and / or intensity value as it relates to each of one or more gray levels and / or colors (e.g., RGB values). For example, each image of the input image elements 205 can be structured as a three-dimensional matrix, where the sizes of the first two dimensions correspond to the width and height of each image in pixels. The size of the third dimension can be based on the color channels associated with each image, for example, the third dimension can be 3, corresponding to the three channels of a color image: red, green, and blue.
[0087] The input image elements 205 are provided as input to the pre-processing subsystem 210 of the CNN architecture, which generates normalized image data across the input image elements 205. Pre-processing may include selecting a subset of images or input image elements 205 for a slice (e.g., coronal, axial, and sagittal slices) or region of the body and performing geometric resampling (e.g., interpolation) of the images or subset of input image elements 205 to uniform pixel spacing (e.g., 1.0 mm) and slice thickness (e.g., 2 mm). Image intensity values for all images may be clipped to a specified range (e.g., -1000 to 3000 Hounsfield units) to remove noise and possible artifacts. Normalization of spacing, slice thickness, and units ensures that each pixel has a consistent area and each voxel has a consistent volume across all images of the input image elements 205. The output from the pre-processing subsystem 210 is a subset of normalized images of a slice (e.g., coronal, axial, and sagittal slices) or region of the body.
[0088] The method may include one or more trained CNN models 215 (e.g., Figure 1 Tumor segmentation is performed using a semantic segmentation model architecture based on a CNN model associated with the classifier subsystem 110a described above. In semantic segmentation, the CNN model 215 identifies the location and shape of different objects (e.g., tumor tissue and normal tissue) in the image by classifying each pixel with a desired label. For example, tumor tissue is labeled as tumor and is marked red, normal tissue is labeled as normal and is marked green, and background pixels are labeled as background and are marked black. Subsets of the standardized images from the preprocessing subsystem 210 are used as input to one or more trained CNN models 215. In some cases, a single trained CNN model is used to process all subsets of the standardized images. In other cases, a set of trained CNN models (e.g., a CNN ensemble) is used, where each CNN model is trained to process images originating from different slices (coronal, axial, and sagittal slices) or regions of the body. For example, a subset of normalized images from coronal slices may be used as input to a first CNN of the set of trained CNN models, while a subset of normalized images from sagittal slices may be used as input to a second CNN of the set of trained CNN models.
[0089] The trained CNN model 215 is a two-dimensional segmentation model, such as a U-Net, which is configured to initially obtain a low-dimensional representation of the normalized image and then upsample the low-dimensional representation to generate a two-dimensional segmentation mask 220 for each image. As described in detail herein, the U-Net includes a contraction path supplemented by an expansion path. The contraction path is divided into different stages operating at different resolutions. Each stage includes one to three convolutional layers that generate a low-dimensional representation. The expansion path upsamples the low-dimensional representation to generate a two-dimensional segmentation mask 220. The pooling operations of successive layers in the expansion path are replaced by upsampling operators, and these successive layers increase the resolution of the two-dimensional segmentation mask 220. The two-dimensional segmentation mask 220 is a high-resolution (as used herein, "high-resolution" refers to an image with more pixels or voxels than the low-dimensional representation processed by the contraction path of the U-Net of the V-Net) mask image, in which all pixels are classified (e.g., some pixel intensity values are zero and others are non-zero). The non-zero pixels represent the location of tissue present in an image or image portion (e.g., a PET-CT or PET-MRI scan) from a standardized image subset. For example, wherever a pixel is classified as background, the pixel intensity within the mask image is set to a background value (e.g., zero). Wherever a pixel is classified as tumor tissue, the pixel intensity within the mask image is set to a tumor value (e.g., non-zero). The two-dimensional segmentation mask 220 shows non-zero pixels representing the location of tumor tissue identified in the PET scan relative to anatomical structures in the underlying CT or MRI scan (see, e.g., FIG. 2 ). Figure 2C The non-zero pixels representing the locations of tumor tissue are grouped and labeled as one or more segmentation instances, which indicate various instances of tumor tissue within the image from the input image element 205.
[0090] The two-dimensional segmentation mask 220 is input to a feature extractor 225 (e.g., Figure 1 120 ), and the feature extractor 225 extracts relevant features from the two-dimensional segmentation mask 220. Feature extraction is a dimensionality reduction process by which the initial data set (e.g., the two-dimensional segmentation mask 220) is reduced to more manageable relevant features for further processing. Relevant features may include texture features (e.g., contrast, dissimilarity, cluster shading, cluster prominence, etc.), shape features (e.g., area, eccentricity, extent, etc.), prognostic signatures (e.g., linking tumor features to possible outcomes or disease courses), etc., which may help further classify the pixels within the two-dimensional segmentation mask 220. Texture features may be extracted using a gray level co-occurrence matrix or similar techniques. Shape features may be extracted using a regional attribute function or similar techniques. Prognostic signatures may be extracted using k-means clustering or similar techniques.
[0091] The relevant features extracted from the feature extractor 225 and the 2D segmentation mask 220 are input to the classifier 230 (e.g., Figure 1 The classifier 130 described above is used to classify the two-dimensional segmentation masks 220, and the classifier 230 converts the relevant features and combines the two-dimensional segmentation masks 220 into a final mask image 235. The classifier 230 can use the relevant features to refine the classification of the pixels in the two-dimensional segmentation mask 220 and combine (e.g., using an average or other statistical function) the two-dimensional segmentation masks 220 to generate the final mask image 235. The final mask image 235 is a high-resolution mask image in which all pixels are classified and segmentation contours, boundaries, transparent blocks, etc. overlap around and / or on a designated segment (e.g., tumor tissue) based on the classified pixels (see, e.g., Figure 2C In some cases, the classifier 230 also converts the data obtained from the final mask image 235 into one or more output metrics that indicate clinical efficacy, estimate TMTV, estimate whole-body PET tumor burden, predict PFS, stage subjects for treatment and select subjects for clinical trials, automatically perform end-of-treatment response assessment (e.g., Lugano), or a combination thereof. In one specific example, the prognostic signature can be classified (e.g., using a classifier) according to the design of the classifier to obtain a desired output format. For example, the classifier 235 can be trained on labeled training data to extract TMTV, predict PFS, stage subjects for treatment, and / or select subjects for clinical trials.
[0092] Figure 2B Representative alternative CNN architectures for tumor segmentation according to various embodiments are shown (e.g., Figure 1 An exemplary schematic diagram 250 of a portion of the described CNN system 105 is shown. Figure 2B The CNN architecture shown is similar to Figure 2A 250 , but it is incorporated into the part detection model 255 and the 3D segmentation model 260. Therefore, for the sake of brevity, similar parts, operations and terms will not be repeated as much as possible in the description of the exemplary schematic 250. In some embodiments, the input image element 205 is obtained from one or more image sources (e.g., as described in relation to Figure 1The input image elements 205 are obtained by the imaging system 160 or the provider system 170 described herein. The input image elements 205 include one or more digital images depicting a portion or region of the body or the entire body (e.g., a whole body scan) of a subject (e.g., a patient) obtained at one or more time points (e.g., baseline or before treatment, during treatment, after treatment, etc.). The underlying image data may include acquired two-dimensional and / or reconstructed three-dimensional PET images, CT images, MRI images, or any combination thereof. The input image elements 205 are provided as input to a pre-processing subsystem 210 of the CNN architecture, which generates standardized image data across the input image elements 205. The output from the pre-processing subsystem 210 is a subset of standardized images of slices (e.g., coronal, axial, and sagittal slices) or regions of the body.
[0093] Tumor segmentation can be performed using a semantic segmentation model architecture that includes one or more trained CNN models 215 (e.g., similar to those described in relation to Figure 1 CNN models associated with the classifier subsystem 110a described above), a part detection model 255, and one or more trained CNN models 260 (e.g., Figure 1 The subset of normalized images from the pre-processing subsystem 210 is used as input to one or more trained CNN models 215. The trained CNN model 215 is a two-dimensional segmentation model, such as a U-Net, which is configured to initially obtain a low-dimensional representation of the normalized image and then upsample the low-dimensional representation to generate a two-dimensional segmentation mask 220.
[0094] The standardized image from the pre-processing subsystem 210 (particularly the standardized image used to generate the two-dimensional segmentation mask 220) is used as input to the part detection model 255. The part detection model 255 automatically evaluates the location of the region or body part captured in the standardized image as a reference point and uses the reference point to divide the region or body into multiple anatomical regions. Thereafter, the part detection model can generate position labels for the multiple anatomical regions and incorporate the position labels into the two-dimensional segmentation mask. For example, a segment above and distal to the lung reference point can be labeled as the head and neck, a segment above and proximal to the liver reference point can be labeled as the chest, and a segment proximal to the lung reference point and the liver reference point can be labeled as the abdomen-pelvis.
[0095] Image data blocks corresponding to segments within the two-dimensional segmentation mask 220 are used as input to one or more trained CNN models 260. Each segment is a pixel-wise or voxel-wise mask of objects classified in the underlying image. Each segment's image data block includes a large amount of image data with fixed-size voxels represented as a (width)×b (height)×c (depth), which is derived from both the pixel size of the block and the slice thickness. The block can be defined by the boundary of the segment (i.e., the pixels or voxels that are classified as part of the segment, such as tumor tissue), a bounding box with generated coordinates containing the segment, or a buffer of the boundary or bounding box plus a predetermined number of pixels or voxels to ensure that the entire segment is contained in the image data block. The system process or CNN model 260 uses the labels provided for multiple anatomical regions as tags to select segments within the two-dimensional segmentation mask 220 for input to the selection CNN model 260 (e.g., a CNN model trained specifically for scans from a specific anatomical region, such as the head and neck region). For example, a subset of blocks corresponding to a subset of segments identified by head and neck position labels can be used as input to a first CNN trained on head and neck scans, a subset of blocks corresponding to a subset of segments identified by chest position labels can be used as input to a second CNN trained on chest scans, and a subset of blocks corresponding to a subset of segments identified by abdomen-pelvis position labels can be used as input to a third CNN trained on abdomen-pelvis scans.
[0096] The trained CNN model 260 is a three-dimensional segmentation model, such as a V-Net, which is configured to initially obtain a low-resolution feature map for each image data block corresponding to a segment, and then upsample the low-resolution feature map to generate a three-dimensional segmentation mask 265 for each image data block. As described in detail herein, the V-Net includes a contraction path supplemented by an expansion path. The contraction path is divided into different stages that operate at different resolutions. Each stage includes one or more convolutional layers. At each stage, a residual function is learned. The input of each stage is used in the convolutional layer, processed nonlinearly, and added to the output of the last convolutional layer of the stage so that the residual function can be learned. This architecture ensures convergence compared to non-residual learning networks such as U-Net. The convolution performed at each stage uses a volume kernel of size n×n×n voxels. A voxel (volume element or volume pixel) represents a value, sample, or data point on a regular grid in three-dimensional space.
[0097] The expansion path extracts features and expands the spatial support of the lower resolution feature maps in order to collect and combine the necessary information to output a three-dimensional segmentation mask 265 for each image data block. At each stage, a deconvolution operation is employed to increase the size of the input, followed by one to three convolutional layers involving half the number of n×n×n kernels employed in the previous layer. A residual function is learned, similar to the contraction part of the network. Each three-dimensional segmentation mask 265 is a high-resolution mask image in which all voxels are classified (e.g., some of the voxel intensity values are zero and others are non-zero). Non-zero voxels represent the location of tissue present in an image or image portion (e.g., PET-CT or PET-MRI scan) from a standardized image subset. For example, wherever a voxel is classified as background, the intensity of the voxel within the mask image will be set to the background value (e.g., zero). Wherever it is classified as tumor tissue, the intensity of the voxel within the mask image will be set to the tumor value (e.g., non-zero). The three-dimensional segmentation mask 265 shows the non-zero voxels representing the location of tumor tissue identified in the PET scan relative to the anatomical structure in the underlying CT or MRI scan. The non-zero voxels representing the location of tumor tissue are grouped and labeled into one or more segments that indicate various instances of tumor tissue within the image data block.
[0098] The 2D segmentation mask 220 and the 3D segmentation mask 265 for each image data block are input to the feature extractor 225 (e.g., as described with respect to Figure 1 The feature extractor 120 described above is input to the feature extractor 120, and the feature extractor 225 extracts relevant features from the two-dimensional segmentation mask 220 and each three-dimensional segmentation mask 265. The extracted relevant features, the two-dimensional segmentation mask 220 and the three-dimensional segmentation mask 265 are input to the classifier 230 (e.g., as described with respect to Figure 1 The classifier 230 converts and combines the relevant features, the two-dimensional segmentation mask 220, and the three-dimensional segmentation mask 265 into a final mask image 235. More specifically, the classifier 230 uses the relevant features to refine the classification of pixels and voxels in each of the two-dimensional segmentation mask 220 and the three-dimensional segmentation mask 265.
[0099] Thereafter, the classifier combines (e.g., averages pixel and / or voxel values) the refined two-dimensional segmentation mask 220 and the three-dimensional segmentation mask 265 to generate a final mask image 235. The final mask image 235 is a high-resolution mask image in which all pixels and voxels are classified and segmentation contours, boundaries, transparent blocks, etc. are overlapped around and / or on segments (e.g., tumor tissue) based on the classified pixels and voxels (see, e.g., Figure 2CIn some cases, the classifier 230 also converts the data obtained from the final mask image 235 into one or more output metrics that indicate clinical efficacy, estimate TMTV, estimate whole-body PET tumor burden, predict PFS, stage subjects for treatment and select subjects for clinical trials, automatically perform end-of-treatment response assessment (e.g., Lugano), or a combination thereof. In one specific example, the prognostic signature can be classified (e.g., using a classifier) according to the design of the classifier to obtain a desired output format. For example, the classifier 235 can be trained on labeled training data to extract TMTV, predict PFS, stage subjects for treatment, and / or select subjects for clinical trials.
[0100] Figure 2C According to various embodiments, Figure 2B Another view of the model architecture of , which includes input image elements 205 (normalized by the pre-processing subsystem 210), a 2D segmentation model 215, a part detection model 255, a 2D segmentation mask 220 with region labels, a 3D segmentation model 265, and a predicted final mask image 235. Figure 2C In the model architecture of , models 215 / 255 / 265 are shown as performing three operations: (i) 2D segmentation 270 generates a first prediction of the tumor within the 2D tumor mask using the 2D segmentation model 215 - as further explained in III.C of this document, (ii) location and region detection 275 identifies anatomical regions within the input image elements 205 (normalized by the preprocessing subsystem 210) and provides location labels for the corresponding anatomical regions within the 2D segmentation mask 220 obtained in step (i) - as further explained in III.D of this document, and (iii) refines the first prediction (from step (i)) of the tumor (identified in step (ii)) in each anatomical region of the 2D segmentation mask 220 using the corresponding 3D segmentation model 265, which performs 3D segmentation 280 on each anatomical region separately to generate multiple 3D tumor masks, which can be combined into a final mask image 235 - as further explained in III.E of this document.
[0101] III.C. Example U-Net for 2D Segmentation
[0102] 2D segmentation uses a modified U-Net to extract features from the input image (e.g., a standardized PET scan, CT scan, MRI scan, or any combination thereof) to generate a 2D segmentation mask with high-resolution features. Figure 3As shown, U-Net 300 includes a contraction path 305 and an expansion path 310, which gives it a U-shaped architecture. The contraction path 305 is a CNN network that includes repeated applications of convolutions (e.g., 3x3 convolutions (unpadded convolutions)), each followed by a rectified linear unit (ReLU) and a maximum pooling operation (e.g., 2x2 maximum pooling with a stride of 2) for downsampling. The input to the convolution operation is a three-dimensional volume (i.e., an input image of size nxnx channels, where n is the number of input features) and a set of "k" filters (also called kernels or feature extractors), each of size (fxfx channels, where f is an arbitrary number, such as 3 or 5). The output of the convolution operation is also a three-dimensional volume (also called an output image or feature map) of size (mxmxk, where M is the number of output features and k is the convolution kernel size).
[0103] Each block 315 of the contraction path 315 includes one or more convolutional layers (represented by gray horizontal arrows), and the number of feature channels varies, for example, from 1 to 64 (e.g., depending on the starting number of channels in the first pass), as the convolution process increases the depth of the input image. The downward gray arrows between each block 315 represent the max pooling process, which halves the size of the input image. With each downsampling step or pooling operation, the number of feature channels may double. During the contraction process, the spatial information of the image data decreases while the feature information increases. Thus, the information that existed in, for example, a 572x572 image before pooling now exists in, for example, a 284x284 image after pooling. Now, when the convolution operation is applied again in a subsequent pass or layer, the filters in that subsequent pass or layer will be able to see a larger context. That is, as the input image progresses deeper into the network, the size of the input image decreases, while the receptive field increases (the receptive field (context) is the area of the input image covered by the kernel or filter at any given point in time). Once block 315 is executed, two more convolutions are performed, but without max pooling, in block 320. The image after block 320 has been resized to, for example, 28x28x1024 (this size is illustrative only, and the size at the end of process 320 may vary depending on the starting size of the input image - a size of nxnx channels).
[0104] The expansion path 310 is a CNN network that combines the features from the contraction path 305 and spatial information (the upsampling of the feature map from the contraction path 305). As described herein, the output of the two-dimensional segmentation is not just a class label or bounding box parameters. Instead, the output (the two-dimensional segmentation mask) is a complete high-resolution image in which all pixels are classified. If a conventional convolutional network with pooling layers and dense layers is used, the CNN network will lose the "where" information and only retain the "what" information which is unacceptable for image segmentation. In the case of image segmentation, both "what" and "where" information are needed. Therefore, it is necessary to upsample the image, that is, convert the low-resolution image to a high-resolution image to recover the "where" information. The transposed convolution represented by the upward white arrow is an exemplary upsampling technique that can be used in the expansion path 310 to upsample the feature map and expand the size of the image.
[0105] After the transposed convolution at block 325, the image is enlarged from 28x28x1024 to 56x56x512, and then the image is concatenated with the corresponding image from the shrinking path (see the horizontal gray bar 330 from the shrinking path 305), together forming an image of size 56x56x1024. The reason for the concatenation is to combine information from previous layers (i.e., combining the high-resolution features from the shrinking path 305 with the upsampled output from the expanding path 310) to obtain more accurate predictions. The process continues as a series of upconvolutions (upsampling operators) that halve the number of channels, concatenation with the corresponding cropped feature maps from the shrinking path 305, repeated application of convolutions followed by rectified linear units (ReLUs) (e.g., two 3x3 convolutions), and the final convolution in block 335 (e.g., one 1x1 convolution) to generate a multi-channel segmentation as a two-dimensional segmentation mask. For localization, U-Net 300 uses the active part of each convolution without any fully connected layers, i.e., the segmentation map only contains pixels for which the full context in the input image is available and uses skip connections that link the contextual features learned in the contraction block with the localization features learned in the expansion block.
[0106] In the case of performing only two-dimensional segmentation, the two-dimensional segmentation mask output from the U-Net is used as the input of the feature extractor, and the feature extractor extracts relevant features from the two-dimensional segmentation mask. The relevant features and the two-dimensional segmentation mask are input into the classifier, and the classifier converts the relevant features and the two-dimensional segmentation mask into a final mask image. The classifier can use the relevant features to refine the classification of pixels in the two-dimensional segmentation mask and generate a final mask image. The final mask image is a high-resolution mask image in which all pixels are classified and segmentation contours, boundaries, transparent blocks, etc. are covered around and / or on the specified segment (e.g., tumor tissue) based on the classified pixels.
[0107] In a traditional U-Net architecture, the blocks of the contraction and expansion paths are simply composed of convolutional layers (e.g., typically two or three layers) for performing convolutions. However, according to various embodiments, a block (e.g., block 315) is a residual block that includes one or more layers that: (i) feed directly into subsequent layers, and (ii) use skip connections to feed directly into layers that are multiple layers away from the one or more layers, which propagate larger gradients to the one or more layers during backpropagation. In some cases, the one or more layers of the residual block are pyramidal layers with separable convolutions performed at one or more dilation levels.
[0108] Figure 4A Shown Figure 3 315 . As shown, the residual block 400 may include multiple convolutional layers 405 (where, in the illustrated embodiment, each individual convolutional layer is replaced by two or more pyramidal layers 320). In a network (e.g., ResNet) that includes the residual block 400, one or more layers 405 (where, in the illustrated embodiment, each layer is a pyramidal layer) feed directly into the next layer (A, B, C, etc.) and directly into further layers, such as, for example, multiple layers (D, E, etc.). Using the residual block 400 in the network helps overcome the degradation problem that occurs with increasing the number of convolutional or pyramidal layers (if the number of layers continues to increase, the accuracy will initially increase, but will begin to saturate and eventually degrade at some point). The residual block 400 uses skip connections, or residual connections, to skip some of these additional layers, which ultimately propagates larger gradients to the initial layers during backpropagation. Skipping effectively simplifies the network by using fewer layers during the initial training phase. This speeds up learning by reducing the effects of vanishing gradients because fewer layers need to be propagated through (i.e., multi-speed residual learning). The network then gradually recovers the skipped layers as it learns the feature space. Empirical evidence shows that residual blocks improve accuracy and are easier to optimize.
[0109] Figure 4B According to various embodiments, Figure 4AA single pyramidal layer 405 is shown. Pyramid layer 405 can use dilated (dark black) separable convolutions ("dilation blocks") at multiple different scales, four levels in this example. Pyramid layer 405 includes the same image at multiple scales to improve the accuracy of detecting objects (e.g., tumors). The separable convolutions in pyramidal layer 405 are separable depthwise convolutions, i.e., spatial convolutions performed independently on each channel of the input, followed by pointwise convolutions, such as 1x1 convolutions, that project the channels output by the depthwise convolutions into a new channel space. Compared to using typical two-dimensional convolutions, separable depthwise convolutions provide significant gains in convergence speed and significantly reduce model size. Dilated (dark black) convolutions dilate the kernel by inserting spaces between kernel elements. This additional parameter, l (dilation rate), which controls the spacing between kernel elements, indicates how much the kernel will be widened. This creates a filter with an "expanded" receptive field, which increases the size of the receptive field relative to the kernel size. In some embodiments, the dilation rate is four levels of dilation. In other embodiments, larger or smaller dilation levels can be used, for example, six levels of dilation. Dilated (dark black) convolutions expand the receptive field without increasing the kernel size and without losing resolution, which is especially effective when multiple dilated convolutions are stacked one after another as in the pyramidal layer 405.
[0110] The convolutional layer outputs 415 are the outputs of the dilation blocks 420 (labeled as dilations 1, 2, 4, and 8). Figure 4B The example shown assumes that there are four dilation blocks, and each dilation block outputs two channels (of the same color), so the total number of output channels is 8. The number of channels output by each dilation block may vary depending on the residual block in question. Figure 4B The example shows Figure 3 In some embodiments, the number of each channel output by each dilation block 415 in the pyramidal layer 410 of the residual block 405 is equal to the number of k filters on the residual block 405 divided by four.
[0111] III.D. Exemplary Site and Region Detection and Labeling
[0112] Due to the different densities and metabolism of organs, PET scans, CT scans, and MRI scans can be highly heterogeneous, depending on the location of the scan in the body. To limit this variability in downstream processing, a part detection model is configured to separate the region or body depicted in the scan into multiple anatomical regions, such as three regions including the head and neck, chest, and abdomen-pelvis. The part detection model automatically evaluates the positions of the regions or parts of the body (e.g., liver and lungs) captured in the two-dimensional segmentation mask as reference points, and uses the reference points to separate the region or body into multiple anatomical regions. Thereafter, the part detection model can generate position labels for the multiple anatomical regions and incorporate the position labels into the two-dimensional segmentation mask. Downstream processing can use the labels as markers to select segments within the two-dimensional segmentation mask for processing by a selected CNN model (e.g., a CNN model trained specifically for scans from the head and neck region). This splitting and labeling reduces the imaging space and allows the selected CNN model to be trained on different anatomical regions, thereby improving the overall learning of the CNN model.
[0113] The liver can be detected using methods such as the following. The part detection model can be performed by thresholding the scan with a threshold (e.g., a threshold of 2.0 SUV) and looking for areas larger than a predetermined area or volume (e.g., 500 mm) after filling holes (e.g., background holes in the image) using morphological closing and opening of structuring elements with a predetermined radius (e.g., 8 mm). 3) to determine if the brain is present in a scan (e.g., a PET scan). Morphological closing is a mathematical operator that consists of dilation followed by erosion, both operations using the same structuring element. Morphological opening is a mathematical operator that consists of erosion followed by dilation, both operations using the same structuring element. The dilation operator takes two pieces of data as input. The first is the image to be dilated. The second is a (usually small) set of coordinates called a structuring element (also called a kernel). It is this structuring element that determines the precise effect of the dilation on the input image. The dilation operator essentially acts by gradually enlarging the boundaries of regions of foreground pixels (e.g., typically white pixels). As a result, the size of the foreground pixel regions increases, while the holes within these regions decrease. The erosion operator takes two pieces of data as input. The first is the image to be eroded. The second is a (usually small) set of coordinates called a structuring element (also called a kernel). It is this structuring element that determines the precise effect of the erosion on the input image. The erosion operator essentially acts by eroding away the boundaries of regions of foreground pixels (e.g., typically white pixels). As a result, the size of the foreground pixel regions decreases, and the holes within these regions become larger. The part detection model examines the lower right portion of the image and fills the holes (using, for example, a closing or opening operation) using a predetermined threshold (e.g., a predetermined threshold of 1.0 SUV), erodes the connected parts using an erosion operator, and examines the topmost connected part with a center of mass in the posterior third of the sagittal axis. The center of mass of this connected part is located within the liver. In other embodiments, alternative values and / or alternative methods for the above terms may be used.
[0114] The centroid of the lungs can be detected using a method such as the following. The site detection model can threshold the image at a predefined scale (e.g., -300 Hounsfield units (HU) for CT scans) to obtain a binary mask and maintain a certain number (e.g., eight) of the largest connected sites identifiable within the image. In each slice (e.g., sagittal, axial, coronal, etc.), the site detection model can remove selected areas adjacent to the slice boundaries, erode the remaining connected sites to avoid any leakage and retain only the two largest connected sites. The model uses the centroid of the remaining two largest connected sites as the centroid of the lungs (inferring that the remaining two largest sites are the lungs). In other embodiments, alternative values and / or alternative methods for the above terms can be used.
[0115] Alternatively, the part detection model can use organ segmentation to estimate the location of a part, such as a region or organ in the body depicted in a PET scan, CT scan, or MRI scan, to obtain one or more reference points for the region or organ in the body. In some cases, organ segmentation can be used additionally or alternatively for organ-specific measurements of one or more organs (such as the spleen, liver, lungs, and kidneys). One or more organs can be segmented using methods such as the following. The part detection model (e.g., a three-dimensional convolutional neural network, such as, for example, a V-Net for three-dimensional organ segmentation) can include downsampling and upsampling subnetworks with skip connections to propagate higher resolution information to the final segmentation. In some cases, the downsampling subnetwork can be a series of multiple dense feature stacks connected by downsampling convolutions, each skip connection can be a single convolution of the corresponding dense feature stack output, and the upsampling network includes bilinear upsampling of the final segmentation resolution. The output of the part detection model will be an organ segmentation mask for the input scan.
[0116] Once the regions or body parts (e.g., liver and lungs) are detected, these parts can be used as reference points within a two-dimensional segmentation mask to divide the region or body into multiple anatomical regions. The part detection model can generate location labels for the multiple anatomical regions and include the location labels within the two-dimensional segmentation mask. As a result, the two-dimensional segmentation mask can include labels for multiple anatomical regions. Downstream processing can use the labels as markers to select segments within the two-dimensional segmentation mask for processing by a selected CNN model (e.g., a CNN model trained specifically for scans from the head and neck region).
[0117] III.D. Exemplary V-Net for 3D Segmentation
[0118] The three-dimensional segmentation based on a volumetric CNN system of multiple different sub-models extracts features separately from the image data blocks of each anatomical segment. The image data blocks correspond to the segments within the two-dimensional segmentation mask. Each block includes a large amount of image data with fixed-size voxels derived from the pixel size of the block and the slice thickness. The block can be defined by the boundary of the segment (i.e., the pixels classified as part of the segment, such as tumor tissue), a bounding box with generated coordinates containing the segment, or a buffer of the boundary or bounding box plus a predetermined number of pixels or voxels to ensure that the entire segment is contained within the image data block. The system processing or CNN system of multiple different sub-models uses the labels provided for multiple anatomical regions as tags to select the segments within the two-dimensional segmentation mask for input into a selected CNN model (e.g., a CNN model trained specifically for scans from a specified anatomical region, such as the head and neck region).
[0119] like Figure 5As shown, for each anatomical region marked in the two-dimensional segmentation mask, a separate V-Net 500a-n can be used to refine the image data within the block associated with each anatomical region. Refining the image data within the block includes classifying the volumetric data of the voxels in each block (i.e., two-dimensional segmentation classifies pixels in two-dimensional space, while three-dimensional segmentation adds classification to the classification for voxels in three-dimensional space). For example, each image data block can be viewed as a matrix of pixel and voxel values, where each pixel and voxel region of the matrix can be assigned a value. In some cases, the image data block includes black and white characteristics with pixel or voxel values ranging from 0 to 1 and / or color characteristics with three assigned RGB pixel or voxel values ranging from 0 to 255. To classify the voxels, each V-Net 500a-n will perform a series of operations on the image data block, including: (1) convolution; (2) nonlinear transformation (e.g., ReLU); (3) pooling and or subsampling; (4) classification (fully connected layer), as described in detail below.
[0120] Each V-Net 500 includes a compression path 505 for downsampling and a decompression path 510 for upsampling, decompressing the signal until it reaches its original size. The compression path 510 is divided into different blocks 515 operating at different resolutions. Each block 515 may include one or more convolutional layers. Appropriate padding may be used to apply convolution within each layer. Each block 515 may be configured to learn a residual function via a residual connection: the input of each block 515 is (i) used in a convolutional layer and processed through a nonlinearity, and (ii) added to the output of the last convolutional layer of the block to enable learning of the residual function. The convolution performed in each block 515 uses a volumetric kernel of a predetermined size (such as 5×5×5 voxels). As the image data travels through different blocks 515 along the compression path 510, the resolution of the image data decreases. This is performed by convolution while applying a kernel of a predetermined size (such as a 2×2×2 voxel-wide kernel) with an appropriate stride (e.g., stride 2). Because the second operation extracts features by considering only non-overlapping volumetric blocks, the size of the resulting feature map is halved (subsampled). This strategy is similar to the purpose of the pooling layer. Replacing the pooling operation with a convolution operation results in a smaller network memory footprint because backpropagation does not require the switch that maps the output of the pooling layer back to its input. Each stage of the compression path 505 computes a multiple of the number of features from the previous layer or block.
[0121] The decompression path 510 is divided into different blocks 520 that are used to extract features and expand the spatial support of the lower-resolution feature maps in order to collect and aggregate the necessary information to output a multi-channel volume segmentation as a 3D segmentation mask. After each block 520 of the decompression path 515, a deconvolution operation can be applied to increase the size of the input, followed by one or more convolutional layers involving half the number of kernels used in the previous layer, such as a 5×5×5 kernel. Similar to the compression path 510, a residual function can be learned in the convolution stage of the decompression path 515. In addition, features extracted from the early stages of the compression path 510 can be forwarded to the decompression path 515, as shown by the horizontal connection 525. The two feature maps computed by the last convolutional layer have an appropriate kernel size, such as a 1×1×1 kernel size, and produce an output of the same size as the input volume (both volumes have the same resolution as the original input image data block), which can be processed by a soft-max layer, which outputs the probability of each voxel belonging to a class (such as foreground or background).
[0122] In image data such as PET scans, CT scans, and MRI scans, it is not uncommon for anatomical structures of interest (e.g., tumors) to occupy only a small region of the scan. This often causes the learning process to get stuck in a local minimum of the loss function, resulting in a network whose predictions are strongly biased towards the background. For example, the average proportion of negative voxels in the volume is 99.5%, while in a single slice it is always above 80%. As a result, foreground regions are often lost or only partially detected. To address the unbalanced nature of the images, an objective function based on the Dice Similarity Coefficient (DSC) and a two-dimensional weighted cross entropy can be used in the soft-max layer, as shown in Equation (2).
[0123]
[0124] Among them, V is the voxel space, T is the set of positive voxels, P refers to the set of predicted positive voxels, y v is the value of voxel v in the 3D segmentation mask, and y_hat v is the value of voxel v in the predicted 3D segmentation mask.
[0125] In three dimensions, DSC can be used together with sensitivity and mean absolute error in the loss function, as shown in Equation (3).
[0126]
[0127] Among them, V is the voxel space, T is the set of positive voxels, P refers to the set of predicted positive voxels, y v is the value of voxel v in the 3D segmentation mask, and y_hat v is the value of voxel v in the predicted 3D segmentation mask.
[0128] It should be understood that although Figure 5 The use of three V-Nets 500 (each with two paths) to refine an image data patch from a two-dimensional segmentation mask is depicted, but different numbers of V-Nets and convolutional layers can be used (e.g., this may have the effect of repeating these operations one or more times for a CNN system). For example, an output can be determined by applying five or more convolutional layers to extract features from the image data patch to determine a current or future prediction related to tumor segmentation.
[0129] When performing two-dimensional segmentation and three-dimensional segmentation, the two-dimensional segmentation mask and the three-dimensional segmentation mask are used as input to a feature extractor, and the feature extractor extracts relevant features from the two-dimensional segmentation mask and the three-dimensional segmentation mask. The relevant features extracted from the feature extractor, the two-dimensional segmentation mask, and the three-dimensional segmentation mask are input to a classifier, and the classifier converts the extracted relevant features, the two-dimensional segmentation mask, and the three-dimensional segmentation mask into a final mask image. For example, the classifier uses the relevant features to refine and classify the pixels in each of the two-dimensional segmentation mask and the three-dimensional segmentation mask. Thereafter, the classifier combines the refined two-dimensional segmentation mask and the three-dimensional segmentation mask to generate a final mask image. For example, the final mask image can be obtained by averaging the refined two-dimensional segmentation mask and the three-dimensional segmentation mask (or applying one or more other statistical operations). The final mask image is a high-resolution mask image in which all pixels and voxels are classified, and segmentation contours, boundaries, transparent blocks, etc. are overlaid around and / or on a designated segment (e.g., tumor tissue) based on the classified pixels and voxels.
[0130] IV. Techniques for Extraction and Prediction
[0131] Figure 6 A process 600 for determining extracted TMTV is shown in accordance with various embodiments.
[0132] Process 600 begins at block 605, where a plurality of PET (or SPECT) scans (e.g., FDG-PET scans) of a subject and a plurality of CT or MRI scans of the subject are accessed. The PET scan and the corresponding CT or MRI scan may depict at least a portion of the subject's body or the subject's entire body. For example, the PET scan and the corresponding CT or MRI scan may depict one or more organs of the body, including the lungs, liver, brain, heart, or any combination thereof. Optionally, at block 610, the PET scan and the corresponding CT or MRI scan may be preprocessed to generate standardized images or a subset of scans of slices (e.g., coronal, axial, and sagittal slices) or regions of the body.
[0133] At box 615, a CNN architecture is used to convert the PET scan and the corresponding CT or MRI scan into an output (e.g., a final mask image). In some embodiments, the CNN architecture includes one or more two-dimensional segmentation models, such as a modified U-Net, as described in detail herein. The two-dimensional segmentation model is configured to generate a two-dimensional segmentation mask from the PET scan and the corresponding CT or MRI scan or a subset of the standardized images or scans. The two-dimensional segmentation model uses multiple residual blocks, each having a separable convolution and multiple dilations, on the PET scan and the corresponding CT or MRI scan or a subset of the standardized images or scans to generate a two-dimensional segmentation mask. The two-dimensional segmentation mask can be refined using a feature extractor and a classifier, and then combined (e.g., using a mean or other statistical function) to generate a final mask image, wherein all pixels are classified and, based on the classified pixels, segmentation contours, boundaries, transparent blocks, etc. are overlaid around and / or on a designated segment (e.g., tumor tissue).
[0134] In other embodiments, the CNN architecture includes one or more two-dimensional segmentation models and multiple three-dimensional segmentation models, as described in detail herein. The two-dimensional segmentation model is configured to generate a two-dimensional segmentation mask from a PET scan and a corresponding CT or MRI scan or a standardized image or a subset of the scan. The three-dimensional segmentation model is configured to generate a three-dimensional segmentation mask from a PET scan and a corresponding CT or MRI scan or a standardized image or a subset of the scan. The three-dimensional segmentation model can use residual connections and multiple losses to generate a three-dimensional segmentation mask for the PET scan and the corresponding CT or MRI scan or a subset of the standardized image or scan. The two-dimensional segmentation mask and the three-dimensional segmentation mask can be refined using a feature extractor and a classifier, and then combined (e.g., using an average or other statistical function) to generate a final mask image, wherein all pixels and voxels are classified, and based on the classified pixels and voxels, segmentation contours, boundaries, transparent blocks, etc. are covered around and / or on a specified segment (e.g., tumor tissue).
[0135] Where the CNN architecture uses one or more two-dimensional segmentation models and multiple three-dimensional segmentation models, the CNN architecture can divide an area or body depicted in a PET scan and a CT or MRI scan, or a subset of standardized images or scans, into multiple anatomical regions, such as three regions including the head and neck, chest, and abdomen-pelvis. The CNN architecture automatically evaluates the positions of the parts of the area or body captured in the PET scan and the CT or MRI scan, or a subset of standardized images or scans, as reference points, and uses the reference points to divide the area or body into multiple anatomical regions. Thereafter, the CNN architecture can generate position labels for the multiple anatomical regions and incorporate the position labels into a two-dimensional segmentation mask. Image patches associated with segments within the two-dimensional segmentation mask can be divided into each of the multiple anatomical regions and images processed by different three-dimensional segmentation models (which can share an architecture but have different learning parameters).
[0136] At block 620, a TMTV can be extracted from the final mask image. In some cases, a CNN architecture segments the tumor in each anatomical region based on features extracted from PET scans and CT or MRI scans (including SUV values from the PET scans). A metabolic tumor volume (MTV) can be determined for each segmented tumor. The TMTV for a given subject can be determined from all segmented tumors and represents the sum of all individual MTVs.
[0137] At block 625, the TMTV is output. For example, the TMTV can be presented locally or transmitted to another device. The TMTV can be output along with an identifier of the subject. In some cases, the TMTV is output along with a final mask image and / or other information that identifies image regions, features, and / or detections that facilitate extraction of the TMTV. Thereafter, a diagnosis can be provided and / or treatment can be administered to the subject, or treatment can be administered to the subject based on the extracted TMTV.
[0138] Figure 7 Illustrated is a process 700 for predicting a subject's likelihood of progression-free survival (PFS), clinical efficacy, a subject's disease stage, and / or a selection decision for inclusion of a subject in a clinical trial, according to various embodiments.
[0139] Process 700 begins at block 705, where a plurality of PET (or SPECT) scans (e.g., FDG-PET scans) of a subject and a plurality of CT or MRI scans of the subject are accessed. The PET scan and the corresponding CT or MRI scan may depict at least a portion of the subject's body or the subject's entire body. For example, the PET scan and the corresponding CT or MRI scan may depict one or more organs of the body, including the lungs, liver, brain, heart, or any combination thereof. Optionally, at block 710, the PET scan and the corresponding CT or MRI scan may be preprocessed to generate standardized images or a subset of scans of slices (e.g., coronal, axial, and sagittal slices) or regions of the body.
[0140] At box 715, a CNN architecture is used to convert the PET scan and the corresponding CT or MRI scan into an output (e.g., a final mask image). In some embodiments, the CNN architecture includes one or more two-dimensional segmentation models, such as a modified U-Net, as described in detail herein. The two-dimensional segmentation model is configured to generate a two-dimensional segmentation mask from the PET scan and the corresponding CT or MRI scan or a subset of the standardized images or scans. The two-dimensional segmentation model uses multiple residual blocks, each having a separable convolution and multiple dilations, on the PET scan and the corresponding CT or MRI scan or a subset of the standardized images or scans to generate a two-dimensional segmentation mask. The two-dimensional segmentation mask can be refined using a feature extractor and a classifier, and then combined (e.g., using a mean or other statistical function) to generate a final mask image, wherein all pixels are classified and, based on the classified pixels, segmentation contours, boundaries, transparent blocks, etc. are overlaid around and / or on a specified segment (e.g., tumor tissue).
[0141] In other embodiments, the CNN architecture includes one or more two-dimensional segmentation models and multiple three-dimensional segmentation models, as described in detail herein. The two-dimensional segmentation model is configured to generate a two-dimensional segmentation mask from a PET scan and a corresponding CT or MRI scan or a standardized image or a subset of the scan. The three-dimensional segmentation model is configured to generate a three-dimensional segmentation mask from a PET scan and a corresponding CT or MRI scan or a standardized image or a subset of the scan. The three-dimensional segmentation model can use residual connections and multiple losses to generate a three-dimensional segmentation mask for the PET scan and the corresponding CT or MRI scan or a subset of the standardized image or scan. The two-dimensional segmentation mask and the three-dimensional segmentation mask can be refined using a feature extractor and a classifier, and then combined (e.g., using an average or other statistical function) to generate a final mask image, wherein all pixels and voxels are classified, and based on the classified pixels and voxels, segmentation contours, boundaries, transparent blocks, etc. are covered around and / or on a specified segment (e.g., tumor tissue).
[0142] Where the CNN architecture uses one or more two-dimensional segmentation models and multiple three-dimensional segmentation models, the CNN architecture can divide an area or body depicted in a PET scan and a CT or MRI scan, or a subset of standardized images or scans, into multiple anatomical regions, such as three regions including the head and neck, chest, and abdomen-pelvis. The CNN architecture automatically evaluates the positions of the parts of the area or body captured in the PET scan and the CT or MRI scan, or a subset of standardized images or scans, as reference points, and uses the reference points to divide the area or body into multiple anatomical regions. Thereafter, the CNN architecture can generate position labels for the multiple anatomical regions and incorporate the position labels into a two-dimensional segmentation mask. Image patches associated with segments within the two-dimensional segmentation mask can be divided into each of the multiple anatomical regions and images processed by different three-dimensional segmentation models (which can share an architecture but have different learning parameters).
[0143] Optionally, at block 720, a separate 3D CNN is used to convert the PET scan and the corresponding CT or MRI scan into outputs associated with organ segmentation. The 3D CNN may include a 3D segmentation model configured to generate 3D organ masks from the PET scan and the corresponding CT scan. The 3D segmentation model may use downsampling and upsampling subnetworks, utilizing skip connections to propagate higher resolution information to generate 3D organ masks.
[0144] At block 725, the TMTV can be extracted from the final mask image. In some cases, the CNN architecture segments the tumor in each anatomical region based on features extracted from PET scans and CT or MRI scans (including SUV values from the PET scans). A metabolic tumor volume (MTV) can be determined for each segmented tumor. The TMTV for a given subject can be determined from all segmented tumors and represents the sum of all individual MTVs. Optionally, organ-specific measurements such as MTV and the number of lesions per organ (e.g., number of lesions>1 ml) can be extracted from the final mask image and the three-dimensional organ mask. Organ involvement can be defined as an automatic organ MTV>0.1 mL for noise reduction purposes.
[0145] At block 730, one or more of the extracted TMTV, the extracted MTV (e.g., for each organ) and the number of lesions for each organ are input into a classifier to generate a clinical prediction for the subject. In some cases, a clinical prediction metric is obtained as the output of the classifier, which uses at least a portion of the final mask output and / or the extracted TMTV as input. In other cases, a clinical prediction metric is obtained as the output of the classifier, which uses at least a portion of the three-dimensional organ mask output, the extracted MTV and / or the number of lesions as input. The clinical prediction metric may correspond to a clinical prediction. In some cases, the clinical prediction is the likelihood of progression-free survival (PFS) of the subject, the subject's disease stage, and / or the selection decision of the subject for inclusion in a clinical trial. Kaplan-Meier analysis can be used to assess PFS, and the Cox proportional hazard model can be used to estimate the prognostic value of organ-specific involvement.
[0146] At block 735, the clinical prediction is output. For example, the clinical prediction can be presented locally or transmitted to another device. The clinical prediction can be output along with an identifier for the subject. In some cases, the clinical prediction is output along with the TMTV, TMV, number of lesions, final mask output, three-dimensional organ mask output, and / or other information identifying image regions, features, and / or detections that contribute to the clinical prediction. Thereafter, a diagnosis can be provided and / or treatment can be administered to the subject, or treatment can be administered to the subject based on the clinical prediction.
[0147] Figure 8 A process 800 is shown for providing automated end-of-treatment response assessment, which may include the Lugano classification for lymphoma staging, in accordance with various embodiments.
[0148] Process 800 begins at block 805, where a plurality of PET (or SPECT) scans (e.g., FDG-PET scans) of a subject and a plurality of CT or MRI scans of the subject are accessed. The PET scans and the corresponding CT or MRI scans may depict at least a portion of the subject's body or the subject's entire body. For example, the PET scans and the corresponding CT or MRI scans may depict one or more organs of the body, including the lungs, liver, brain, heart, or any combination thereof. Optionally, at block 810, the PET scans and the corresponding CT or MRI scans may be preprocessed to generate standardized images or slices (e.g., coronal, axial, and sagittal slices) of the body or a subset of scans of a region.
[0149] At box 815, a CNN architecture is used to convert the PET scan and the corresponding CT or MRI scan into an output (e.g., a final mask image). In some embodiments, the CNN architecture includes one or more two-dimensional segmentation models, such as a modified U-Net, as described in detail herein. The two-dimensional segmentation model is configured to generate a two-dimensional segmentation mask from the PET scan and the corresponding CT or MRI scan or a subset of the standardized images or scans. The two-dimensional segmentation model uses multiple residual blocks, each having a separable convolution and multiple dilations, on the PET scan and the corresponding CT or MRI scan or a subset of the standardized images or scans to generate a two-dimensional segmentation mask. The two-dimensional segmentation mask can be refined using a feature extractor and a classifier and then combined (e.g., using a mean or other statistical function) to generate a final mask image in which all pixels are classified and, based on the classified pixels, segmentation contours, boundaries, transparent blocks, etc. are overlaid around and / or on a designated segment (e.g., tumor tissue).
[0150] In other embodiments, the CNN architecture includes one or more two-dimensional segmentation models and multiple three-dimensional segmentation models, as described in detail herein. The two-dimensional segmentation model is configured to generate a two-dimensional segmentation mask from a PET scan and a corresponding CT or MRI scan or a standardized image or a subset of the scan. The three-dimensional segmentation model is configured to generate a three-dimensional segmentation mask from a PET scan and a corresponding CT or MRI scan or a standardized image or a subset of the scan. The three-dimensional segmentation model can use residual connections and multiple losses to generate a three-dimensional segmentation mask for the PET scan and the corresponding CT or MRI scan or a subset of the standardized image or scan. The two-dimensional segmentation mask and the three-dimensional segmentation mask can be refined using a feature extractor and a classifier, and then combined (e.g., using an average or other statistical function) to generate a final mask image, wherein all pixels and voxels are classified, and based on the classified pixels and voxels, segmentation contours, boundaries, transparent blocks, etc. are covered around and / or on a specified segment (e.g., tumor tissue).
[0151] Where the CNN architecture uses one or more two-dimensional segmentation models and multiple three-dimensional segmentation models, the CNN architecture can divide an area or body depicted in a PET scan and a CT or MRI scan, or a subset of standardized images or scans, into multiple anatomical regions, such as three regions including the head and neck, chest, and abdomen-pelvis. The CNN architecture automatically evaluates the positions of the parts of the area or body captured in the PET scan and the CT or MRI scan, or a subset of standardized images or scans, as reference points, and uses the reference points to divide the area or body into multiple anatomical regions. Thereafter, the CNN architecture can generate position labels for the multiple anatomical regions and incorporate the position labels into a two-dimensional segmentation mask. Image patches associated with segments within the two-dimensional segmentation mask can be divided into each of the multiple anatomical regions and images processed by different three-dimensional segmentation models (which can share an architecture but have different learning parameters).
[0152] At block 820, a TMTV can be extracted from the final mask image. In some cases, a CNN architecture segments the tumor in each anatomical region based on features extracted from PET scans and CT or MRI scans (including SUV values from the PET scans). A metabolic tumor volume (MTV) can be determined for each segmented tumor. The TMTV for a given subject can be determined from all segmented tumors and represents the sum of all individual MTVs.
[0153] At block 825, the extracted TMTV is input into a classifier to generate an automatic end-of-treatment response assessment based on the TMTV. The automatic end-of-treatment response assessment may correspond to a predicted current or future occurrence, rate of progression, or magnitude of a tumor. The automatic end-of-treatment response assessment may include, for example, a progression score along a specific scale (e.g., tumor grade). In some cases, the automatic end-of-treatment response assessment and / or output includes the difference between the progression score at a predefined time point and a baseline time point. In some cases, the automatic end-of-treatment response assessment and / or output includes a binary indicator, such as a binary value representing a prediction of whether the subject's tumor will progress by at least a predefined amount within a predefined time period. In some cases, the automatic end-of-treatment response assessment includes a prediction of the Lugano classification for lymphoma staging.
[0154] At block 830, the automatic end-of-treatment response assessment is output. For example, the automatic end-of-treatment response assessment can be presented locally or transmitted to another device. The automatic end-of-treatment response assessment can be output along with the subject's identifier. In some cases, the automatic end-of-treatment response assessment is output along with the TMTV, the final mask output, and / or other information that identifies image regions, features, and / or detections that facilitate the automatic end-of-treatment response assessment. Thereafter, a diagnosis can be provided and / or treatment can be administered to the subject, or treatment can be administered to the subject based on the automatic end-of-treatment response assessment.
[0155] V. Examples
[0156] The systems and methods implemented in the various embodiments may be better understood with reference to the following examples.
[0157] VA Example 1.-Fully Automated Measurement of Total Metabolic Tumor Burden in Diffuse Large B-Cell Lymphoma and Follicular Lymphoma
[0158] Baseline TMTV from FDG PET / CT scans has been shown to predict progression-free survival in lymphomas such as diffuse large B-cell lymphoma (DLBCL) and follicular lymphoma (FL).
[0159] VB Data
[0160] The dataset includes a total of 3,506 whole-body (including thoracoabdominal and pelvic) FDG-PET / CT scans collected from multiple sites, representing over 1.5 million images of subjects with lymphoma and lung cancer. The dataset contains scans for 1,595 subjects with non-Hodgkin's lymphoma, 1,133 subjects with DLBCL, and 562 subjects with FL with complete ground truth, as well as 158 subjects with non-small cell lung cancer (NSCLC) with partial ground truth. The data is stored in DICOM format.
[0161] Data were obtained from two Phase 3 clinical trials in NHL subjects (Goya, N=1401, NCT01287741; and Gallium, N=595, NCT01332968). FDG-PET images and semi-automatically defined 3D tumor contours were used. After preprocessing, a total of 870 (Goya only) baseline scans and associated segmentation masks were used for algorithm training. In addition, a separate set of 400 randomly selected datasets (250 Goya and 150 Gallium) was provided for testing purposes.
[0162] VC preprocessing
[0163] Preprocessing consisted of aligning the PET and CT scans, resampling the scans to obtain an isotropic voxel size of 2x2x2 mm, and deriving the SUV for PET using information from the DICOM header. The segmentation mask was reconstructed from the RTStruct file and used as the ground truth for training the CNN architecture.
[0164] 1133 DLBCL subjects were used as the training dataset. This included a total of 861,053 coronal slices, 770,406 sagittal slices, and 971,265 axial slices, as well as 13,942 individual tumors. The test set included a total of 1066 scans from FL subjects and 316 scans from NSCLC subjects.
[0165] There are two reasons for adopting this split and retaining such a large amount of data for testing. One issue is being able to verify that the model can be extended to other types of cancer. Therefore, all subjects with follicular lymphoma were retained in the test set. In addition, the lung cancer subjects in the dataset only had up to five lesions segmented. Therefore, these scans were excluded from the training set to avoid training on data with false negatives, and sensitivity was used to verify the algorithm's performance on these scans.
[0166] VD Training Program
[0167] Experiments were conducted to determine the optimal set of hyperparameters. The learning rate was varied (coarse-fine tuning) and a variable learning rate (cosine annealing) was tested for each network. For the 2D CNN, experiments included testing two kernel sizes, 3x3 and 5x5; a kernel size of 5x5 did not result in improved performance and slowed the model down. Experiments were also conducted to determine the optimal depth of the U-Net. Increasing the depth from 6 to 7 did not improve performance metrics. Predictions for axial slices were removed as they resulted in a large number of false positives with high activity (e.g., kidney, heart, bladder).
[0168] A two-dimensional network associated with processing images or scans from the coronal plane and a two-dimensional network associated with processing images or scans from the sagittal plane were trained on two Nvidia Quadro P6000 graphics processing units (GPUs) using the RMSProp optimizer (160,000 iterations with a batch size of 8). The learning rate was set to 1e-5 for 80,000 iterations and then divided by 2 every 20,000 iterations. More than 80% of the slices did not contain any tumor. To avoid convergence to null predictions, the dataset was rebalanced to achieve a percentage of healthy slices of approximately 10% (98,000 training slices per view).
[0169] V-Net was trained using an optimizer (e.g., the optimizer disclosed in Bauer C, Sun S, Sun W et al. Automated measurement of uptake in cerebellum, liver, and aortic arch in full-body FDGPET / CT scans. Med Phys. 2012; 39(6): 3112-23) with a learning rate of 1e-4 for 200,000 iterations, and the learning rate was set to 1e-4 to 100,000 iterations, 1e-4 = 2 for 50,000 iterations, and 1e-4 = 4 for 50,000 iterations.
[0170] Tumors were manually segmented and peer-reviewed by board-certified radiologists. Compared to radiologist-derived tumor segmentations, the model reported an average voxel-wise sensitivity of 92.7% on a test set of 1,470 scans and an average 3D DICE score of 88.6% on 1,064 scans.
[0171] VE segmentation results
[0172] To perform 3D segmentation, the model uses image data patches associated with segments identified in the 2D segmentation mask obtained from the 2D segmentation model discussed in this article. Both FDG-PET and CT or MRI are used as input to the CNN architecture. Connected sites in the 2D segmentation mask are labeled according to their relative position to the liver and chest references. For each of these anatomical regions, a separate V-Net is used to refine the 2D segmentation. In one example embodiment, the network contains 4 downsampling blocks and 3 upsampling blocks, and the layers use ReLu activation and a 3x3x3 kernel size. In this example embodiment, the patches are 32x32x32x2 in the head or neck, 64x64x64x2 in the chest, and 96x96x96x2 in the abdomen.
[0173] The segmentation results are listed in Table 1 and 9A to 9F In. 9A to 9F In the example embodiment, for each sub-graph, the left side is the ground truth and the right side is the prediction. Both CT and SUV are used as input to exploit the structural and metabolic information provided by each modality. 9A to 9F In the example shown in , the input size is 448x512x2. The number of convolutions in the first layer is 8, and is multiplied by two along the downsampling block. A separate network is used for each coronal and sagittal plane. For these example lung cancer scans, only up to 5 lesions were segmented, so the sensitivity of these scans is reported. The method is more accurate than traditional algorithms that rely on manual intervention (DSC of 0.732 and 0.85, respectively). The method is applicable to whole-body FDG-PET / CT scans, and the model trained on DLBCL subjects can be transferred to FL subjects and NSCLC subjects.
[0174] Table 1. Segmentation results
[0175]
[0176]
[0177] VF total metabolic tumor volume and SUV max Comparison
[0178] Figure 10A and Figure 10B The tumor volume and SUV of the prediction mask according to this example embodiment are shown respectively.max Tumor volume and SUVmax have been shown to have prognostic value. Specifically, K-means clustering based on the algorithm's predicted TMTV values identified a signature with slower to faster PFS. The K-means method identified four distinct clusters based on the predicted baseline TMTV. Figure 11A and Figure 11B Kaplan-Meier estimates of these clusters are shown, showing the correlation of TMTV with prognosis, according to example embodiments. At the subject level, these clusters are more discriminatory than clustering using simple TMTV quartiles or maximum SUV or total lesion glycolysis. Figure 12A and Figure 12B Automated TMTV was shown to provide prognostic metrics at baseline that were consistent with manual TMTV assessment. The ability to automatically and accurately quantify these prognostic metrics enables rapid integration with other clinical markers and may facilitate clinical trial stratification as well as save time and costs. Figure 13A and Figure 13B Baseline TMTV was shown to be prognostic in NSCLC and melanoma.
[0179] VG Example 2.-Automated Assessment of the Independent Prognostic Value of Extranodal Involvement in Diffuse Large B-Cell Lymphoma and Follicular Lymphoma
[0180] The presence of extranodal disease detected by FDG-PET / CT in DLBCL and FL is associated with poor outcome. Accurate and reproducible quantitative image interpretation tools are needed for tumor detection and assessment of metabolic activity in lymphomas using FDG-PET / CT. Based on various aspects discussed herein, a model architecture is provided for fully automated tumor and organ segmentation in PET / CT images based on organ-specific (liver, spleen, and kidney) metabolic tumor burden and prognosis of subjects with DLBCL and FL.
[0181] VH data
[0182] The dataset includes a total of 1,139 pre-treatment PET / CT scans from the GOYA study in DLBCL (NCT01287741) and 541 pre-treatment scans from the GALLIUM study in FL (NCT01332968). The data are stored in DICOM format.
[0183] VI method
[0184] An image processing pipeline consisting of two-dimensional and three-dimensional cascaded convolutional neural networks was trained on the GOYA dataset and tested on the GALLIUM dataset for tumor segmentation. The three-dimensional cascaded convolutional neural network was also trained on publicly available liver, spleen, and kidney segmentation datasets (validation DSC = 0.94, 0.95, and 0.91, respectively). Segmentation allowed the extraction of total metabolic tumor volume (TMTV) and organ-specific measurements (metabolic tumor volume [MTV] and number of lesions >1 mL) for the spleen, liver, and kidney. Organ involvement was defined as an automatic organ MTV > 0.1 mL for noise reduction purposes. Kaplan-Meier analysis was used to assess progression-free survival (PFS), while the Cox proportional hazards model was used to assess the prognostic value of organ-specific involvement.
[0185] VJ Results
[0186] Automated analysis of pretreatment PET / CT scans from the GOYA study showed that the presence of ≥2 lesions >1 mL in the liver and / or spleen was associated with inferior PFS in univariate analysis (hazard ratio (HR) = 1.73; 95% confidence interval (CI) = 1.29–2.32; p = 0.0002). This association was maintained in multivariate analysis after adjusting for TMTV > median in the liver / spleen and ≥2 extranodal lesions (HR = 1.52; 95% CI = 1.10–2.07; p = 0.009) and after adjusting for the International Prognostic Index (IPI), cell of origin (COO), and imaging-derived features (≥2 extranodal sites: HR = 1.49; 95% CI = 1.02–2.18; p = 0.037). Kaplan-Meier analysis also showed that extranodal involvement (≥2 extranodal lesions in the liver and / or spleen) was significantly associated with worse PFS in GOYA ( Figure 14A Both liver and kidney involvement were prognostic in DLBCL on univariate analysis (HR = 1.48; 95% CI = 1.13–1.94; p = 0.004 and HR = 1.44; 95% CI = 1.08–1.91; p = 0.013, respectively). Figure 14B and Figure 14C ); however, splenic involvement was not prognostic. Multivariate analysis also confirmed the prognostic value of liver and kidney involvement for PFS when adjusting for imaging-derived factors (HR = 1.40; 95% CI = 1.06–1.85; p = 0.017 and HR = 1.34; 95% CI = 1.00–1.80; p = 0.049, respectively).
[0187] In FL subjects from the GALLIUM study, ≥2 lesions >1 mL in the liver and / or spleen were associated with PFS by univariate analysis (HR=1.61; 95% CI=1.09–2.38; p=0.017) and by Kaplan-Meier analysis ( Figure 14D Liver and spleen involvement also had prognostic significance by univariate analysis (HR = 1.64; 95% CI = 1.12-2.38; p = 0.010 and HR = 1.67; 95% CI = 1.16-2.40; p = 0.006, respectively). Figure 14E and Figure 14F ), but renal involvement was not prognostic for FL. When adjusting for imaging-derived features, multivariate analysis showed that splenic involvement remained prognostic for PFS (HR = 1.51; 95% CI = 1.03–2.21; p = 0.034); however, liver involvement was no longer significantly associated (HR = 1.44; 95% CI = 0.97–2.14; p = 0.068). After adjusting for the Follicular Lymphoma International Prognostic Index (FLIPI), liver involvement remained prognostic (HR = 1.52; 95% CI = 1.03–2.23; p = 0.036).
[0188] In subjects with DLBCL from the GOYA cohort, automated analysis of PET / CT demonstrated that the presence of ≥2 extranodal lesions in the liver and / or spleen was an independent prognostic factor and added prognostic value to TMTV > median, IPI, and COO. Splenic involvement alone was not prognostic in DLBCL. In subjects with FL, extranodal involvement (≥2 lesions in the liver and / or spleen) and the presence of splenic lesions were independent prognostic factors.
[0189] VI. Other Notes
[0190] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes disclosed herein.
[0191] The terms and expressions used are used as terms of description and not of limitation, and in the use of such terms and expressions, there is no intention to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications can be made within the scope of the invention claimed. Therefore, it should be understood that although the invention claimed has been specifically disclosed by way of embodiments and optional features, modifications and variations of the concepts disclosed herein may be made by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined in the appended claims.
[0192] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the preferred exemplary embodiments will provide those skilled in the art with a feasible description for implementing various embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.
[0193] Specific details are provided in the following description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.
Claims
1. A method comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of the subject; preprocessing the PET scan and the CT or MRI scan to generate a first subset of standardized images for a first plane or region of the subject and a second subset of standardized images for a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; generating a first two-dimensional segmentation mask using a first two-dimensional segmentation model implemented as part of a convolutional neural network architecture and having as input a first subset of the normalized image, wherein the first two-dimensional segmentation model is trained to process the image for the first plane or region using a first residual block comprising a first layer that: (i) feeds directly into subsequent layers and (ii) feeds directly into a layer a plurality of layers away from the first layer using a skip connection; generating a second two-dimensional segmentation mask using a second two-dimensional segmentation model implemented as part of the convolutional neural network architecture and having as input a second subset of the normalized image, wherein the second two-dimensional segmentation model is trained to process the image for the second plane or region using a second residual block comprising a second layer that: (i) feeds directly into subsequent layers and (ii) feeds directly into a layer a plurality of layers away from the second layer using a skip connection; as well as A final imaging mask is generated by combining information from the first two-dimensional segmentation mask and the second two-dimensional segmentation mask.
2. The method of claim 1, wherein the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
3. The method according to claim 1 or 2, further comprising: determining a total metabolic tumor burden TMTV using the final imaging mask; and providing the TMTV.
4. The method according to claim 1 or 2, further comprising: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, the three-dimensional organ segmentation model taking the PET scan and the CT or MRI scan as input; Determining a metabolic tumor burden (MTV) and a number of lesions for one or more organs in the three-dimensional organ segmentation using the final imaging mask and the three-dimensional organ mask; and The MTV and the number of lesions of the one or more organs are provided.
5. The method according to claim 3, further comprising: using a classifier that takes as input one or more of total metabolic tumor burden TMTV, metabolic tumor burden MTV, and the number of lesions, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: The probability of progression-free survival (PFS) of the subject; the subject's disease stage; and The selection decision to include the subject in the clinical trial.
6. The method according to any one of claims 1 to 2, further comprising: inputting a plurality of PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; as well as One or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions are received on a display of a computing device.
7. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform operations comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of the subject; preprocessing the PET scan and the CT or MRI scan to generate a first subset of standardized images for a first plane or region of the subject and a second subset of standardized images for a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; generating a first two-dimensional segmentation mask using a first two-dimensional segmentation model implemented as part of a convolutional neural network architecture and having as input a first subset of the normalized image, wherein the first two-dimensional segmentation model is trained to process the image for the first plane or region using a first residual block comprising a first layer that: (i) feeds directly into subsequent layers and (ii) feeds directly into a layer a plurality of layers away from the first layer using a skip connection; generating a second two-dimensional segmentation mask using a second two-dimensional segmentation model implemented as part of the convolutional neural network architecture and having as input a second subset of the normalized image, wherein the second two-dimensional segmentation model is trained to process the image for the second plane or region using a second residual block comprising a second layer that: (i) feeds directly into subsequent layers and (ii) feeds directly into a layer a plurality of layers away from the second layer using a skip connection; as well as A final imaging mask is generated by combining information from the first two-dimensional segmentation mask and the second two-dimensional segmentation mask.
8. The computer program product of claim 7, wherein the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
9. The computer program product of claim 7 or 8, wherein the operations further comprise: determining a total metabolic tumor burden TMTV using the final imaging mask; and providing the TMTV.
10. The computer program product of claim 7 or 8, wherein the operations further comprise: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, the three-dimensional organ segmentation model taking the PET scan and the CT or MRI scan as input; Determining a metabolic tumor burden (MTV) and a number of lesions for one or more organs in the three-dimensional organ segmentation using the final imaging mask and the three-dimensional organ mask; and The MTV and the number of lesions of the one or more organs are provided.
11. The computer program product of claim 9, wherein the operations further comprise: using a classifier that takes as input one or more of total metabolic tumor burden TMTV, metabolic tumor burden MTV, and the number of lesions, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: The probability of progression-free survival (PFS) of the subject; the subject's disease stage; and The selection decision to include the subject in the clinical trial.
12. The computer program product of any one of claims 7 to 8, wherein the operations further comprise: inputting a plurality of PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; as well as One or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions are received on a display of a computing device.
13. The computer program product of any one of claims 7 to 8, wherein the operations further comprise administering a treatment to the subject based on one or more of the final imaging mask, total metabolic tumor burden (TMTV), metabolic tumor burden (MTV), and number of lesions.
14. The computer program product of any one of claims 7 to 8, wherein the operations further comprise providing a diagnosis for the subject based on one or more of the final imaging mask, total metabolic tumor burden (TMTV), metabolic tumor burden (MTV), and number of lesions.
15. A system comprising: one or more data processors; as well as A non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of the subject; preprocessing the PET scan and the CT or MRI scan to generate a first subset of standardized images for a first plane or region of the subject and a second subset of standardized images for a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; generating a first two-dimensional segmentation mask using a first two-dimensional segmentation model implemented as part of a convolutional neural network architecture and having as input a first subset of the normalized image, wherein the first two-dimensional segmentation model is trained to process the image for the first plane or region using a first residual block comprising a first layer that: (i) feeds directly into subsequent layers and (ii) feeds directly into a layer a plurality of layers away from the first layer using a skip connection; generating a second two-dimensional segmentation mask using a second two-dimensional segmentation model implemented as part of the convolutional neural network architecture and having as input a second subset of the normalized image, wherein the second two-dimensional segmentation model is trained to process the image for the second plane or region using a second residual block comprising a second layer that: (i) feeds directly into subsequent layers and (ii) feeds directly into a layer a plurality of layers away from the second layer using a skip connection; as well as A final imaging mask is generated by combining information from the first two-dimensional segmentation mask and the second two-dimensional segmentation mask.
16. The system of claim 15, wherein the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
17. The system of claim 15 or 16, wherein the operations further comprise: determining a total metabolic tumor burden TMTV using the final imaging mask; and providing the TMTV.
18. The system of claim 15 or 16, wherein the operations further comprise: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, the three-dimensional organ segmentation model taking the PET scan and the CT or MRI scan as input; Determining a metabolic tumor burden (MTV) and a number of lesions for one or more organs in the three-dimensional organ segmentation using the final imaging mask and the three-dimensional organ mask; and The MTV and the number of lesions of the one or more organs are provided.
19. The system of claim 17, wherein the operations further comprise: using a classifier that takes as input one or more of total metabolic tumor burden TMTV, metabolic tumor burden MTV, and the number of lesions, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: The probability of progression-free survival (PFS) of the subject; the subject's disease stage; and The selection decision to include the subject in the clinical trial.
20. The system according to any one of claims 15 to 16, wherein The operations further include: inputting a plurality of PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; and One or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions are received on a display of a computing device.
21. The system of any one of claims 15 to 16, wherein the operations further comprise administering a treatment to the subject based on one or more of the final imaging mask, total metabolic tumor burden (TMTV), metabolic tumor burden (MTV), and number of lesions.
22. The system of any one of claims 15 to 16, wherein the operations further comprise providing a diagnosis for the subject based on one or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions.
23. A method comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of the subject; preprocessing the PET scan and the CT or MRI scan to generate a standardized image that incorporates information from the PET scan and the CT or MRI scan; generating one or more two-dimensional segmentation masks using one or more two-dimensional segmentation models implemented as part of a convolutional neural network architecture and having the normalized image as input; generating one or more three-dimensional segmentation masks using one or more three-dimensional segmentation models implemented as part of the convolutional neural network architecture and taking as input patches of image data associated with segments from the two-dimensional segmentation mask; as well as generating a final imaging mask by combining information from the one or more two-dimensional segmentation masks and the one or more three-dimensional segmentation masks, wherein the standardized images include a first subset of standardized images for a first plane or region of the subject and a second subset of standardized images for a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; The one or more two-dimensional segmentation models include a first two-dimensional segmentation model and a second two-dimensional segmentation model; and Generating the one or more two-bit segmentation masks includes: generating a first two-dimensional segmentation mask using the first two-dimensional segmentation model implemented to take as input the first subset of the normalized images, wherein the first two-dimensional segmentation model is trained to process images for the first plane or region; as well as A second two-dimensional segmentation mask is generated using the second two-dimensional segmentation model having the second subset of the normalized images as input, wherein the second two-dimensional segmentation model is trained to process images for the second plane or region.
24. The method of claim 23, wherein: The one or more three-dimensional segmentation models include a first three-dimensional segmentation model and a second three-dimensional segmentation model; The image data blocks include: a first image data block associated with a first section; and a second image data block associated with a second section; and Generating the one or more 3D segmentation masks includes: generating a first 3D segmentation mask using the first 3D segmentation model with the first image data block as input, and generating a second 3D segmentation mask using the second 3D segmentation model with the second image data block as input.
25. The method of claim 24, further comprising: evaluating the position of the region or part of the body captured in the standardized image as a reference point; dividing the region or body into a plurality of anatomical regions based on the reference points; generating location labels for the plurality of anatomical regions; Incorporating the position labels into the two-dimensional segmentation mask; determining, based on the location tag, that the first segment is located in a first anatomical region of the plurality of anatomical regions; determining, based on the location tag, that the second segment is located in a second anatomical region of the plurality of anatomical regions; inputting the first image data block associated with the first segment into the first three-dimensional segmentation mask based on a determination that the first segment is located in the first anatomical region; as well as Based on a determination that the second segment is located in the second anatomical region, the second image data block associated with the second segment is input into the second three-dimensional segmentation mask.
26. The method of claim 23, 24 or 25, wherein: The first two-dimensional segmentation model uses a first residual block comprising a first layer that: (i) feeds directly into subsequent layers, and (ii) feeds directly into a layer that is a plurality of layers away from the first layer using a skip connection; and The second 2D segmentation model uses a second residual block comprising a second layer that: (i) feeds directly into subsequent layers, and (ii) feeds directly into a layer a plurality of layers away from the second layer using a skip connection.
27. The method of claim 26, wherein the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
28. The method according to any one of claims 23 to 25, further comprising: determining a total metabolic tumor burden TMTV using the final imaging mask; and providing the TMTV.
29. The method according to any one of claims 23 to 25, further comprising: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, the three-dimensional organ segmentation model taking the PET scan and the CT or MRI scan as input; Determining a metabolic tumor burden (MTV) and a number of lesions for one or more organs in the three-dimensional organ segmentation using the final imaging mask and the three-dimensional organ mask; and The MTV and the number of lesions of the one or more organs are provided.
30. The method of claim 28, further comprising: using a classifier that takes as input one or more of total metabolic tumor burden TMTV, metabolic tumor burden MTV, and the number of lesions, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: The probability of progression-free survival (PFS) of the subject; the subject's disease stage; and The selection decision to include the subject in the clinical trial.
31. The method according to any one of claims 23 to 25, further comprising: inputting a plurality of PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; as well as One or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions are received on a display of a computing device.
32. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform operations comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of the subject; preprocessing the PET scan and the CT or MRI scan to generate a standardized image that incorporates information from the PET scan and the CT or MRI scan; generating one or more two-dimensional segmentation masks using one or more two-dimensional segmentation models implemented as part of a convolutional neural network architecture and having the normalized image as input; generating one or more three-dimensional segmentation masks using one or more three-dimensional segmentation models implemented as part of the convolutional neural network architecture and taking as input patches of image data associated with segments from the two-dimensional segmentation mask; as well as generating a final imaging mask by combining information from the one or more two-dimensional segmentation masks and the one or more three-dimensional segmentation masks, wherein the standardized images include a first subset of standardized images for a first plane or region of the subject and a second subset of standardized images for a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; The one or more two-dimensional segmentation models include a first two-dimensional segmentation model and a second two-dimensional segmentation model; and Generating the one or more two-bit segmentation masks includes: generating a first two-dimensional segmentation mask using the first two-dimensional segmentation model implemented to take as input the first subset of the normalized images, wherein the first two-dimensional segmentation model is trained to process images for the first plane or region; as well as A second two-dimensional segmentation mask is generated using the second two-dimensional segmentation model having the second subset of the normalized images as input, wherein the second two-dimensional segmentation model is trained to process images for the second plane or region.
33. The computer program product of claim 32, wherein: The one or more three-dimensional segmentation models include a first three-dimensional segmentation model and a second three-dimensional segmentation model; The image data blocks include: a first image data block associated with a first section; and a second image data block associated with a second section; and Generating the one or more 3D segmentation masks includes: generating a first 3D segmentation mask using the first 3D segmentation model with the first image data block as input, and generating a second 3D segmentation mask using the second 3D segmentation model with the second image data block as input.
34. The computer program product of claim 33, wherein the operations further comprise: evaluating the position of the region or part of the body captured in the standardized image as a reference point; dividing the region or body into a plurality of anatomical regions based on the reference points; generating location labels for the plurality of anatomical regions; Incorporating the position labels into the two-dimensional segmentation mask; determining, based on the location tag, that the first segment is located in a first anatomical region of the plurality of anatomical regions; determining, based on the location tag, that the second segment is located in a second anatomical region of the plurality of anatomical regions; inputting the first image data block associated with the first segment into the first three-dimensional segmentation mask based on a determination that the first segment is located in the first anatomical region; as well as Based on a determination that the second segment is located in the second anatomical region, the second image data block associated with the second segment is input into the second three-dimensional segmentation mask.
35. A computer program product according to claim 32, 33 or 34, wherein: The first two-dimensional segmentation model uses a first residual block comprising a first layer that: (i) feeds directly into subsequent layers, and (ii) feeds directly into a layer that is a plurality of layers away from the first layer using a skip connection; and The second 2D segmentation model uses a second residual block comprising a second layer that: (i) feeds directly into subsequent layers, and (ii) feeds directly into a layer a plurality of layers away from the second layer using a skip connection.
36. The computer program product of claim 35, wherein the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
37. The computer program product of any one of claims 32 to 34, wherein the operations further comprise: determining a total metabolic tumor burden TMTV using the final imaging mask; and providing the TMTV.
38. The computer program product of any one of claims 32 to 34, wherein the operations further comprise: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, the three-dimensional organ segmentation model taking the PET scan and the CT or MRI scan as input; Determining a metabolic tumor burden (MTV) and a number of lesions for one or more organs in the three-dimensional organ segmentation using the final imaging mask and the three-dimensional organ mask; and The MTV and the number of lesions of the one or more organs are provided.
39. The computer program product of claim 37, wherein the operations further comprise: using a classifier that takes as input one or more of total metabolic tumor burden TMTV, metabolic tumor burden MTV, and the number of lesions, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: The probability of progression-free survival (PFS) of the subject; the subject's disease stage; and The selection decision to include the subject in the clinical trial.
40. The computer program product of any one of claims 32 to 34, wherein the operations further comprise: inputting a plurality of PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; as well as One or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions are received on a display of a computing device.
41. The computer program product of any one of claims 32 to 34, wherein the operations further comprise administering a treatment to the subject based on one or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions.
42. The computer program product of any one of claims 32 to 34, wherein the operations further comprise providing a diagnosis for the subject based on one or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions.
43. A system comprising: one or more data processors; as well as A non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: obtaining multiple positron emission tomography (PET) scans and multiple computed tomography (CT) or magnetic resonance imaging (MRI) scans of the subject; preprocessing the PET scan and the CT or MRI scan to generate a standardized image that incorporates information from the PET scan and the CT or MRI scan; generating one or more two-dimensional segmentation masks using one or more two-dimensional segmentation models implemented as part of a convolutional neural network architecture and having the normalized image as input; generating one or more three-dimensional segmentation masks using one or more three-dimensional segmentation models implemented as part of the convolutional neural network architecture and taking as input patches of image data associated with segments from the two-dimensional segmentation mask; as well as generating a final imaging mask by combining information from the one or more two-dimensional segmentation masks and the one or more three-dimensional segmentation masks, wherein the standardized images include a first subset of standardized images for a first plane or region of the subject and a second subset of standardized images for a second plane or region of the subject, wherein the first subset of standardized images and the second subset of standardized images incorporate information from the PET scan and the CT or MRI scan; The one or more two-dimensional segmentation models include a first two-dimensional segmentation model and a second two-dimensional segmentation model; and Generating the one or more two-bit segmentation masks includes: generating a first two-dimensional segmentation mask using the first two-dimensional segmentation model implemented to take as input the first subset of the normalized images, wherein the first two-dimensional segmentation model is trained to process images for the first plane or region; as well as A second two-dimensional segmentation mask is generated using the second two-dimensional segmentation model having the second subset of the normalized images as input, wherein the second two-dimensional segmentation model is trained to process images for the second plane or region.
44. The system of claim 43, wherein: The one or more three-dimensional segmentation models include a first three-dimensional segmentation model and a second three-dimensional segmentation model; The image data blocks include: a first image data block associated with a first section; and a second image data block associated with a second section; and Generating the one or more 3D segmentation masks includes: generating a first 3D segmentation mask using the first 3D segmentation model with the first image data block as input, and generating a second 3D segmentation mask using the second 3D segmentation model with the second image data block as input.
45. The system of claim 44, wherein the operations further comprise: evaluating the position of the region or part of the body captured in the standardized image as a reference point; dividing the region or body into a plurality of anatomical regions based on the reference points; generating location labels for the plurality of anatomical regions; Incorporating the position labels into the two-dimensional segmentation mask; determining, based on the location tag, that the first segment is located in a first anatomical region of the plurality of anatomical regions; determining, based on the location tag, that the second segment is located in a second anatomical region of the plurality of anatomical regions; inputting the first image data block associated with the first segment into the first three-dimensional segmentation mask based on a determination that the first segment is located in the first anatomical region; as well as Based on a determination that the second segment is located in the second anatomical region, the second image data block associated with the second segment is input into the second three-dimensional segmentation mask.
46. The system of claim 43, 44 or 45, wherein: The first two-dimensional segmentation model uses a first residual block comprising a first layer that: (i) feeds directly into subsequent layers, and (ii) feeds directly into a layer that is a plurality of layers away from the first layer using a skip connection; and The second 2D segmentation model uses a second residual block comprising a second layer that: (i) feeds directly into subsequent layers, and (ii) feeds directly into a layer a plurality of layers away from the second layer using a skip connection.
47. The system of claim 46, wherein the first layer and the second layer are pyramidal layers with separable convolutions performed at one or more dilation levels.
48. The system of any one of claims 43 to 45, wherein the operations further comprise: determining a total metabolic tumor burden TMTV using the final imaging mask; and providing the TMTV.
49. The system of any one of claims 43 to 45, wherein: The operations further include: generating a three-dimensional organ mask using a three-dimensional organ segmentation model, the three-dimensional organ segmentation model taking the PET scan and the CT or MRI scan as input; Determining a metabolic tumor burden (MTV) and a number of lesions for one or more organs in the three-dimensional organ segmentation using the final imaging mask and the three-dimensional organ mask; and The MTV and the number of lesions of the one or more organs are provided.
50. The system of claim 48, wherein the operations further comprise: using a classifier that takes as input one or more of total metabolic tumor burden TMTV, metabolic tumor burden MTV, and the number of lesions, generating a clinical prediction for the subject based on one or more of the TMTV, the MTV, and the number of lesions, wherein the clinical prediction is one of: The probability of progression-free survival (PFS) of the subject; the subject's disease stage; and The selection decision to include the subject in the clinical trial.
51. The system of any one of claims 43 to 45, wherein: The operations further include: inputting a plurality of PET scans and CT or MRI scans of the subject into a data processing system comprising the convolutional neural network architecture; providing the final imaging mask; and One or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions are received on a display of a computing device.
52. The system of any one of claims 43 to 45, wherein the operations further comprise administering treatment to the subject based on one or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions.
53. The system of any one of claims 43 to 45, wherein the operations further comprise providing a diagnosis for the subject based on one or more of the final imaging mask, total metabolic tumor burden TMTV, metabolic tumor burden MTV, and number of lesions.
Citation Information
Patent Citations
Image segmentation method, apparatus, computer device, and storage medium
CN109410220A