Region-specific lesion segmentation enabling machine learning
By resampling and adjusting PET-CT scans, region-specific scans are generated to improve the signal-to-noise ratio and background uniformity. This solves the problems of low signal-to-noise ratio and background inhomogeneity in PET-CT lesion segmentation, and achieves more accurate lesion identification and disease assessment.
Patent Information
- Application Number
- CN202480041434.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-23
- Filing Date
- 2024-06-21
- Publication Date
- 2026-01-23
AI Technical Summary
Existing positron emission tomography (PET) and computed tomography (CT) scans suffer from low signal-to-noise ratio and background inhomogeneity in lesion segmentation, leading to performance degradation of segmentation models when covering large areas of the body.
By resampling PET-CT scans to generate region-specific scans, unnecessary body areas are excluded, and voxel sizes and values are adjusted to improve the signal-to-noise ratio and background uniformity. A segmentation model is then trained based on the training dataset.
It improves the performance of lesion segmentation models, especially in terms of precision and recall within specific body regions, enabling more accurate identification of lesions and supporting the determination of disease staging, treatment response, and progression.
Smart Images

Figure CN121399656A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 509,981, entitled “MACHINE LEARNING ENABLED REGION-SPECIFIC LESIONSEGMENTATION”, filed on June 23, 2023, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] The topics described in this article generally involve machine learning, and more specifically machine learning-based lesion segmentation in positron emission tomography (PET) and computed tomography (CT) scans. Background Technology
[0003] Medical imaging refers to the techniques and processes used to acquire data characterizing the internal anatomy and pathophysiology of a subject. This data includes images generated, for example, by detecting radiation passing through the body (e.g., X-rays) or radiation emitted by administered radiopharmaceuticals (e.g., gamma rays from intravenously administered radiotracers). By revealing internal anatomy obscured by other tissues such as skin, subcutaneous fat, and bone, medical imaging is an integral part of numerous medical diagnoses and / or treatments. Examples of forms of medical imaging include two-dimensional imaging, such as plain X-ray films, bone scintillation imaging, and thermal imaging. Examples of three-dimensional imaging include magnetic resonance imaging (MRI), computed tomography (CT), cardiac staptin-99m (sestamibi) scans, and positron emission tomography (PET). Summary of the Invention
[0004] Systems, methods, and articles (including computer program products) are provided for machine learning-enabled lesion segmentation in positron emission tomography (PET) and computed tomography (CT) scans. In one aspect, a system for segmenting PET and CT scans is provided. The at least one memory may include program code that provides operation when executed by at least one processor. The operation may include: identifying a first region including a first lesion within a first positron emission tomography and computed tomography (PET-CT) scan depicting multiple regions of the body; extracting a first portion of a second region depicting the first region but not the multiple regions from the first PET-CT scan to generate a second PET-CT scan having one or more initial sizes; adjusting one or more initial sizes of the second PET-CT scan to one or more target sizes, the adjustment including adding one or more voxels to the second PET-CT scan and determining the value of each voxel in the one or more voxels added to the second PET-CT scan; training a segmentation model based at least on a training dataset including the second PET-CT scan adjusted to one or more target sizes; and applying the trained segmentation model to identify one or more lesions present in a region-specific PET-CT scan depicting the first region but not the second region of the multiple regions of the body.
[0005] On the other hand, a method for segmenting positron emission tomography (PET) and computed tomography (CT) scans is provided. The method may include: identifying a first region comprising a first lesion within a first PET-CT scan depicting multiple regions of a body; extracting a first portion of a second region from the first PET-CT scan that depicts the first region but not the multiple regions, to generate a second PET-CT scan having one or more initial sizes; adjusting one or more initial sizes of the second PET-CT scan to one or more target sizes, the adjustment including adding one or more voxels to the second PET-CT scan and determining the value of each voxel in the one or more voxels added to the second PET-CT scan; training a segmentation model based at least on a training dataset including the second PET-CT scan adjusted to one or more target sizes; and applying the trained segmentation model to identify one or more lesions present in a region-specific PET-CT scan depicting the first region but not the second region of the multiple regions of a body.
[0006] In another aspect, a computer program product is provided, comprising a non-transitory computer-readable medium storing instructions. These instructions can cause operations that can be executed by at least one data processor. The operations may include: identifying a first region comprising a first lesion within a first positron emission tomography (PET-CT) scan depicting multiple regions of the body; extracting a first portion of a second region depicting the first region but not the multiple regions from the first PET-CT scan to generate a second PET-CT scan having one or more initial sizes; adjusting one or more initial sizes of the second PET-CT scan to one or more target sizes, the adjustment including adding one or more voxels to the second PET-CT scan and determining the value of each voxel in the one or more voxels added to the second PET-CT scan; training a segmentation model at least based on a training dataset including the second PET-CT scan adjusted to one or more target sizes; and applying the trained segmentation model to identify one or more lesions present in a region-specific PET-CT scan depicting the first region but not the second region of the multiple regions of the body.
[0007] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.
[0008] In some variations, a first tumor mask may be identified, which identifies a first plurality of voxels depicting one or more lesions present in a first PET-CT scan. In response to determining that the overlap between the first tumor mask and a first region of the body satisfies a first threshold, a second PET-CT scan may be generated by extracting at least a first portion of the first PET-CT scan, based at least on the first PET-CT scan.
[0009] In some variations, in response to determining that the first tumor mask satisfies the second threshold, a second PET-CT scan can be generated based at least on the first PET-CT scan.
[0010] In some variations, the second threshold may include at least one of tumor volume and distance to another lesion.
[0011] In some variations, an organ mask can be identified, which identifies a second plurality of voxels depicting a target organ present in a first region of the body. In response to determining that the overlap between the first tumor mask and the organ mask satisfies a first threshold, a second PET-CT scan can be generated, at least based on the first PET-CT scan.
[0012] In some variations, a second tumor mask may be identified, which identifies a second plurality of voxels depicting one or more lesions present in a multi-region PET-CT scan. In response to determining that (i) the overlap between the second tumor mask and a first region of the body fails to meet a first threshold or (ii) the second tumor mask fails to meet a second threshold including at least one of tumor volume and distance to another lesion, a second PET-CT scan may be generated based on a first PET-CT scan rather than a multi-region PET-CT scan.
[0013] In some variations, a training dataset can be generated to include a first benchmark ground truth annotation that identifies a first plurality of voxels depicting a first lesion in a second PET-CT scan.
[0014] In some variations, a second lesion present in a second PET-CT scan can be identified as failing to meet one or more thresholds for at least one of tumor volume and distance to another lesion. In response to determining that a second lesion fails to meet one or more thresholds, a training dataset can be generated to exclude second benchmark ground truth annotations that identify a second plurality of voxels depicting the second lesion in the second PET-CT scan.
[0015] In some variations, adjustments to one or more initial dimensions of the second PET-CT scan may include: performing a first adjustment to a first dimension of the second PET-CT scan; and performing a second adjustment to a second dimension of the second PET-CT scan and a third adjustment to a third dimension of the second PET-CT scan, based at least on the first adjustment.
[0016] In some variations, the first adjustment of the first size of the second PET-CT scan may include: adding at least a first voxel to a first number of voxels along the first size of the second PET-CT scan to increase the first number of voxels to a maximum number of voxels; and determining a first value for the first voxel.
[0017] In some variations, the first value of the first voxel can be determined at least based on the second value of the second voxel within a threshold distance of the first voxel.
[0018] In some variations, the first value of the first voxel can correspond to the first metabolic activity level and the first tissue density at the first position of the first voxel. The second value of the second voxel can correspond to the second metabolic activity level and the first tissue density at the second position of the second voxel.
[0019] In some variations, a second adjustment to the second dimension of the second PET-CT scan may include adding at least a second voxel to a second number of voxels along the second dimension of the second PET-CT scan to increase the second number of voxels to a maximum number of voxels. A third adjustment to the third dimension of the second PET-CT scan may include adding at least a third voxel to a third number of voxels along the third dimension of the second PET-CT scan to increase the third number of voxels to a maximum number of voxels.
[0020] In some variants, each of the second and third voxels can be assigned a first value for the level of metabolic activity and a second value for tissue density or X-ray attenuation.
[0021] In some variations, the first value can be 0, and the second value can be -1024.
[0022] In some variations, the maximum number of voxels can be 64, 128, 256, or 512.
[0023] In some variations, the first PET-CT scan can be a whole-body scan. A second PET-CT scan can be generated to depict some, rather than all, of the regions depicted in the whole-body scan.
[0024] In some variations, the multiple regions of the body depicted in the first PET-CT scan may include multiple organs in the body. A second PET-CT scan may be generated to depict some, rather than all, of the multiple organs in the body.
[0025] In some variations, a first region in a second PET-CT scan may depict at least a first organ, and a second region removed from a first PET-CT scan may depict at least a second organ.
[0026] In some variations, the extraction of a first portion of a first region of the body depicted in a first PET-CT scan may include identifying the multiple voxels based at least on the range of intensity values exhibited by the multiple voxels constituting the first portion of the first PET-CT scan.
[0027] In some variations, the extraction of a first portion of a first region of the body depicted in a first PET-CT scan may include identifying the multiple voxels based at least on the difference between the maximum and minimum intensity values exhibited by the multiple voxels constituting the first portion of the first PET-CT scan.
[0028] In some variations, extraction of a first portion of a first region of the body depicted in a first PET-CT scan may include: determining a first range of intensity values exhibited by a first portion of the first PET-CT scan excluding voxels; determining a second range of intensity values exhibited by a first portion of the first PET-CT scan including voxels; identifying voxels to be included in the first portion of the first PET-CT scan based at least on the difference between the first and second intensity ranges satisfying one or more thresholds; and determining that the voxel is excluded from the first portion of the first PET-CT scan based at least on the difference between the first and second intensity ranges failing to satisfy one or more thresholds.
[0029] In some variations, the first region of the multiple regions may include one of the small intestine, large intestine, lungs, thyroid gland, prostate, pancreas, and cervix.
[0030] In some variations, the segmentation model may include an encoder and a decoder.
[0031] In some variations, the segmentation model may further include one or more transform blocks that couple the encoder's output to the decoder's input.
[0032] In some variants, trained segmentation models can identify multiple voxels depicting one or more lesions present in a first region of the body within a region-specific PET-CT scan.
[0033] In some variants, the metabolic tumor volume of one or more lesions present in the first region of the body can be determined based on at least multiple voxels.
[0034] In some variants, at least one of the following can be determined based on metabolic tumor volume: response to disease stage, disease grade, treatment for disease, disease progression, and disease burden.
[0035] In some variations, the first and second PET-CT scans can be three-dimensional volumes.
[0036] In some variations, the first and second PET-CT scans can be a series of two-dimensional slices.
[0037] Specific implementations of the present subject matter may include, but are not limited to, methods consistent with the descriptions provided herein, and articles of art comprising a tangible, machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to perform operations implementing one or more of the described features. Similarly, computer systems comprising one or more processors and one or more memories coupled to the one or more processors are also described. Memory that may include a non-transitory computer-readable or machine-readable storage medium may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implementation methods consistent with one or more implementations of the present subject matter may be implemented by one or more data processors existing in a single computing system or multiple computing systems. Such multiple computing systems may be interconnected and may exchange data and / or commands or other instructions via one or more connections, including, for example, direct connections between one or more of the multiple computing systems via a network (e.g., the Internet, wireless wide area network, local area network, wide area network, wired network, etc.).
[0038] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Further features and advantages of the subject matter described herein will become apparent from the description, drawings, and claims. While certain features of the currently disclosed subject matter have been described for illustrative purposes in relation to fluorodeoxyglucose-high affinity (FDG-high affinity) carcinomas such as certain types of small bowel tumors, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter. Attached Figure Description
[0039] The accompanying drawings, incorporated in and forming part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the specification, help explain some principles associated with the disclosed embodiments. In the drawings, Figure 1 depicts a system diagram illustrating an example of a machine learning-based medical imaging analysis system according to some exemplary implementations; Figure 2A depicts an example of a process for performing machine learning-enabled lesion segmentation in positron emission tomography and computed tomography (PET-CT) scans, according to some example embodiments; Figure 2B depicts a schematic diagram illustrating another example of a process for machine learning-enabled lesion segmentation in positron emission tomography and computed tomography (PET-CT) scans, according to some example embodiments. Figure 2C depicts a schematic diagram illustrating an example of a workflow for machine learning-enabled lesion segmentation in positron emission tomography and computed tomography (PET-CT) scans, according to some example embodiments. Figure 3 depicts a flowchart illustrating an example of a process for machine learning-enabled lesion segmentation in positron emission tomography and computed tomography (PET-CT) scans, according to some example embodiments; Figure 4 depicts an example flowchart illustrating the process for preprocessing positron emission tomography and computed tomography (PET-CT) scans according to some example embodiments; Figure 5 depicts an example flowchart illustrating a process for identifying one or more positron emission tomography (PET) and computed tomography (PET-CT) scans to be included in a training dataset, according to some example embodiments; Figure 6 depicts a schematic diagram illustrating an example architecture for implementing a segmentation model according to some example embodiments; Figure 7A depicts a comparison of error bars and 95% confidence intervals for Dice on the Goya test set, illustrating region-specific (e.g., organ-focused) segmentation methods and multi-region (e.g., whole-body) segmentation methods according to some example embodiments. Figure 7B depicts a comparison of error bars and 95% confidence intervals for Dice on the Gallium test set, illustrating region-specific (e.g., organ-focused) segmentation methods and multi-region (e.g., whole-body) segmentation methods according to some example embodiments. Figure 8A depicts a graph comparing the automated total metabolic tumor volume (TMTV) of a patient with diffuse large B-cell lymphoma (DLBCL) according to some example embodiments with the baseline true value; Figure 8A depicts a graph showing a comparison of the automated total metabolic tumor volume (TMTV) of patients with follicular lymphoma (FL) according to some example embodiments with the baseline true value; Figure 9 depicts a comparison of the first performance of a segmentation model trained to perform lesion segmentation on organ-focused positron emission tomography (PET-CT) scans, according to some example embodiments, and the second performance of a segmentation model trained to perform lesion segmentation on whole-body PET-CT scans; and Figure 10 depicts a block diagram illustrating an example of a computing system according to some exemplary embodiments.
[0040] In practical applications, similar reference numerals indicate similar structures, features, or elements. Detailed Implementation
[0041] Various forms of medical imaging can be used to obtain data characterizing a subject's internal anatomy and pathophysiology. Computed tomography (CT) is an example of a three-dimensional imaging form, in which a series of X-rays is captured to generate cross-sectional images (e.g., image patches, slices, etc.) of bones, blood vessels, and soft tissues within the body. A CT scan can be a three-dimensional volume formed by a series of two-dimensional images (or slices), where each pixel is associated with an intensity value indicating tissue density or X-ray attenuation at a corresponding location in the subject's body. Another example of a three-dimensional imaging form is positron emission tomography (PET), which captures radioactive signals indicating cellular metabolic activity within a subject's body. A PET scan can be a three-dimensional volume formed by a series of two-dimensional images (or slices), where each pixel is associated with an intensity value indicating the level of cellular metabolic activity (e.g., glucose uptake) at a corresponding location in the subject's body. In some cases, a single gantry combining a positron emission tomography (PET) scanner and a computed tomography (CT) scanner can acquire both PET and CT scans during the same session. The resulting PET and CT scans can be combined into a single overlay (e.g., co-registered) image (e.g., a PET-CT scan) where the spatial distribution of metabolic activity depicted in the PET scan is aligned with the anatomical structures depicted in the CT scan.
[0042] In some example embodiments, the analysis controller can perform lesion segmentation by applying a segmentation model to identify one or more voxels depicting the lesion within positron emission tomography (PET) and computed tomography (PET-CT) scans. For example, in some cases, a PET-CT scan can be a three-dimensional volume formed by a series of two-dimensional images (or slices). A single two-dimensional image (or slice) in a PET-CT scan can include multiple pixels forming, for example, a two-dimensional plane. Simultaneously, in some cases, a PET-CT scan can include multiple voxels, wherein each voxel in the three-dimensional volume of the PET-CT scan corresponds to a single pixel in one of the two-dimensional images (or slices) forming the PET-CT scan. A single voxel in the PET scan forming the PET-CT scan can be associated with an intensity value (e.g., a standard update value (SUV)) indicating the level of metabolic activity at the corresponding location in the patient's body, while a corresponding voxel in a co-registered computed tomography (CT) scan can be associated with an intensity value indicating tissue density or X-ray attenuation at the same location. In some cases, the segmentation model described above can determine whether a voxel in a PET-CT scan is part of a lesion based at least on the voxel's intensity value and the intensity values of one or more neighboring voxels. That is, the segmentation model can determine whether a voxel is part of a lesion based at least on the level of metabolic activity and tissue density (or X-ray attenuation) present at the corresponding location in the patient's body. In this case, the segmentation model can generate a tumor mask in which voxels are assigned a first value (e.g., "1") to indicate that the voxel is part of a lesion or a second value (e.g., "0") to indicate that the voxel is not part of a lesion.
[0043] Segmentation models may perform poorly when trained to operate on PET-CT scans depicting large, general areas of the body (e.g., whole-body scans). Voxels spanning large areas of the body (such as in whole-body scans) can exhibit a wider range of intensity values than voxels depicting lesions. Therefore, in some cases, the poor performance of segmentation models trained on PET-CT scans covering large areas of the body (e.g., whole-body scans) may be at least partly attributable to the low signal-to-noise ratio (SNR) of these scans. That is, in PET-CT scans depicting large areas of the body (e.g., whole-body scans), the signals associated with voxels depicting lesions tend to be masked by the high noise levels imposed by background voxels (or voxels not depicting lesions).
[0044] Anatomical differences across the body can also cause significant variations in the range of voxel intensity values across different regions of the body. However, voxels within a single region of the body may be more homogeneous, at least because voxels within the same region of the body tend to exhibit more uniform intensity values (or intensity values from a more limited range of intensity values). For example, in some cases, a first background voxel (or a voxel that does not depict a lesion) from a first region of the body may exhibit a first intensity value that is more similar to a second background voxel from the first region of the body than a third background voxel from a second region of the body has a third intensity value. Therefore, in some example embodiments, the analysis controller can improve the performance of the segmentation model by at least the following: training the segmentation model to perform lesion segmentation based on PET-CT scans that have been resampled to exclude at least one region of the body. For example, in some cases, the analysis controller may train the segmentation model to perform lesion segmentation based on region-specific PET-CT scans that have been resampled to depict a first region of the body but not a second region. As described in more detail below, the performance of segmentation models can be improved, at least because the signal-to-noise ratio (SNR) and background homogeneity of region-specific PET-CT scans can be increased by excluding at least one region of the body from these scans.
[0045] In some example embodiments, the analysis controller may resample a first PET-CT scan depicting multiple regions of the body by at least the following: identifying a first region including a first lesion within the first PET-CT scan. The analysis controller may extract a first portion of the first PET-CT scan depicting the first region of the body but not a second region to generate a second PET-CT scan. In some cases, resampling may further include the analysis controller adjusting one or more initial sizes of the second PET-CT scan to one or more target sizes. For example, in some cases, the second PET-CT scan may be adjusted to one or more target sizes by adding at least one or more voxels to a first portion of the first PET-CT scan and determining the value of each voxel in the one or more voxels added to the second PET-CT scan. Furthermore, the analysis controller may generate a training dataset to include the second PET-CT scan adjusted to one or more target sizes and benchmark ground truth annotations identifying one or more voxels depicting the first lesion in the second PET-CT scan. In some cases, the analysis controller may train a segmentation model based at least on a training dataset including the second PET-CT scan adjusted to said one or more target sizes. The trained segmentation model can be applied to identify one or more lesions present in a region-specific PET-CT scan, such as one that depicts the first region of the body.
[0046] In some example embodiments, the analysis controller may generate a training dataset based on one or more multi-region PET-CT scans in which lesions present in a first region of the body satisfy one or more thresholds. For example, in some cases, the analysis controller may determine a first tumor mask that identifies a first plurality of voxels depicting one or more lesions present in a multi-region PET-CT scan (such as a first PET-CT scan). When the overlap between the first tumor mask and the first region of the body satisfies a first threshold (e.g., 20%, etc.), the analysis controller may generate a region-specific scan, such as a second PET-CT scan, based at least on the first PET-CT scan, for example by extracting at least a first portion of the first PET-CT. Where the first region includes or is limited to one or more specific target organs in the body (e.g., small intestine, large intestine, lung, thyroid gland, prostate, pancreas, cervix, etc.), the analysis controller may generate a second PET-CT scan based on the first PET-CT scan if the overlap between the first tumor mask and an organ mask that includes one or more organs satisfies a first threshold (e.g., 20%, etc.). In contrast, if the overlap between the first region of the body and the second tumor mask identifying one or more lesions present in a multi-region PET-CT scan fails to meet a first threshold (e.g., 20%), the analysis controller can omit the multi-region PET-CT scan when generating the training dataset for the segmentation model. In this case, the analysis controller can ensure that the training dataset for training segmentation includes resampled PET-CT scans in which at least one lesion is present in the first region of the body.
[0047] In some example embodiments, in addition to having sufficient overlap with a first region of the body, the analysis controller may also exclude one or more PET-CT scans from the resampled PET-CT scans in the training dataset used for segmentation models that depict lesions more likely to be artifacts. For example, in some cases, the analysis controller may determine whether a first tumor mask identifying one or more lesions present in a first region of the body meets a second threshold. In some cases, the second threshold may include at least one of tumor volume and distance to another lesion. Thus, if one or more lesions identified by the first tumor mask exceed a threshold volume (e.g., 0.8 ml, etc.) and / or are located within a threshold distance (e.g., 10 voxels, etc.) from another lesion, a second PET-CT scan may be determined at least based on the first PET-CT scan, for example, by extracting at least a first portion of the first PET-CT scan rather than a second portion. In contrast, if one or more lesions identified by the first tumor mask are smaller than a threshold volume (e.g., 0.8 ml, etc.) and are located at a threshold distance from another lesion (e.g., 10 voxels, etc.), the analysis controller can omit the first PET-CT scan when generating the training dataset for the segmentation model.
[0048] In some example embodiments, when extracting a first portion of a first PET-CT scan, the analysis controller can adjust one or more initial dimensions of the resulting second PET-CT scan by performing at least a first adjustment to a first dimension of the second PET-CT scan. Furthermore, the analysis controller can perform a second adjustment to a second dimension of the second PET-CT scan and a third adjustment to a third dimension of the second PET-CT scan, based at least on the first adjustment. In some cases, the first adjustment to the first dimension of the first PET-CT scan may include adding at least a first voxel to a first number of voxels along the first dimension of the second PET-CT scan to increase the first number of voxels to a maximum number of voxels (e.g., 64, 128, 256, 512, etc.). In this case, the first value of the first voxel can be determined at least based on a second value of a second voxel within a threshold distance of the first voxel. In some cases, a second adjustment to the second size of the second PET-CT scan may include adding at least a second voxel to a second number of voxels along the second size of the second PET-CT scan to increase the second number of voxels to a maximum number of voxels, while a third adjustment to the third size of the second PET-CT scan may include adding at least a third voxel to a third number of voxels along the third size of the second PET-CT scan to increase the third number of voxels to a maximum number of voxels. The second and third voxels may each be assigned a first value for a metabolic activity level (e.g., "0") and a second value for tissue density or X-ray attenuation (e.g., "-1024").
[0049] In some example embodiments, the analysis controller may extract a first portion of a first PET-CT scan depicting a first region of the body to maximize the uniformity of background voxels (or voxels not depicting lesions) included in the first portion of the first PET-CT scan. For example, in some cases, the analysis controller may extract the first portion of the first PET-CT scan by identifying at least a plurality of voxels forming the first portion of the first PET-CT scan. In some cases, the plurality of voxels forming the first portion of the first PET-CT scan may be identified at least based on the range of intensity values exhibited by the voxels (including, for example, the difference between the maximum and minimum intensity values of the voxels selected for inclusion in the first portion of the first PET-CT scan). For example, in some cases, a voxel may be identified to be included among the plurality of voxels forming the first portion of the first PET-CT if the intensity value of a pixel meets one or more criteria. In some cases, if the intensity value of a voxel is within a specific range of intensity values, it may be determined that the intensity value of the voxel meets one or more criteria. Alternatively, if including a voxel in the first portion of a first PET-CT scan does not result in an increase above a threshold in the range of intensity values present in the first portion of the first PET-CT scan, then the voxel's intensity value can be determined to meet one or more criteria. By excluding voxels that skew the range of intensity values present in the first portion of the first PET-CT scan, the analysis controller can increase the signal-to-noise ratio (SNR) and background homogeneity of the second PET-CT scan generated from it. Training this segmentation model to operate on region-specific PET-CT scans with high SNR and background homogeneity (such as PET-CT scans that have been resampled to depict a specific region of the body while excluding other regions) can maximize the performance of the segmentation model when applied to region-specific PEC-CT scans.
[0050] In some example embodiments, when applied to region-specific PET-CT scans that depict a first region of the body but not a second region, a trained segmentation model can identify one or more lesions present in the region-specific PET-CT scan by identifying at least multiple voxels depicting one or more lesions. In some cases, the analysis controller can determine the metabolic tumor volume (or total metabolic tumor volume (TMTV)) of one or more lesions based at least on multiple voxels depicting one or more lesions in the region-specific PET-CT scan. Furthermore, in some cases, the analysis controller can determine at least one of the following based at least on the metabolic tumor volume (or total metabolic tumor volume (TMTV)): response to disease stage, disease grade, treatment for the disease, disease progression, and disease burden. For example, in some cases, the analysis controller can determine the response to treatment for the disease (e.g., complete metabolic response (CMR), objective response (OR), four-category assessment, etc.) based at least on the metabolic tumor volume (or total metabolic tumor volume (TMTV)). Alternatively and / or additionally, the analysis controller can determine disease progression based at least on the metabolic tumor volume (or total metabolic tumor volume (TMTV)) from two or more different time points.
[0051] Figure 1 depicts a system diagram illustrating an example of a machine learning-based medical imaging analysis system 100 according to some exemplary embodiments. Referring to Figure 1, the machine learning-based medical imaging analysis system 100 may include an analysis controller 110, one or more imaging devices 120, and a client device 130. As shown in Figure 1, the analysis controller 110, one or more imaging devices 120, and the client device 130 may be communicatively connected via a network 140. The one or more imaging devices 120 may include, for example, a computed tomography (CT) scanner 121 and a positron emission tomography (PET) scanner 123. The client device 130 may be a processor-based device, including, for example, a smartphone, tablet computer, wearable device, virtual assistant, Internet of Things (IoT) device, etc. The network 140 may be a wired network and / or a wireless network, including, for example, a wide area network (WAN), a local area network (LAN), a virtual local area network (VLAN), a public land mobile network (PLMN), the Internet, etc.
[0052] In some example embodiments, the analysis controller 110 may perform segmentation by training a segmentation model 113 to operate on region-specific positron emission tomography (PET) and computed tomography (PET-CT) scans extracted from PET-CT scans depicting multiple regions of the body. In some example embodiments, the analysis controller 110 may receive a first PET-CT scan generated by one or more imaging devices 120 (e.g., computed tomography scanner 121, positron emission tomography (PET) scanner 123, etc.). The first PET-CT scan may be a three-dimensional volume, wherein a three-dimensional computed tomography (CT) scan captured by CT scanner 121 is co-registered with a three-dimensional positron emission tomography (PET) scan captured by PET scanner 123. Furthermore, in some cases, the first PET-CT scan received from one or more imaging devices 120 (e.g., a whole-body scan, etc.) may depict multiple regions of the body. As described in more detail below, the analysis controller 110 can resample a first PET-CT scan depicting multiple regions of the body to generate a second PET-CT scan, which includes the first region of the body depicted in the first PET-CT scan but excludes a second region of the body. In this case, the second PET-CT scan depicting one or more specific regions of the body can exhibit a higher signal-to-noise ratio (SNR) and background uniformity than the first PET-CT scan. Training the segmentation model 113 to operate on the second PET-CT scan instead of the first PET-CT scan can improve the performance of the segmentation model 113 (e.g., precision, recall, F1 score, etc.).
[0053] Figure 2A depicts a schematic diagram illustrating an example of a process 200 for training a segmentation model 113 to perform lesion segmentation in positron emission tomography (PET) and computed tomography (CT) scans, according to some example embodiments. As shown in Figure 2A, an analysis controller 110 may receive a first PET-CT scan 210 from one or more imaging devices 120. In some cases, the first PET-CT scan 210 may include a PET scan 213 co-registered with a computed tomography (CT) scan 215. Furthermore, in some cases, the PET scan 213 and the CT scan 215 may each be a three-dimensional volume formed by a series of two-dimensional images (or slices). Thus, the first PET-CT scan 210 may be a three-dimensional volume comprising a plurality of voxels, wherein each voxel corresponds to a pixel in one of the two-dimensional images (or slices) forming the first PET-CT scan 210.
[0054] In some cases, each of the two-dimensional images (or slices) forming the first PET-CT scan 210 can be, for example, along... x shaft and y The axis is a two-dimensional plane formed by multiple pixels. The three-dimensional volume of the first PET-CT scan 210 can be obtained by following, for example... z Two-dimensional images (or slices) are formed by stacking axes. A single voxel in the three-dimensional volume of the first PET-CT scan 210 can correspond to a pixel in one of the constituent two-dimensional images (or slices) of the first PET-CT scan 210. For example, the first PET-CT scan 210 has a first three-dimensional coordinate ( x 1, y 1, z 1) The first voxel can be in the first two-dimensional image (or slice) in z 1 location has two-dimensional coordinates ( x 1, y 1) corresponds to the first pixel and has a second three-dimensional coordinate ( x 2, y 2, z 2) The second voxel can be in the second two-dimensional image (or slice) in z Two locations have two-dimensional coordinates ( x 2, y2) Corresponding to the second pixel. In addition, each voxel in the first PET-CT scan 210 may have a first intensity value (e.g., standard update value (SUV)) indicating the level of metabolic activity at the corresponding location in the patient's body and a second intensity value indicating the tissue density or X-ray attenuation at the same location.
[0055] In some example embodiments, the analysis controller 110 may generate a second PET-CT scan 220 based at least on a first PET-CT scan 210, which is then used as a training sample for training the segmentation model 113. In the example shown in FIG2A, the analysis controller 110 may include a preprocessing engine 111 that generates the second PET-CT scan 220 by resampling at least one first PET-CT scan 210 received from one or more imaging devices 120. For example, in some cases, the first PET-CT scan 210 may depict multiple regions of the body. Therefore, the preprocessing engine 111 may extract a first portion of a first region of the body depicted in the first PET-CT scan 210, rather than a second region of the body, from the first PET-CT scan 210 depicting multiple regions of the body. That is, the preprocessing engine 111 may generate the second PET-CT scan 220 to exclude at least one region among the regions depicted in the first PET-CT scan 210. The first PET-CT scan 210 can have a low signal-to-noise ratio (SNR) and background homogeneity, at least in part due to the large differences in voxel intensity values across multiple different regions of the body. Therefore, by eliminating one or more regions present in the first PET-CT scan 210, such as regions that do not contain the target organ, the resulting second PET-CT scan 220 can exhibit a higher signal-to-noise ratio (SNR) and background homogeneity than the first PET-CT scan 210.
[0056] In some example embodiments, the preprocessing engine 111 may further generate a second PET-CT scan 220 by resampling at least a first portion extracted from the first PET-CT scan 210. For example, in some cases, extracting a first portion of the first PET-CT scan 210 that depicts a first region of the body but not a second region may generate a second PET-CT scan 220 with one or more initial sizes. As described in more detail below, resampling the first portion of the first PET-CT scan 210 may include adjusting one or more initial sizes of the second PET-CT scan 220 to one or more target sizes. For example, in some cases, one or more initial sizes of the second PET-CT scan 220 may be adjusted by adding one or more voxels to the second PET-CT scan 220 to achieve one or more target sizes. Furthermore, in some cases, one or more initial sizes of the second PET-CT scan 220 may be adjusted by at least determining the values (e.g., intensity values, etc.) of one or more voxels added to the second PET-CT scan 220.
[0057] In some example embodiments, the analysis controller 110 may train the segmentation model 113 based at least on a training dataset including a second PET-CT scan 220 adjusted to one or more target sizes. As noted above, training the segmentation model 113 to operate on the second PET-CT scan 220 depicting one or more specific regions of the body can improve the performance of the segmentation model 113. For example, as shown in FIG2A, the analysis controller 110 may train the segmentation model 113 by at least the following: applying the segmentation model 113 to generate a tumor mask 225 that identifies a first plurality of voxels depicting one or more lesions present in the second PET-CT scan 220. Training the segmentation model 113 may further include adjusting the segmentation model 113 (e.g., one or more weights imposed by the segmentation model 113) to minimize the difference (e.g., Dice coefficients, etc.) between the tumor mask 225 and a benchmark ground truth tumor mask 227 that identifies a second plurality of voxels depicting one or more actual lesions present in the second PET-CT scan 220.
[0058] For further illustration, Figure 6 depicts a schematic diagram illustrating an example architecture for implementing segmentation model 113 according to some example embodiments. In the example shown in Figure 6, segmentation model 113 may include cascaded two-dimensional and three-dimensional artificial neural networks, such as convolutional neural networks. For example, as shown in Figure 6, segmentation model 113 may include an encoder trained to generate an encoding of, for example, a second PET-CT scan 220. Furthermore, segmentation model 113 may also include a decoder trained to decode the encoding of the second PET-CT scan 220. In some cases, the output of the decoder decoding the encoding of the second PET-CT scan 220 may be a tumor mask 225. Furthermore, in the example shown in Figure 6, segmentation model 113 may couple the output of the encoder to one or more transform blocks of the input of the decoder. Therefore, in the example shown in Figure 6, the encoding of the second PET-CT scan 220 output by the encoder may undergo one or more transforms before being ingested by the decoder and decoded into the tumor mask 225.
[0059] In some cases, the example of segmentation model 113 shown in Figure 6 can implement a two-step approach involving both 2D and 3D segmentation. For example, in some cases, individual slices of a PET-CT scan 220 can undergo 2D segmentation using an adapted UNet before employing an improved 3D VNet architecture to enhance 2D segmentation. The tumor mask 225 can be derived by averaging the tumor masks generated from the 2D segmentation performed by the UNet and the subsequent 3D segmentation performed by the VNet.
[0060] Figure 2B illustrates an example of a process 250 for machine learning-enabled lesion segmentation in positron emission tomography (PET) and computed tomography (CT) scans, according to some example embodiments. Referring to Figures 1 and 2A through 2B, once trained, the segmentation model 113 can be applied to identify one or more lesions present in, for example, a region-specific PET-CT scan 260. In some cases, the region-specific PET-CT scan 260 may also be preprocessed, for example, by a preprocessing engine 111 to depict a first region of the body but not a second region. For example, as shown in Figure 2B, the analysis controller 110 may receive multi-region PET-CT scans 270 depicting multiple regions of the body (e.g., including a PET scan 273 co-registered with a CT scan 275) from one or more imaging devices 120. The preprocessing engine 111 can generate a region-specific PET-CT scan 260 by at least the following: extracting a portion of a first region of the body, but not a second region of the body, from a multi-region PET-CT scan 270. Furthermore, the preprocessing engine 111 can generate a region-specific PET-CT scan 260 by at least resizing the region-specific PET-CT scan 260 to one or more target sizes.
[0061] Referring again to Figure 2B, the analysis controller 110 can apply a trained segmentation model 113 to identify one or more lesions present in a region-specific PET-CT scan 260 adjusted to one or more target sizes. For example, in the example shown in Figure 2B, the segmentation model 113 can generate a tumor mask 277 that identifies multiple voxels depicting lesions in the region-specific PET-CT scan 260. As shown in Figure 2B, in some cases, the analysis controller 110 may include an assessment engine 115 that determines a tumor volume 280 (e.g., metabolic tumor volume, total metabolic tumor volume (TMTV), etc.) based at least on the tumor mask 277. In some cases, the assessment engine 115 may further determine at least one of the following based at least on the tumor volume 280: disease stage, disease grade, response to treatment for the disease, disease progression, and disease burden.
[0062] Figure 2C illustrates an example of a workflow 290 for machine learning-enabled lesion segmentation in positron emission tomography (PET) and computed tomography (PET-CT) scans, according to some example embodiments. In the example shown in Figure 2C, a multi-region PET-CT scan 291 (e.g., a whole-body scan, etc.) that may depict multiple regions of the body can be processed prior to segmentation. As shown in Figure 2C, preprocessing of the multi-region PET-CT scan 291 may include extracting a region-specific PET-CT scan 293 from the multi-region PET-CT scan 291 that depicts a first region of the body but not a second region. In some cases, preprocessing may further include removing background from the region-specific PET-CT scan 293 extracted from the multi-region PET-CT scan 291 to increase background uniformity of the region-specific PET-CT scan 293. In some cases, preprocessing may also include resampling the region-specific PET-CT scan 293, for example, adding and / or removing a number of voxels along one or more dimensions of the region-specific PET-CT scan 293, such that the resulting region-specific PET-CT scan 293 includes a threshold number of voxels along each dimension (e.g., 128 × 128 × 128). In some cases, the region-specific PET-CT scan 293 may be resampled to increase the signal-to-noise ratio (SNR) between the signal associated with a first number of voxels depicting the lesion and the noise associated with a second number of background voxels (or voxels not depicting the lesion). As shown in Figure 2C, the region-specific PET-CT scan 293 generated from the preprocessing of the multi-region PET-CT scan 291 may subsequently undergo segmentation. In some cases, segmentation may include applying a segmentation model 113 to the region-specific PET-CT scan 293 to generate a tumor mask 295. In some cases, segmentation model 113 can generate tumor mask 295 by assigning at least a first value (e.g., "1") to each voxel depicting a lesion and a second value (e.g., "0") to each voxel not depicting a lesion, to locate one or more lesions present in a multi-region PET-CT scan 291.
[0063] Figure 3 depicts an example flowchart illustrating a process 300 for machine learning-enabled lesion segmentation in positron emission tomography (PET) and computed tomography (CT) scans, according to some example embodiments. Referring to Figures 1, 2A through 2B, and 3, process 300 can be executed by an analysis controller 110 to train and apply a segmentation model 113 to perform lesion segmentation in one or more region-specific PET-CT scans.
[0064] At 302, the analysis controller 110 can preprocess a first PET-CT scan depicting multiple regions of the body to generate a second PET-CT scan depicting a first region but not a second region among the multiple regions of the body. In some exemplary embodiments, the preprocessing engine 111 of the analysis controller 110 can preprocess, for example, the first PET-CT scan 210 depicting multiple regions of the body by at least: extracting a first portion of the first PET-CT scan 210 depicting the first region but not the second region of the body. In this case, the preprocessing engine 111 can generate a second PET-CT scan 220 having one or more initial sizes. As described in more detail below, the second PET-CT scan 220 can be further preprocessed by at least adjusting the second PET-CT scan 220 from one or more initial sizes to one or more target sizes. The first PET-CT scan 210 is preprocessed to generate a second PET-CT scan 220, which exhibits a higher signal-to-noise ratio (SNR) and background uniformity than the first PET-CT scan 210. Therefore, the segmentation model 113 can perform better when trained on the second PET-CT scan 220 than when trained on the first PET-CT scan 210 (e.g., with better precision, recall, F1 score, etc.).
[0065] At 304, the analysis controller 110 may train the segmentation model based at least on a training dataset including the second PET-CT scan. In some example embodiments, the analysis controller 110 may train the segmentation model 113 to perform lesion segmentation based at least on a training dataset including the second PET-CT scan 220. For example, in the example shown in FIG2A, training the segmentation model 113 may include applying the segmentation model 113 to generate a tumor mask 225 that identifies a first plurality of voxels depicting one or more lesions in the second PET-CT scan 220. That is, in some cases, the segmentation model 113 may generate the tumor mask 225 by assigning each voxel in the second PET-CT scan 220 a first value (e.g., "1") to identify the voxel as depicting a lesion and a second value (e.g., "0") to identify the voxel as not depicting a lesion. To train segmentation model 113, analysis controller 110 may adjust segmentation model 113 (e.g., by applying one or more weights from the weights) to minimize the error of the output of segmentation model 113. In the example shown in Figure 2A, the error of the output of segmentation model 113 may include the difference between tumor mask 225 and a baseline ground truth tumor mask 227, which identifies a second plurality of voxels depicting one or more actual lesions present in the second PET-CT scan 220. It should be understood that segmentation model 113 may undergo one or more training iterations, each training iteration including segmentation model 113 being adjusted to minimize the error of the output of segmentation model 113 applied to one or more training samples of region-specific PET-CT scans (such as the second PET-CT scan 220).
[0066] In some example embodiments, the parameters of segmentation model 113 (e.g., weights, biases, etc.) may be initially set to values randomly selected from a normal distribution. In some cases, gradient descent, such as stochastic gradient descent, may be performed to adjust the parameters of segmentation 113 to reduce (or minimize) the error in the output of segmentation model 113 by a loss function. Equation (1) below is an example of a loss function that includes cross-entropy loss and Dice loss to address class imbalances present in the training samples. In Equation (1), V This refers to a single voxel within a PET-CT scan. T This represents the set of voxels (positive pixels) that depict the lesions in the baseline truth. P This represents the set of voxels predicted by segmentation model 113 to depict the lesion. y vVoxels in the baseline true value tumor mask v The value, and Voxels in the tumor mask predicted by segmentation model 113 v The value of . In the context of lesion segmentation, class imbalance may refer to a significantly higher proportion of voxels that do not depict lesions (or negative voxels) relative to voxels that actually depict lesions (or positive voxels). Loss functions that combine the two losses (such as equation (1)) mitigate the effects of class imbalance and encourage segmentation model 113 to produce accurate pixel-level probabilities and well-defined lesion boundaries.
[0067] At 306, the analysis controller 110 may apply a trained segmentation model to identify one or more lesions present in a region-specific PET-CT scan depicting a first region but not a second region of the body across multiple regions. In some example embodiments, the analysis controller 110 may apply a trained segmentation model 113 to identify one or more lesions present in, for example, a region-specific PET-CT scan 260. As shown in FIG2B, in some cases, the region-specific PET-CT scan 260 may be a region-specific PET-CT scan generated by the preprocessing engine 111 preprocessing the multi-region PET-CT scan 270 by resampling at least a portion of the first region of the body depicting but not the second region extracted from the multi-region PET-CT scan 270. In some cases, the trained segmentation model 113 may perform lesion segmentation on the region-specific PET-CT scan 260 by generating at least a tumor mask 277 that identifies multiple voxels depicting lesions in the region-specific PET-CT scan 260. For example, in some cases, the trained segmentation model 113 can generate the tumor mask 277 by at least assigning a first value (e.g., "1") to each voxel in the region-specific PET-Ct scan 260 to identify the voxel as depicting a lesion and a second value (e.g., "0") to identify the voxel as not depicting a lesion. In the example shown in Figure 2B, in some cases, the analysis controller 110 may include an assessment engine 115 that determines the tumor volume 280 (e.g., metabolic tumor volume, total metabolic tumor volume (TMTV), etc.) based at least on the tumor mask 265. In some cases, the assessment engine 115 may further determine at least one of the following based at least on the tumor volume 280: disease stage, disease grade, response to treatment for the disease, disease progression, and disease burden.
[0068] Figure 4 depicts an example flowchart illustrating process 400 for preprocessing positron emission tomography and computed tomography (PET-CT) scans according to some example embodiments. Referring to Figures 1, 2A-2B, and 3-4, in some cases, process 400 may be performed by an analysis controller 110 (e.g., a preprocessing engine 111). Furthermore, in some cases, process 400 may implement operation 302 of process 300 shown in Figure 3. In some example embodiments, process 400 may be performed to generate a second PET-CT scan depicting a first region of the multiple regions of the body but not a second region, based at least on a first PET-CT scan depicting multiple regions of the body. It should be understood that the second PET-CT scan may exhibit a higher signal-to-noise ratio (SNR) and background uniformity than the first PET-CT scan, at least because the exclusion of the second region of the body reduces the voxel intensity value differences present in the first PET-CT scan.
[0069] At 402, the preprocessing engine 111 can identify a first region including one or more lesions within a first PET-CT scan depicting multiple regions of the body. In some example embodiments, the analysis controller 110 can receive a first PET-CT scan 210 from one or more imaging devices 120 that can depict multiple regions of the body. For example, in some cases, the first PET-CT scan 210 can depict two or more of the head, neck, upper limbs, chest, abdomen, pelvis, and lower limbs. Furthermore, in some cases, the first PET-CT scan 210 can depict multiple organs. In some cases, the preprocessing engine 111 can preprocess the first PET-CT scan by identifying at least one region of the body at least within the first PET-CT scan 210. For example, in some cases, the preprocessing engine 111 can identify one or more regions of the body within the first PET-CT scan 210 where lesions may be present. Alternatively and / or additionally, the preprocessing engine 111 may identify one or more regions of the body within the first PET-CT scan 210, which contain at least one target organ, such as, for example, the small intestine, large intestine, lungs, thyroid gland, prostate, pancreas, cervix, etc.
[0070] At 404, the preprocessing engine 111 can extract a portion of a second region depicting a first region but not a second region among multiple regions from the first PET-CT scan to generate a second PET-CT scan with one or more initial dimensions. In some example embodiments, the preprocessing engine 111 can extract a first portion of the first PET-CT scan 210 that depicts a first region among multiple regions of the body depicted in the first PET-CT scan 210 but not a second region. In this case, the preprocessing engine 111 can generate a second PET-CT scan 220 that depicts the first region of the body but not the second region. That is, in some cases, the preprocessing engine 111 can generate the second PET-CT scan 220 by at least removing a second portion of the second region of the body depicted in the first PET-CT scan 210. In some cases, although a first PET-CT scan 210 depicts multiple organs of the body, a first portion extracted from the first PET-CT scan 210 to generate a second PET-CT scan 220 may depict some, but not all, of the organs depicted in the first PET-CT scan 210. For example, in some cases, although a first region of the body depicted in the first portion of the first PET-CT scan 210 extracted to form the second PET-CT scan 220 may depict a first organ, a second region of the body excluded from the second PET-CT scan 220 may depict a second organ instead of the first organ.
[0071] In some example embodiments, the preprocessing engine 111 may extract a first portion of a first region depicting the body from the first PET-CT scan 210 by identifying multiple voxels to form the first portion of the first PET-CT scan 210 based at least on the intensity value range exhibited by multiple voxels. For example, in some cases, the preprocessing engine 111 may extract the first portion of the first PET-CT scan 210 by identifying multiple voxels based on the difference between the maximum and minimum intensity values exhibited by multiple voxels to form the first portion of the first PET-CT scan 210. Therefore, when extracting the first portion of the first PET-CT scan 210, the preprocessing engine 111 may avoid including voxels whose intensity values skew the intensity value range present in the first portion of the first PET-CT scan 210, thereby maximizing the signal-to-noise ratio (SNR) and background uniformity of the second PET-CT scan 220 generated from the first portion. For example, in some cases, the preprocessing engine 111 can determine a first intensity value range exhibited by a first portion of the first PET-CT scan 210 excluding the voxel and a second intensity value range exhibited by a first portion of the first PET-CT scan 210 including the voxel. If the difference between the first intensity value range and the second intensity value range satisfies one or more thresholds, the preprocessing engine 111 can extract the voxel when extracting the first portion of the first PET-CT scan 210. Alternatively, if the difference between the first intensity value range and the second intensity value range does not satisfy one or more thresholds, the preprocessing engine 111 can exclude the voxel when extracting the first portion of the first PET-CT scan 210.
[0072] At 406, the preprocessing engine 111 can adjust one or more initial sizes of the second PET-CT scan to one or more target sizes. In some example embodiments, the second PET-CT scan 220 generated by extracting a first portion of the first PET-CT scan 210 through the preprocessing engine 111 may have one or more initial sizes. However, to further optimize the performance of the segmentation model 113, the preprocessing engine 111 may resample the second PET-CT scan 220 to increase the resolution of the first portion extracted from the first PET-CT scan 210. For example, in some cases, the preprocessing engine 111 may resample the second PET-CT scan 220 by at least adjusting one or more initial sizes of the second PET-CT scan 220 to one or more target sizes.
[0073] In some example embodiments, the preprocessing engine 111 may adjust one or more initial dimensions of the second PET-CT scan 220 by performing at least a first adjustment to a first dimension of the second PET-CT scan 220. For example, in some cases, the first adjustment may include adding at least a first voxel to a first number of voxels along the first dimension of the second PET-CT scan 220 to increase the first number of voxels to a maximum number of voxels (e.g., 64, 128, 256, 512, etc.). Furthermore, the preprocessing engine 111 may determine a first value for the first voxel added to the first dimension of the second PET-CT scan 220. For example, in some cases, the preprocessing engine 111 may determine the first value of the first voxel based at least on a second value of a second voxel within a threshold distance of the first voxel. In some cases, the first value of the first voxel may correspond to a first metabolic activity level and a first tissue density at a first location of the first voxel, while the second value of the second voxel may correspond to a second metabolic activity level and a second tissue density at a second location of the second voxel.
[0074] In some example embodiments, the preprocessing engine 111 may adjust one or more initial dimensions of the second PET-CT scan 220 by at least the following: performing a second adjustment to a second dimension of the second PET-CT scan 220 and a third adjustment to a third dimension of the second PET-CT scan 220, based at least on a first adjustment to a first dimension of the second PET-CT scan 220. In some cases, the second adjustment to the second dimension of the second PET-CT scan 220 may include adding at least a second voxel to a second number of voxels along the second dimension of the second PET-CT scan 220 to increase the second number of voxels to a maximum number of voxels (e.g., 64, 128, 256, 512, etc.). Simultaneously, the third adjustment to the third dimension of the second PET-CT scan 220 may include adding at least a third voxel to a third number of voxels along the third dimension of the second PET-CT scan 220 to increase the third number of voxels to a maximum number of voxels (e.g., 64, 128, 256, 512, etc.). In some cases, the second voxel of the second size added to the second PET-CT scan 220 and the third voxel of the third size added to the second PET-CT scan 220 may each be assigned a first value (e.g., "0") corresponding to the level of metabolic activity and a second value (e.g., "-1024") corresponding to tissue density or X-ray attenuation.
[0075] As noted above, training segmentation model 113 to operate on the second PET-CT scan 220 can improve the performance of segmentation model 113, at least because the second PET-CT scan 220 can exhibit a higher signal-to-noise ratio (SNR) and background uniformity than the first PET-CT scan 210 from which the analysis controller 110 (e.g., preprocessing engine 111) extracts the second PET-CT scan 220. In some cases, this improvement in the performance of segmentation model 113 can be attributed at least in part to the exclusion of at least a second portion of the first PET-CT scan 210 occupied by voxels, which exhibit a different range of intensity values than the voxels in the first portion of the first PET-CT scan 210 that are extracted and resampled to form the second PET-CT scan 220.
[0076] Figure 5 depicts an example flowchart illustrating process 500 for identifying one or more positron emission tomography (PET) and computed tomography (PET-CT) scans to be included in a training dataset, according to some example embodiments. Referring to Figures 1, 2A-2B, 3, and 5, in some cases, process 500 may be executed by analysis controller 110 to identify one or more PET-CT scans to be preprocessed, for example, by preprocessing engine 111 to generate one or more corresponding training samples to be included in the training dataset for segmentation model 113. In some example embodiments, analysis engine 110 may execute process 500 to avoid preprocessing PET-CT scans that do not depict sufficiently large lesions in a specific region of the body.
[0077] At 502, the analysis controller 110 can determine a tumor mask that identifies a first plurality of voxels depicting one or more lesions in a first PET-CT scan. In some example embodiments, the analysis controller 110 can determine a tumor mask that identifies a first plurality of voxels depicting one or more lesions in a first PET-CT scan 210. In some cases, the first PET-CT scan 210 can depict multiple regions of the body. Therefore, the tumor mask can identify lesions present in multiple regions of the body.
[0078] At 504, the analysis controller 110 can determine the overlap between the tumor mask and one or more specific regions of the body. In some example embodiments, the analysis controller 110 can determine the overlap between the tumor mask and a first region of the body. In some cases, the first region of the body may contain one or more target organs (small intestine, large intestine, lung, thyroid gland, prostate, pancreas, cervix, etc.). Therefore, in some cases, the analysis controller 110 can determine the overlap between the tumor mask and an organ mask that identifies a second plurality of voxels depicting one or more target organs present in the first region of the body. As described in more detail below, if the first PET-CT scan 210 does not include one or more lesions in the first region of the body, the analysis controller 110 can avoid preprocessing the first PET-CT scan 210, for example, to generate a second PET-CT scan 220, to be included in the training dataset for the segmentation model 113.
[0079] At 505-N, the analysis controller 110 can determine that the overlap between the tumor mask and one or more specific regions of the body fails to meet a first threshold. Therefore, at 506, the analysis controller 110 can determine to omit the first PET-CT scan to avoid including it in the training dataset used for the segmentation model. For example, in some cases, if the first PET-CT scan 210 does not include any lesions whose overlap with a first region of the body (or one or more target organs in the first region of the body) meets a first threshold (e.g., 20%, etc.), the analysis controller 110 can avoid preprocessing the first PET-CT scan 220, for example, to generate a second PET-CT scan 220, to include it in the training dataset used for the segmentation model 113.
[0080] Alternatively, at 505-Y, the analysis controller 110 may determine that the overlap between the tumor mask and one or more specific regions of the body satisfies a first threshold. Therefore, at 507, the analysis controller 110 may determine whether the tumor mask satisfies a second threshold including at least one of tumor volume and distance to another lesion. In some cases, the analysis controller 110 may determine that the first PET-CT scan 210 includes at least one lesion whose overlap with a first region of the body (or one or more target organs within the first region of the body) satisfies a first threshold (e.g., 20%, etc.). Therefore, the analysis controller 110 may further verify whether at least one lesion present in the first region of the body (or one or more specific organs within one or more regions of the body) is an actual lesion or an artifact (e.g., associated with one or more imaging devices 120). For example, in some cases, upon determining that the overlap between the tumor mask and the first region of the body satisfies a first threshold (e.g., 20%, etc.), the analysis controller 110 may further determine whether the tumor mask satisfies a second threshold including at least one of tumor volume and distance to another lesion. If a lesion present in a first region of the body is below a threshold volume (e.g., 8 ml) and beyond a threshold distance from another lesion (e.g., 10 voxels), applying a second threshold ensures that the analysis controller 110 avoids preprocessing the first PET-CT scan 210. Small lesions far from another lesion are often artifacts rather than actual lesions. Therefore, if the first PET-CT scan 210 does not contain any lesions that are large enough or close enough to another lesion in the first region of the body to not be artifacts, the analysis controller 110 can avoid preprocessing the first PET-CT scan 210 to generate a second PET-CT scan 220 to be included in the training dataset for the segmentation model 113.
[0081] At 507-Y, the analysis controller 110 can determine that the tumor mask meets a second threshold. Therefore, at 508, the analysis controller 110 can determine, at least based on the first PET-CT scan, to generate a second PET-CT scan to be included in the training dataset for the segmentation model. In some example embodiments (where the analysis controller 110 determines that the tumor mask meets the second threshold), the analysis controller 110 can determine further preprocessing of the first PET-CT scan 210 to generate, for example, a second PET-CT scan 220 to be included in the training dataset for the segmentation model 113. Where the tumor mask associated with the first PET-CT scan 210 meets both the first and second thresholds, the analysis controller 110 can determine that the first PET-CT scan 210 not only depicts lesions in a first region of the body (or one or more specific organs in one or more regions of the body), but that these lesions are large enough or close enough to be not artifacts (e.g., associated with one or more imaging devices 120). In this case, the first PET-CT scan 210 can undergo preprocessing, for example by preprocessing engine 110, in which a first portion of the first PET-CT scan 210 depicting a first region of the body but not a second region of the body can be extracted and resampled to generate a second PET-CT scan 220.
[0082] Alternatively, at 507-N, the analysis controller 110 may determine that the tumor mask fails to meet a second threshold, including at least one of tumor volume and distance to another lesion. Therefore, process 500 may continue at operation 504, where the analysis controller 110 determines to omit the first PET-CT scan to avoid including it in the training dataset for the segmentation model. In some cases, even if the first PET-CT scan 210 depicts one or more lesions in a first region of the body (or one or more target organs within the first region of the body), these lesions may still be too small (e.g., less than 8 ml in volume) and / or too far from another lesion (e.g., more than 10 voxels) to be considered actual lesions. Therefore, if the tumor mask of the first PET-CT scan 210 meets the first threshold but not the second threshold, the analysis controller 110 may also avoid preprocessing the first PET-CT scan 210 to generate a second PET-CT scan 220 for inclusion in the training dataset for the segmentation model 113.
[0083] The performance of segmentation model 113 was evaluated using the Dice score to assess the accuracy of region-specific (e.g., organ-focused) methods relative to multi-region (e.g., whole-body) methods. At the lesion level, the performance of segmentation model 113 was evaluated by calculating precision, recall, and F1 score. Total metabolic volume (TMV) was also calculated for tumor masks predicted by implementing both region-specific (e.g., organ-focused) and multi-region (e.g., whole-body) methods using segmentation model 113, and these tumor masks were then compared to corresponding benchmark ground truth tumor masks to demonstrate that region-specific (e.g., organ-focused) methods produce tumor masks with greater relevance to benchmark ground truth tumor masks when calculating the metabolically active disease burden in patients with lymphoma. Spearman correlation coefficients were also calculated for each method to provide a statistical measure of the strength of the relationship between the predicted outcome and the benchmark ground truth in each case.
[0084] To calculate the precision, recall, and F1 score for the lesion level analysis performed by segmentation model 113, lesions identified in the baseline ground truth tumor mask are compared with lesions in the predicted tumor mask. A lesion in the predicted tumor mask is classified as a true positive (TP) if it has a threshold overlap (e.g., 20 percent) with a lesion in the baseline ground truth tumor mask. Conversely, a lesion in the predicted tumor mask is classified as a false positive (FP) if it exhibits less than a threshold overlap (e.g., 20 percent) compared to a lesion in the baseline ground truth tumor mask. False negatives (FNs) are lesions in the baseline ground truth tumor mask that have less than a threshold overlap (e.g., 20 percent) with a lesion in the predicted tumor mask. Equation (2) below shows the calculation of precision, recall, and F1 score based on the incidence of true positives (TP), false positives (FP), and false negatives (FN). It should be understood that a higher F1 score indicates a better balance between precision and recall.
[0085] The segmentation model 113 was applied to both the Goya and Gallium test sets. Quantitative results for the Goya test set are shown in Table 1, while those for the Gallium test set are shown in Table 2. Both the Dice score and the F1 score, a measure of lesion level, indicate that region-specific (e.g., organ-focused) methods outperform multi-regional (e.g., whole-body) methods.
[0086] Table 1 Table 2 Figures 7A and 7B depict mean Dice scores with standard deviations and confidence intervals for both region-specific (e.g., organ-focused) and multi-region (e.g., whole-body) methods. As shown in Figures 7A and 7B, the region-specific (e.g., organ-focused) method achieves statistically significant better Dice scores (e.g., in Goya) for both test sets. p <10 -5 And in Gallium for p <10 -3 Figures 7A and 7B also show that, compared to multi-region (e.g., whole-body) methods, region-specific (e.g., organ-focused) methods are able to produce intestinal tumor segmentation results with less variability and greater consistency (e.g., smaller standard deviation) across different cases.
[0087] For example, Figure 700 in Figure 7A includes a leftmost (vertical) error bar showing the dispersion of the average Dice score for the multi-region (e.g., whole-body) method. On the far right of Figure 700 is another (vertical) error bar showing the dispersion of the average Dice score for the region-specific (e.g., organ-focused) method. A horizontal line connecting the median Dice scores for the multi-region (e.g., whole-body) and region-specific (e.g., organ-focused) methods illustrates the region-specific (e.g., organ-focused) method. Figure 700 shows the region-specific (e.g., organ-focused) method as achieving not only a higher average Dice score with less overall variability, but also a higher median Dice score for the Goya test set compared to the multi-region (e.g., whole-body) method. Figure 725 compares the 95% confidence intervals for the Dice scores for the multi-region (e.g., whole-body) and region-specific (e.g., organ-focused) methods. As shown in Figure 725, region-specific (e.g., organ-focused) methods achieve significantly higher Dice scores within the 95% confidence interval of the Goya test set.
[0088] In Figure 7B, Figure 750 includes: a leftmost (vertical) error bar showing the dispersion of the average Dice score for the multi-region (e.g., whole-body) method; and a rightmost (vertical) error bar showing the dispersion of the average Dice score for the region-specific (e.g., organ-focused) method. A horizontal line connecting the median Dice scores for the multi-region (e.g., whole-body) and region-specific (e.g., organ-focused) methods illustrates the region-specific (e.g., organ-focused) method. Figure 750 shows the region-specific (e.g., organ-focused) method as achieving not only a higher average Dice score with less overall variability, but also a higher median Dice score for the Gallium test set compared to the multi-region (e.g., whole-body) method. Figure 750 compares the 95% confidence intervals for the Dice scores for the multi-region (e.g., whole-body) and region-specific (e.g., organ-focused) methods. As shown in Figure 750, region-specific (e.g., organ-focused) methods also achieve significantly higher Dice scores within the 95% confidence interval of the Gallium test set.
[0089] In the Gallium test set, region-specific (e.g., organ-focused) methods also demonstrated better compatibility with the benchmark ground truth tumor mask in terms of Dice scoring compared to multi-region (e.g., whole-body) methods (e.g., Table 2, 0.70 vs. 0.58). p <10 -3 A higher Dice score indicates better overall overlap and similarity between the predicted tumor mask and the baseline truth tumor mask. In the context of lesion segmentation, a higher Dice score is desirable because it indicates accurate capture of the true region while minimizing the incidence of false negatives.
[0090] For each of the predicted tumor masks generated by the two methods, the intestinal metabolic tumor volume was calculated and compared with the volume in the baseline true tumor mask. Figures 8A and 8B depict the comparison of the predicted total metabolic tumor volume with the corresponding baseline true value for patients with diffuse large B-cell lymphoma (DLBCL) and follicular lymphoma (FL) using region-specific (e.g., organ-focused) and multi-region (e.g., whole-body) methods. In Figure 8A for patients with diffuse large B-cell lymphoma (DLBCL), Figure 800 shows the relationship between the reported (baseline true) metabolic tumor volume and the predicted total metabolic tumor volume determined using the multi-region (e.g., whole-body) method, while Figure 825 shows the relationship between the reported (baseline true) metabolic tumor volume and the predicted total metabolic tumor volume determined using the region-specific (e.g., organ-focused) method. In Figure 8B for patients with follicular lymphoma (FL), Figure 850 shows the relationship between the reported (benchmark true) metabolic tumor volume and the predicted total metabolic tumor volume determined using a multi-region (e.g., whole-body) approach, while Figure 875 shows the relationship between the reported (benchmark true) metabolic tumor volume and the predicted total metabolic tumor volume determined using a region-specific (e.g., organ-focused) approach.
[0091] As shown in Figures 8A and 8B, region-specific (e.g., organ-focused) methods produce results for estimating metabolic tumor burden that correlate well with the baseline truth, while multi-region (e.g., whole-body) methods produce larger data dispersion. The Spearman correlations for region-specific (e.g., organ-focused) and multi-region (e.g., whole-body) methods are 0.86 and 0.93 for the Goya test set, and 0.77 and 0.88 for the Gallium test set, respectively. Furthermore, the multi-region (e.g., whole-body) and region-specific (e.g., organ-focused) methods produce coefficients of determination of 0.84 and 0.91, respectively, for the Goya test set. R 2 The Spearman correlation and coefficient of determination yielded coefficients of 0.50 and 0.89 for the Gallium test set, respectively. R 2 The values above indicate that, compared to multi-region (e.g., whole-body) methods, region-specific (e.g., organ-focused) methods are more consistent with the baseline truth.
[0092] Figure 9 illustrates another comparison of the first performance of a segmentation model 113 trained to perform lesion segmentation on region-specific PET-CT scans (such as the second PET-CT scan 220) and the second performance of a segmentation model 113 trained to perform lesion segmentation on multi-region PET-CT scans (such as the first PET-CT scan 210). In the example shown in Figure 9, the performance of segmentation model 113 is measured based on the Dice score, which is a quantitative similarity measure of voxel-level consistency between the predicted segmentation and the corresponding benchmark ground truth segmentation. When evaluated against a benchmark ground truth tumor mask, segmentation model 113 trained to operate on organ-focused PET-CT scans achieves a significantly higher Dice score than segmentation model 113 trained to operate on whole-body PET-CT scans.
[0093] Figure 10 depicts a block diagram illustrating an example of a computing system 1000 consistent with the implementation of the present topic. Referring to Figures 1 through 1010, the computing system 1000 may be used to implement an analysis controller 110, one or more imaging devices 120, a client device 130, and / or any of its components.
[0094] As shown in Figure 1010, the computing system 1000 may include a processor 1010, a memory 1020, a storage device 1030, and an input / output device 1040. The processor 1010, memory 1020, storage device 1030, and input / output device 1040 may be interconnected via a system bus 1050. The processor 1010 is capable of processing instructions for execution within the computing system 1000. Such execution instructions may implement one or more components, such as an analysis controller 110, one or more imaging devices 120, and a client device 130. In some exemplary embodiments, the processor 1010 may be a single-threaded processor. Alternatively, the processor 1010 may be a multi-threaded processor. The processor 1010 is capable of processing instructions stored on the memory 1020 and / or storage device 1030 to display graphical information for a user interface provided via the input / output device 1040.
[0095] Memory 1020 is a computer-readable medium, such as a volatile or non-volatile computer-readable medium, that stores information within computing system 1000. For example, memory 1020 may store data structures representing a configuration object database. Storage device 1030 provides persistent storage for computing system 1000. Storage device 1030 may be a solid-state drive, floppy disk drive, hard disk drive, optical disk drive, magnetic tape drive, or other suitable persistent storage medium. Input / output device 1040 provides input / output operations for computing system 1000. In some exemplary embodiments, input / output device 1040 includes a keyboard and / or a pointing device. In various specific embodiments, input / output device 1040 includes a display unit for displaying a graphical user interface.
[0096] According to some exemplary embodiments, input / output device 1040 may provide input / output operations for network devices. For example, input / output device 1040 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0097] In some exemplary embodiments, the computing system 1000 can be used to execute various interactive computer software applications that can be used to organize, analyze, and / or store data in various formats. Alternatively, the computing system 1000 can be used to execute any type of software application. These applications can be used to perform various functionalities, such as planning functionalities (e.g., generating, managing, and editing spreadsheet documents, word processing documents, and / or any other objects), computing functionalities, communication functionalities, etc. Applications may include various additional functionalities or may be standalone computing products and / or functionalities. Once activated within an application, the functionality can be used to generate a user interface provided via the input / output device 1040. The user interface can be generated by the computing system 1000 and presented to the user (e.g., on a computer screen monitor, etc.).
[0098] One or more aspects or features of the subject matter described herein can be implemented as digital electronic circuits, integrated circuits, specially designed ASICs, field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These aspects or features may be implemented in one or more computer programs that are executable and / or interpretable on a programmable system, which includes at least one programmable processor (which may be dedicated or general-purpose, coupled to receive and send data and instructions to it), a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Typically, clients and servers are remotely configured to each other and generally interact via a communication network. Relationships between clients and servers arise from computer programs running on their respective computers and the client-server relationships between them.
[0099] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, including machine instructions for a programmable processor, and may be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. Machine-readable media may (such as, for example, non-transitory solid-state memory or magnetic hard disk drive or any equivalent storage medium) non-transitory store such machine instructions. Machine-readable media may alternatively or additionally store such machine instructions transiently (such as, for example, a processor cache or other random access memory associated with one or more physical processor cores).
[0100] To provide interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or a light-emitting diode (LED) monitor for displaying information to the user) and a keyboard and pointing device (such as, for example, a mouse or trackball, through which the user can provide input to the computer). Other kinds of devices can also be used to provide interaction with the user. For example, the loop provided to the user can be any form of sensory loop, such as, for example, a visual loop, an auditory loop, or a tactile loop; input from the user can be received in any form, including sound, speech, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice identification hardware and software, optical scanners, optical indicators, digital image capture devices, and associated interpretation software, etc.
[0101] In the foregoing description and claims, phrases such as “at least one” or “one or more” may appear, followed by a list of combinations of elements or features. The term “and / or” may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to mean any element or feature listed alone, or any other recounted element or feature in combination with any other recounted element or feature. For example, the phrases “at least one of A and B”; “one or more of A and B”; and “A and / or B” are each intended to mean “A alone, B alone, or A and B together”. A similar interpretation applies to lists comprising three or more items. For example, the phrases “at least one of A, B, and C”; “one or more of A, B, and C”; and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together”. The use of the term "based on" in the above and claims is intended to mean "at least partially based on," so that undescribed features or elements are also permissible.
[0102] Depending on the desired construction, the subject matter described herein can be embodied in systems, apparatuses, methods, and / or articles of manufacture. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, other features and / or variations may be provided in addition to those features and / or variations set forth herein. For example, the above embodiments may be provided for various combinations and sub-combinations of the disclosed features and / or for combinations and sub-combinations of several further features disclosed above. Furthermore, the logical flows depicted in the drawings and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other embodiments may be within the scope of the following claims.
Claims
1. A computer-implemented method comprising: determining, within a first positron emission tomography and computed tomography (PET-CT) scan depicting a plurality of regions of a body, a first region including a first lesion; extracting, from the first PET-CT scan, a first portion of the first PET-CT scan that depicts the first region but does not depict a second region of the plurality of regions to generate a second PET-CT scan having one or more initial dimensions; adjusting the one or more initial dimensions of the second PET-CT scan to one or more target dimensions, the adjusting including adding one or more voxels to the second PET-CT scan and determining a value for each of the one or more voxels added to the second PET-CT scan; training a segmentation model based at least on a training dataset including the second PET-CT scan adjusted to the one or more target dimensions; and applying the trained segmentation model to identify one or more lesions present in a region-specific PET-CT scan that depicts the first region of the plurality of regions of the body but does not depict the second region.
2. The method of claim 1, further comprising: determining a first tumor mask that identifies a first plurality of voxels that depict one or more lesions present in the first PET-CT scan; and in response to determining that an overlap between the first tumor mask and the first region of the body satisfies a first threshold, generating the second PET-CT scan based at least on the first PET-CT scan by at least extracting the first portion of the first PET-CT scan.
3. The method of claim 2, further comprising: in response to determining that the first tumor mask satisfies a second threshold, generating the second PET-CT scan based at least on the first PET-CT scan.
4. The method of claim 3, wherein the second threshold includes at least one of a tumor volume and a distance to another lesion.
5. The method of any one of claims 2-4, further comprising: determining an organ mask that identifies a second plurality of voxels that depict a target organ present in the first region of the body; and in response to determining that an overlap between the first tumor mask and the organ mask satisfies the first threshold, generating the second PET-CT scan based at least on the first PET-CT scan.
6. The method of any one of claims 2-5, further comprising: determining a second tumor mask that identifies a second plurality of voxels that depict one or more lesions present in a multi-region PET-CT scan; and In response to determining that (i) an overlap between the second tumor mask and the first region of the body fails to satisfy the first threshold or (ii) the second tumor mask fails to satisfy a second threshold comprising at least one of a tumor volume and a distance to another lesion, generating the second PET-CT scan based on the first PET-CT scan rather than the multi-region PET-CT scan.
7. The method of any of claims 1-6, further comprising: generating the training data set to include first ground truth annotations that identify a first plurality of voxels that depict the first lesion in the second PET-CT scan.
8. The method of claim 7, further comprising: determining that a second lesion present in the second PET-CT scan fails to satisfy one or more thresholds comprising at least one of a tumor volume and a distance to another lesion; and in response to determining that the second lesion fails to satisfy the one or more thresholds, generating the training data set to exclude second ground truth annotations that identify a second plurality of voxels that depict the second lesion in the second PET-CT scan.
9. The method of any of claims 1-8, wherein the adjusting of the one or more initial dimensions of the second PET-CT scan comprises performing a first adjustment to a first dimension of the second PET-CT scan, and based at least on the first adjustment, performing a second adjustment to a second dimension of the second PET-CT scan and a third adjustment to a third dimension of the second PET-CT scan.
10. The method of claim 9, wherein the first adjustment to the first dimension of the second PET-CT scan comprises: adding at least a first voxel to a first number of voxels along the first dimension of the second PET-CT scan to increase the first number of voxels to a maximum number of voxels, and determining a first value for the first voxel.
11. The method of claim 10, wherein the first value for the first voxel is determined based at least on a second value for a second voxel within a threshold distance of the first voxel.
12. The method of claim 11, wherein the first value for the first voxel corresponds to a first metabolic activity level and a first tissue density at a first location of the first voxel, and wherein the second value for the second voxel corresponds to a second metabolic activity level and a second tissue density at a second location of the second voxel.
13. The method of claim 12, wherein the second adjustment to the second dimension of the second PET-CT scan comprises: adding at least a second voxel to a second number of voxels along the second dimension of the second PET-CT scan to increase the second number of voxels to the maximum number of voxels, and wherein the third adjustment to the third dimension of the second PET-CT scan comprises adding at least a third voxel to a third number of voxels along the third dimension of the second PET-CT scan to increase the third number of voxels to the maximum number of voxels.
13. The method of any of claims 1-12, wherein the one or more initial dimensions of the second PET-CT scan comprise a first dimension, a second dimension, and a third dimension.
14. The method of any of claims 1-13, wherein the first dimension of the second PET-CT scan is a slice dimension, the second dimension of the second PET-CT scan is a row dimension, and the third dimension of the second PET-CT scan is a column dimension.
15. The method of any of claims 1-14, wherein the first PET-CT scan is a first slice of a first row of a first column of a first volume of the first PET-CT scan, and wherein the second PET-CT scan is a second slice of a second row of a second column of a second volume of the second PET-CT scan.
16. The method of any of claims 1-15, wherein the first PET-CT scan is a first slice of a first row of a first column of a first volume of the first PET-CT scan, and wherein the second PET-CT scan is a second slice of a second row of a second column of a second volume of the second PET-CT scan.
17. The method of any of claims 1-16, wherein the first PET-CT scan is a first slice of a first row of a first column of a first volume of the first PET-CT scan, and wherein the second PET-CT scan is a second slice of a second row of a second column of a second volume of the second PET-CT scan.
14. The method of claim 13, wherein each of the second voxel and the third voxel is assigned a first value for a level of metabolic activity and a second value for tissue density or x-ray attenuation.
15. The method of claim 14, wherein the first value is 0 and the second value is -1024.
16. The method of any one of claims 9 to 15, wherein the maximum number of voxels is 64, 128, 256, or 512.
17. The method of any one of claims 1 to 16, wherein the first PET-CT scan is a whole body scan of the body, and wherein the second PET-CT scan is generated to depict some, but not all, of the plurality of regions depicted in the whole body scan of the body.
18. The method of any one of claims 1 to 17, wherein the plurality of regions of the body depicted in the first PET-CT scan comprises a plurality of organs in the body, and wherein the second PET-CT scan is generated to depict some, but not all, of the plurality of organs in the body.
19. The method of any one of claims 1 to 18, wherein the first region of the plurality of regions that constitutes the second PET-CT scan depicts at least a first organ, and wherein the second region of the plurality of regions that is removed from the first PET-CT scan depicts at least a second organ.
20. The method of any one of claims 1 to 19, wherein the extracting of the first portion of the first scan of the first region of the body of the first PET-CT scan comprises: identifying the plurality of voxels based at least on a range of intensity values exhibited by a plurality of voxels that constitute the first portion of the first PET-CT scan.
21. The method of any one of claims 1 to 20, wherein the extracting of the first portion of the first scan of the first region of the body of the first PET-CT scan comprises: identifying the plurality of voxels based at least on a difference between a maximum intensity value and a minimum intensity value exhibited by a plurality of voxels that constitute the first portion of the first PET-CT scan.
22. The method of any one of claims 1 to 21, wherein the extracting of the first portion of the first PET-CT scan that depicts the first region of the body comprises: determining a first range of intensity values exhibited by the first portion of the first PET-CT scan that excludes the voxel; determining a second range of intensity values exhibited by the first portion of the first PET-CT scan that includes the voxel; identifying the voxel to include in the first portion of the first PET-CT scan based at least on a difference between the first range of intensity values and the second range of intensity values satisfying one or more thresholds; and determining the voxel to exclude from the first portion of the first PET-CT scan based at least on the difference between the first range of values and the second range of values failing to satisfy the one or more thresholds.
23. The method of any one of claims 1 to 22, wherein the first region of the plurality of regions comprises one of small intestine, large intestine, lung, thyroid, prostate, pancreas, and cervix.
24. The method of any one of claims 1 to 23, wherein the segmentation model comprises an encoder and a decoder.
25. The method of claim 24, wherein the segmentation model further comprises one or more transform blocks that couple the output of the encoder to the input of the decoder.
26. The method of any one of claims 1 to 25, wherein the trained segmentation model identifies a plurality of voxels depicting the one or more lesions present in the first region of the body within the region-specific PET-CT scan.
27. The method of claim 26, further comprising: The metabolic tumor volume of the one or more lesions present in the first region of the body is determined based at least on the plurality of voxels.
28. The method of claim 27, further comprising: Based at least on the metabolic tumor volume, determine at least one of the following: response to disease stage, disease grade, treatment for the disease, disease progression, and disease burden.
29. The method according to any one of claims 1 to 28, wherein the first PET-CT scan and the second PET-CT scan are three-dimensional volumes.
30. The method according to any one of claims 1 to 29, wherein the first PET-CT scan and the second PET-CT scan each comprise a series of two-dimensional slices.
31. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 1 to 30.
32. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 1 to 30.