Automated tumor identification and segmentation from medical images

The multi-stage neural network approach for tumor detection and segmentation in medical images addresses the limitations of RECIST by providing a comprehensive and objective assessment of tumor burden and progression, enhancing the accuracy of disease evaluation.

JP2026027229APending Publication Date: 2026-02-18GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025169834
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-16
Filing Date
2025-10-08
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Existing tumor detection methods, such as RECIST, are limited in assessing overall disease burden and suffer from intra- and inter-reader variability due to inconsistent lesion selection and heterogeneous tumor appearances, particularly in metastatic cancers with multiple lesions.

Method used

An automated method using multi-stage neural networks for tumor detection and segmentation in medical images, including bounding box detection and organ-specific segmentation, to provide comprehensive tumor assessment.

Benefits of technology

Accurately tracks tumor progression and treatment efficacy by identifying and segmenting multiple lesions across different organs, reducing variability and providing objective, consistent disease burden evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026027229000001_ABST
    Figure 2026027229000001_ABST
Patent Text Reader

Abstract

To provide an automated method of tumor detection and measurement that accounts for the overall disease burden of a subject.SOLUTION: The medical image is input to the detection network to generate a mask identifying a set of regions in the medical image, and the detection network predicts that each region identified in the mask includes a depiction of a tumor of the one or more tumors in the subject. For each region, the region of the medical image is processed using a tumor segmentation network to generate one or more tumor segmentation boundaries of a tumor present in the subject. For each tumor and by using a plurality of organ-specific segmentation networks, an organ in which at least a part of the tumor is located is determined. The output is generated based on the one or more tumor segmentation boundaries and a location of an organ in which at least a portion of the one or more tumors is located.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 62 / 952,008, filed December 20, 2019, and U.S. Provisional Patent Application No. 62 / 990,348, filed March 16, 2020, each of which is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] Medical imaging (e.g., CT scans, X-rays, or MRI scans) is widely used for tumor detection to aid in the diagnosis and treatment of cancer (e.g., lung cancer, breast cancer, etc.). Medical professionals often evaluate the effectiveness of drugs and / or treatment regimens by measuring changes in tumor size or volume. The Response Evaluation Criteria in Solid Tumors (RECIST) is a standardized method for assessing treatment response in cancer subjects and is part of the regulatory criteria for the approval of new oncology drugs. RECIST requires a significant amount of time from trained experts (e.g., radiologists). Specifically, the annotator manually (e.g., by a radiologist) identifies up to five target lesions and up to ten non-target lesions. The annotator identifies the perimeter of each target lesion in each scan, where a cross-section of the target lesion is depicted, and records the cross-sectional diameter of each target lesion. A quantitative metric (e.g., the sum of the longest diameters) is then determined for all target lesions. Non-target lesions are qualitatively evaluated to indicate whether they are observed on the scan and whether there is any clear change. Scans can be collected at multiple time points, and metrics for target and non-target lesions can be determined for each time point. Changes in the metrics over time can then be used to assess the extent to which the disease is progressing and / or being effectively treated.

[0003] However, RECIST has several limitations. Because RECIST frequently measures only a small subset of tumors (e.g., fewer than 5–10) for each subject, the method does not account for the overall disease burden. Given that only up to five tumors are tracked, this technique cannot accurately assess disease progression and / or treatment efficacy in subjects with cancers that have metastasized to include multiple lesions (e.g., more than five lesions). Furthermore, variability in lesion selection can lead to inconsistencies in target lesion selection, which can cause significant intra- and inter-reader variability, leading to different assessments of tumor burden even within the same subject. For example, different sets of lesions may be identified (e.g., inadvertently) across different time points. Furthermore, many tumors often have a heterogeneous appearance on CT, which can vary by location, size, and shape. For example, lung lesions can be cavitary or calcified, and bone metastases can be (for example) lytic (destroying skeletal tissue) or blastic (abnormal bone growth), and each lesion type is associated with a different structural and visual appearance, making it difficult to assess the disease stage and / or type of each lesion without obtaining a complete readout. Therefore, it would be advantageous to identify automated techniques that assess tumor growth and / or metastasis using more comprehensive data sets and more objective techniques.

[0004] The present disclosure attempts to address at least the above limitations by providing an automated method of tumor detection and measurement that is consistent and accounts for a subject's overall disease burden. Summary of the Invention

[0005] The technology described herein discloses methods for identifying and segmenting biological objects using one or more medical images.

[0006] In some embodiments, a computer-implemented method is provided for accessing at least one or more medical images of a subject, the one or more medical images comprising: The detection network generates one or more masks that identify a set of regions in the one or more images. The detection network predicts that each region of the set of regions identified in the one or more masks contains a depiction of a tumor in the subject. For each region of the set of regions, the region of the one or more medical images is processed using the tumor segmentation network to generate one or more tumor segmentation boundaries of the tumor present in the subject. An organ in which at least a portion of the tumor is located is determined for each tumor of the one or more tumors by using multiple organ-specific segmentation networks. An output is then generated based on the one or more tumor segmentation boundaries and the organ location.

[0007] In some embodiments, another computer-implemented method is provided for accessing one or more medical images of a subject. A set of organ locations for a set of tumor lesions present in the one or more medical images is also accessed. The one or more medical images and the set of organ locations are input into a network associated with one of a plurality of therapeutic treatments to generate a score representing whether the subject is a good candidate for a particular treatment compared to other previous subjects who have received the treatment. The score is then returned for evaluation of the subject's survival rate and each of the plurality of treatments.

[0008] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0009] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein. [Brief explanation of the drawings]

[0010] The present disclosure is described in conjunction with the accompanying drawings, in which:

[0011] [Figure 1A] 1 illustrates an exemplary interactive system for using, acquiring, and processing medical images using a multi-stage neural network platform.

[0012] [Figure 1B] 1 shows an example image stack containing a set of patches and bounding boxes of detected biological objects.

[0013] [Figure 2] 1 illustrates an exemplary system for generating one or more pairwise comparisons between two or more subjects.

[0014] [Figure 3] 1 illustrates an exemplary method for processing medical images using a multi-stage neural network platform.

[0015] [Figure 4] 1 shows an example set of images for tumor detection: the leftmost panel shows a whole-body scan of axial slices after preprocessing, and the right panel shows detected bounding boxes, automatically generated and labeled by a bounding box detection network, for the lung, liver, and mediastinal regions in the axial slices.

[0016] [Figure 5] Figure 1 shows an example of tumor segmentation using axial CT scans. Each of the upper panels shows the determined area of ​​the tumor. The corresponding lower panels show an example segmentation boundary of the tumor.

[0017] [Figures 6A-6B]Plots comparing manual assessment using RECIST with the automated method on an exemplary training set are shown: Panel A: Comparison for several identified lesions, Panel B: Comparison for the sum of longest diameters (SLD) determined.

[0018] [Figures 7A-7B] Figure 1 shows plots comparing manual assessment using RECIST with automated methods for an exemplary test set. Panel A: Comparison for some identified lesions, Panel B: Comparison for determined SLD.

[0019] [Figure 8] 1 shows a plot comparing the number of lesions identified using a full reading performed by a radiologist with the number of lesions identified using the automated method for an exemplary training set.

[0020] [Figure 9] 1 shows a plot comparing the volume of lesions identified using a full reading performed by one or more radiologists with the volume of lesions identified using the automated method for an exemplary training set.

[0021] [Figures 10A-10B] Figure 1 shows plots comparing the mean and median volumes of lesions identified using full interpretation with the volumes of lesions identified using the automated method for an exemplary training set. Panel A: Mean volume data. Panel B: Median volume data.

[0022] [Figures 11A-11C]Kaplan-Meier curves are shown for an exemplary training set. Panel A: SLD derived by manually assessed RECIST, divided into quartiles based on the derived SLD. Panel B: Number of lesions derived by manually assessed RECIST, divided into quartiles based on the number of lesions. Panel C: Total SLD derived by the automated method, divided into quartiles based on the derived total SLD.

[0023] [Figures 12A-12B] Kaplan-Meier curves for an exemplary training set are shown. Panel A: Total volume derived by the automated method, divided by quartiles. Panel B: Number of lesions derived by the automated method, divided by quartiles.

[0024] [Figures 13A-13B] Kaplan-Meier curves using lesions located within the lung region for an exemplary training set are shown. Panel A: Volume of lung lesions derived by the automated method, divided by quartiles. Panel B: Number of lung lesions derived by the automated method, divided by quartiles.

[0025] [Figures 14A-14B] Kaplan-Meier curves for an exemplary training set are shown. Panel A: Measure of liver involvement derived by the automated method, divided by quartiles. Panel B: Measure of bone involvement derived by the automated method, divided by quartiles.

[0026] [Figures 15A-15B] Kaplan-Meier curves for an exemplary validation set are shown. Panel A: SLD derived by manually assessed RECIST, divided by quartiles. Panel B: Number of lesions derived by manually assessed RECIST, divided by quartiles.

[0027] [Figures 16A-16C] Kaplan-Meier curves are shown for an exemplary validation set. Panel A: SLD derived by manually assessed RECIST, divided by quartiles; Panel B: total SLD derived by the automated method, divided by quartiles; Panel C: total volume derived by the automated method, divided by quartiles.

[0028] [Figures 17A-17B] Kaplan-Meier curves for an exemplary validation set are shown. Panel A: Total tumor volume derived by the automated method, divided by quartiles. Panel B: Number of lesions derived by the automated method, divided by quartiles.

[0029] [Figures 18A-18B] Kaplan-Meier curves using lesions located within the lung region for an exemplary validation set are shown. Panel A: Volume of lung lesions derived by the automated method, divided by quartiles. Panel B: Number of lung lesions derived by the automated method, divided by quartiles.

[0030] [Figure 19] Kaplan-Meier curves for measures of renal involvement derived by the automated method for the exemplary validation set are shown. Data for renal involvement were divided by quartiles.

[0031] [Figure 20]This figure shows examples of tumor detection and segmentation from axial CT scans using the automated detection and segmentation method. The upper left panel shows three lesions detected in the liver, with the associated lesion segmentations in the lower plots. Similarly, the upper right panel shows four lesions detected in the lung / mediastinum, along with their associated segmentations. The two examples in the lower panel show detected lesions in the kidney and lung space, respectively.

[0032] [Figure 21] Row by row, we show segmentation examples from left to right: radiologist annotations, problem UNetβ=10, problem UNetβ==2, combined tumor segmentation network implemented as a probabilistic UNet.

[0033] [Figures 22A-22B] Kaplan-Meier curves are shown for another exemplary study set: Panel A: SLD derived by manually assessed RECIST, divided by quartiles; Panel B: SLD by the automated method, divided by quartiles.

[0034] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. When only a first reference label is used in this specification, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION

[0035] I. Overview Recent image analysis efforts have focused on developing automated algorithms that can assist radiologists' workflow by detecting and segmenting tumors. Recent methods focus on detecting and / or segmenting RECIST lesions in single axial CT sections. These recent efforts have been limited to segmenting tumors only on a single section or in a single organ (e.g., in the lung) for tumor screening, as opposed to advanced-stage subjects, which suffer from higher and more variable tumor burdens.

[0036] As described herein, the techniques are used to analyze one or more image scans of a subject (e.g., CT or MRI scans). Each image scan can include a set of images corresponding to different slices (e.g., different axial slices). A first neural network can be used to detect regions containing depictions of a particular type of biological object (e.g., a tumor) for each image in the image scan. The first neural network (i.e., a bounding box detection neural network) can include a convolutional neural network such as RetinaNet and / or a three-dimensional neural network. The first neural network can be configured to define each region as a bounding box containing the depicted biological object and, optionally, a predetermined size of padding (e.g., the width of the box is defined to be the estimated maximum width of the biological object depiction plus two times the padding). The first neural network can process the image scans by individual focus (e.g., to define regions for each individual image), but can also be configured to use a scan depicting an upper slice of each scan and another scan depicting a lower slice of each scan to provide context.

[0037] A second neural network (e.g., a segmentation neural network) can be configured to process smaller portions of the image scan to segment individual objects. More specifically, one or more cropped portions of the image processed by the first neural network can be input to the second neural network. Each cropped portion can correspond to a bounding box defined for a particular image. The cropped portion can have (for example) an area equal to the area of ​​the bounding box or the area of ​​the bounding box plus padding. The second neural network may be configured to receive corresponding portions from other images representing adjacent slices. The second neural network can include a convolutional and / or three-dimensional neural network, such as UNet. The output of the second neural network can identify, for each box, a set of pixels estimated to correspond to the circumference or area of ​​a cross-section of the object cross-section depicted in the image.

[0038] In some cases, the object segmentations are aligned and / or smoothed across the images, and then three-dimensional representations of the individual objects can be obtained.

[0039] A neural network (e.g., the first neural network, the second neural network, or another neural network) can be configured to estimate the environment of an object. For example, the network can output a probability that a biological object is within the subject's lung, liver, bone, mediastinum, or other location. The probabilities can be evaluated independently (e.g., in which case the probabilities do not need to sum to 1 across various probabilities). Predicting the context can facilitate segmentation, registration, and / or other processing. For example, a particular type of biological object (e.g., a tumor) may generally have different characteristics in different environments. Thus, the environment prediction can inform what type of image features are used to generate object segmentation and / or perform other image processing. In some cases, the network outputs an estimated probability of an image truly depicting a particular type of object.

[0040] In some cases, the third neural network can determine the environment of the biological object by performing a second segmentation of the location of interest in the image. For example, the third neural network can output a segmentation of the lungs, liver, kidneys, and / or another location corresponding to the subject (e.g., in the form of a two-dimensional and / or three-dimensional mask). In some cases, the third neural network can be trained to segment a single location of interest, and additional neural networks can be configured to segment additional locations of interest. For example, the third neural network can output a segmentation of the lungs, the fourth neural network can output a segmentation of the liver, and the fifth neural network can output a segmentation of the kidneys.

[0041] Using either two-dimensional or three-dimensional segmentation, one or more object-specific statistics can be generated to characterize each estimated object representation. The one or more object-specific statistics can include (for example) area, longest dimension, or circumference. One or more scan-specific statistics can be generated for each scan. The scan-specific statistics can include (for example) the number of objects detected per scan, a statistic based on the number of objects detected per scan (e.g., mean, median, or maximum), a statistic based on an object-specific statistic (e.g., mean, median, or maximum), or a statistic based on the amount of objects detected across each scan (e.g., mean, median, or maximum). Additional object-level statistics can be generated, such as (for example) the total number of objects detected across all scans (e.g., associated with a given object), the sum of the longest dimension of objects detected across all scans, and / or the cumulative amount of objects detected across all scans.

[0042] Scan-specific, object-specific, and / or subject-level statistics can be output. In some cases, the statistics can be stored in association with a time point and a subject identifier. The statistics can then be tracked and compared over time to estimate the extent to which a medical condition is progressing, the effectiveness of a given treatment, and / or the prognosis for a given subject. II. Definition

[0043] As used herein, "medical image" refers to an image of the inside of a subject's body. Medical images can include CT, MRI, and / or X-ray images. Medical images can depict a portion of a subject's tissue, organ, and / or entire anatomical region. Medical images can depict a portion of a subject's torso, chest, abdomen, and / or pelvis. Medical images can depict a subject's entire body. Medical images can include two-dimensional images.

[0044] As used herein, "whole-body imaging" refers to collecting a set of images that collectively depict the entire body of a subject. The set of images can include images associated with a virtual "slice" extending from a first end (e.g., anterior end) to a second end (e.g., posterior end) of the subject. The set of images can include virtual slices of at least the brain, thoracic, abdominal, and pelvic regions of the subject.

[0045] As used herein, an "image stack" refers to a set of images depicting a set of adjacent virtual slices. Thus, the set of images can be associated with different depths (for example). An image stack can include at least two images or at least three images (for example). The image stack can include a bottom image, a middle image, and a top image, where the depth associated with the middle image is between the depths of the bottom image and the top image. The bottom image and the top image can be used to provide context information related to the processing of the middle image.

[0046] As used herein, a "biological object" (e.g., also referred to as an "object") refers to a biological structure and / or one or more regions of interest associated with a biological structure. Exemplary biological structures can include one or more biological cells, organs, and / or tissues of a subject. An object can include, but is not limited to, any of these identified biological structures and / or similar structures within or connected to the identified biological structure (e.g., multiple tumor cells and / or tissues identified within normal cells, organs, and / or tissues of a subject's larger body).

[0047] As used herein, a "mask" refers to an image or other data file that represents the surface area of ​​a detected object or other region of interest. A mask may include non-zero intensity pixels that indicate one or more regions of interest (e.g., one or more detected objects) and zero intensity pixels that indicate the background.

[0048] As used herein, a "binary mask" refers to a mask in which each pixel value is set to one of two values ​​(e.g., 0 or 1). An intensity value of zero can indicate that the corresponding pixel is part of the background, and a non-zero intensity value (e.g., a value of 1) can indicate that the corresponding pixel is part of the region of interest.

[0049] As used herein, a "3D mask" refers to the complete surface area of ​​an object in a three-dimensional image. Multiple binary masks of an object can be combined to form a 3D mask. The 3D mask can further provide information about the volume, density, and location in space of the object or other region of interest.

[0050] As used herein, "segmentation" refers to determining the location and shape of an object or region of interest within an image (two-dimensional or three-dimensional) or other data file. Segmentation can include determining a set of pixels that describe the region or perimeter of an object within an image. Segmentation can include generating a binary mask of the object. Segmentation can further include processing multiple binary masks corresponding to the object to generate a 3D mask of the object.

[0051] As used herein, "segmentation boundary" refers to an estimated perimeter of an object in an image. The segmentation boundary may be generated during a segmentation process in which image features are analyzed to determine the location of the object's edges. The segmentation boundary may also be represented by a binary mask. As used herein, "treatment" refers to prescribing or applying a therapy, drug, and / or radiation, and / or prescribing or performing a surgical procedure for the purpose of treating a medical condition (e.g., to slow the progression of a medical condition, to halt the progression of a medical condition, to reduce the severity and / or extent of a medical condition, and / or to cure a medical condition). III. Exemplary Dialogue System

[0052] 1A illustrates an exemplary interactive system for using, acquiring, and processing medical images using a multi-stage neural network platform. In this particular example, the interactive system is specifically configured to identify and segment depictions of tumor biological structures and organs within medical images. A. Input Data

[0053] One or more imaging systems 101 (e.g., CT, MRI, and / or X-ray devices) can be used to generate one or more sets of medical images 102 (e.g., CT, MRI, and / or X-ray images). The imaging system 101 can be configured to iteratively adjust focus and / or position as multiple images are collected, such that each image in the set of images is associated with a different depth, position, and / or viewpoint relative to the other images in the set. The imaging system 201 can include a light source (e.g., a motorized and / or X-ray source), a photodetector (e.g., a camera), lenses, an objective lens, filters, a magnet, shim coils (e.g., to correct for magnetic field inhomogeneities), a gradient system (e.g., to localize magnetic resonance signals), and / or an RF system (e.g., to excite a sample and detect the resulting nuclear magnetic resonance signals).

[0054] Each set of images 102 can correspond to an imaging session, a session date, and a subject. The subject can include a human or animal subject. The subject may have been diagnosed with a particular disease (e.g., cancer) and / or may have one or more tumors.

[0055] Each set of images 102 may depict the interior of a corresponding object. In some cases, each image depicts at least a region of interest of the object (e.g., one or more organs, a chest region, an abdominal region, and / or a pelvic region).

[0056] Each image in the set of images 102 may further have the same viewing angle, such that each image depicts a plane parallel to other planes depicted in the other images in the set. In some cases, each image in the set may correspond to a different distance along an axis non-parallel (e.g., perpendicular) to the plane. For example, the set of images 102 may correspond to a set of horizontal virtual slices corresponding to different positions along the anterior-posterior axis of the subject. The set of images 102 may be preprocessed (e.g., collectively or individually). For example, preprocessing may include normalizing pixel intensity, aligning images to each other or to another reference point / image, cropping images to a uniform size, and / or adjusting contrast to distinguish between light and dark pixels. In some cases, the set of images 102 may be processed to generate a three-dimensional (3D) image structure. The 3D image structure may then be used to generate another set of images corresponding to a different angle of the virtual slice. B. Training Data

[0057] Some medical images collected by at least one of the imaging systems 101 may be included in a training dataset to include training images for training one or more neural networks (e.g., a bounding box detection network and a segmentation network). The training images may be associated with other subjects compared to the subjects for which the trained networks are used.

[0058] Each training image can have one or more characteristics of the medical images 102 described herein and can be associated with annotation data that indicates whether and / or where the image depicts a tumor and / or organ. To identify this annotation data, images collected by the imaging system 101 can be utilized (e.g., transmitted) to the annotation device 103.

[0059] The image may be presented on the annotation device 103, and the annotation user (e.g., a radiologist, etc.) may provide input using (e.g., a mouse, trackpad, stylus, and / or keyboard indicating (e.g., whether the image depicts any tumors (or one or more specific types of organs); the number of tumors shown in the image; the number of tumors annotated (e.g., outlined) by the annotator; and the perimeter of each of the one or more tumors and / or one or more specific types of organs.

[0060] The annotation device 103 can convert the input into (for example) label data 104. Each label dataset can be associated with a corresponding image dataset. The label data 104 can indicate whether the image includes a tumor and / or one or more specific types of organs. The label data 104 can further indicate the location of tumors and / or organs located within the image by identifying spatial features (e.g., perimeters and / or regions) of the tumors and / or organs. For example, the label data 104 can include sets of coordinates identifying coordinates associated with the perimeters of each of a set of depicted tumors. As another example, the label data 104 can include instructions regarding which pixels (or voxels) in the training images correspond to the perimeters and / or regions of the depicted tumors.

[0061] Spatial features can be further identified for multiple objects. In some cases, the label data 104 may (but need not) identify spatial features of all tumors, organs, and / or other biological objects depicted in the training images. For example, if the training images depict 10 tumors, the label data 104 may identify the perimeter for each of the 10 tumors or for only two of the depicted tumors. In such cases, an incomplete subset of objects may (but need not) be selected based on predetermined selection criteria. For example, the annotation user may be instructed to mark only depictions of tumors that meet a threshold tumor length and / or threshold tumor volume and / or are within a region of interest (e.g., within one or more specific organs).

[0062] The label data 104 may further identify a tumor classification, which may represent the type, location, and / or size of the tumor identified based on input from the annotator. For example, a particular label may indicate that the depicted tumor is within a region of the image 102 corresponding to a particular organ (e.g., the liver). The label data 104 may further include a probability that a particular label actually corresponds to the tumor or organ of interest. The probability value may be calculated based on tumor length, tumor volume, location relative to the subject, and / or the number of annotator users who identify a particular label as corresponding to a tumor or organ. The label data 104 can be used to train one or more neural networks to detect regions containing depictions of tumors or organs for each image in the image scan. The trained neural networks can be configured to delineate each region identified as containing the indicated tumor or organ by processing the image scans by individual focus points using the image stack corresponding to each respective scan (e.g., to define specific regions for each individual image). C. Bounding Box Detection Network

[0063] The neural network processing system 120 can be configured to receive one or more sets of images 102 and corresponding label data 104. Each image in the one or more sets of images can first be preprocessed by the preprocessing controller 105. For example, one or more images showing different regions of an object can be stitched together to generate an aggregate image showing all of the different regions. In some cases, the aggregate image depicts a “full-body” view of the object. As another example, one or more images can be scaled and / or cropped to a predetermined size. In yet another example, one or more images can be aligned to another image or a reference image included in the set (e.g., using alignment markings in the image, correlation-based techniques, or entropy-based techniques). In another example, pixel intensities of one or more images can be adjusted by a normalization or standardization method. In some cases, the set of images 102 does not undergo preprocessing techniques.

[0064] The preprocessed images may be utilized by the bounding box detection controller 106, which may control and / or perform all of the functions and operations of the bounding box detection network, as described herein. The bounding box detection network may be a convolutional neural network, a deconvolutional neural network, or a three-dimensional neural network configured to identify regions (e.g., bounding boxes) within the set of images 102 that contain depictions of tumors. The regions identified by the bounding box detection neural network may include one or more rectangular or hyper-rectangular regions.

[0065] The bounding box detection controller 106 can use the training images and corresponding annotations to train a bounding box detection network to learn a set of detection parameters 107. The detection parameters 107 can include weights between nodes in the convolutional network. A penalty function can be configured to introduce a penalty if any portion of the detected bounding box does not completely contain a depiction of the tumor and / or if the padding between the additional horizontal and / or vertical points is less than a lower threshold and / or greater than an upper threshold. In some cases, the penalty function is configured to penalize bounding boxes that are larger or smaller than a predetermined zoom range. The penalty function can include a focus loss. Focal loss (as defined in Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollar, P. "Focal loss for dense object detection." ICCV 2017, pp. 2980-2988 (2017), the entirety of which is incorporated by reference herein for all purposes) can be used to address class imbalance as well as to "refocus" training of the detection task towards difficult-to-predict cases due to tag-perceptual variability in tumors.

[0066] The bounding box detection network can be trained and / or defined using one or more fixed hyperparameters, such as a learning rate, a number of nodes per layer, a number of layers, etc.

[0067] The bounding box detection network can detect one or more bounding boxes 108 corresponding to potential tumor delineations in each image 102. Detecting the bounding boxes can include using the image stack for each image to locate the bounding boxes. For example, if 100 images were collected during a particular imaging session (numbered sequentially according to imaging depth), detecting a bounding box in the seventh image can define an image stack to include the sixth, seventh, and eighth images. The image stack can include two or more adjacent images in one or more directions (e.g., including the third through eleventh images when detecting a bounding box in the seventh image).

[0068] The features of the image stack are used to provide contextual information when determining whether and / or where one or more regions contain a tumor and / or organ. The features can include three-dimensional features that extend across the images in the image stack. For example, if a feature (e.g., a learned feature) is present in similar locations throughout the image stack (e.g., a combination of the top virtual slice, the bottom virtual slice, and the center virtual slice), the bounding box detection network can determine that the image region corresponding to (e.g., containing) the feature represents a tumor bounding box. Alternatively, if a feature in the center slice of the image stack is not present in either the top or bottom slices of the image stack, the bounding box detection network can determine that the image region corresponding to the feature corresponds to the image background (i.e., any biological structure other than the tumor) and does not represent a bounding box. In some cases, the bounding box detection network can further assign a probability value to each detected bounding box. If the probability value of a bounding box does not exceed a threshold, the bounding box may be discarded as background.

[0069] The bounding box detection network can further process each detected bounding box 108 so that the margins of the bounding box include at least a certain amount of padding (e.g., 10px, 15px, or another suitable amount) from each edge of the region corresponding to the tumor. In some cases, the amount of padding is predefined (e.g., by generating an initial box that intersects the pixels furthest to the left, top, right, and bottom of the detected object depiction and expanding the box using a predetermined padding or until an image boundary is encountered). In other examples, varying degrees of padding are added to maintain a uniform bounding box size.

[0070] The bounding box data 108 may include a definition of each bounding box (e.g., as two or more corner coordinates, one or more edge coordinates, etc.) and / or one or more identifiers of the corresponding image or set of images (e.g., image identifier, subject, capture date, etc.).

[0071] It will be appreciated that the location of a bounding box in one image can be related to the location of a bounding box in another image. An image stack may be used to convey this dependency, although other processes may also or alternatively be used. For example, the input to a bounding box detection neural network may include the identification of each of one or more bounding boxes detected from previously processed images (corresponding to the same imaging session and the same subject). As another example, the bounding box output may be post-processed to modify (e.g., translate, resize, remove, or add) the bounding box detection corresponding to one image based on the bounding box detections from one or more other adjacent images.

[0072] FIG. 1B illustrates an exemplary image stack depicting a set of bounding boxes for a single biological object 125. The image stack can include at least image 121, image 122, and image 123, with each image in the image stack showing a different axial perspective of the region of interest. In some cases, the image stack can include additional images not shown. Each image in the image stack can further include a bounding box depicting a possible location of the biological object 125 in that particular image. As a result, each bounding box can be related to corresponding bounding boxes included in other images in the image stack, as each bounding box identifies the presence of the same biological object 125. For example, image 121 includes a bounding box 121A covering at least a portion of image 121, and image 122 includes a bounding box 122A covering at least a corresponding portion of image 122. As a result, bounding box 121A and bounding box 122A are related bounding boxes that contain regions predicted to depict first and second possible locations of the biological object 125 from a first and second axial perspective, respectively. In other examples, the biological object 125 may not be detected in at least a subset (e.g., one or more) of the images in the image stack, and therefore the subset of images in the image stack may not include an associated bounding box of the biological object 125.

[0073] Additionally, there may be differences in the exact location (e.g., represented by a set of coordinates), surface area, and / or shape of the associated bounding boxes within the image stack. In this example, the surface area of ​​bounding box 121A may be smaller than the surface area of ​​bounding box 122A because it is estimated that the majority of the biological object 125 is located within image 122. Each location of the associated bounding boxes may further include one or more transformations (e.g., in the x-plane, y-plane, or both) that describe the corresponding location of the same biological object 125 from one or more different axial perspectives of the images in the image stack.

[0074] In some cases, in response to identifying a set of related bounding boxes for an image stack, a detection region is determined for each of the related bounding boxes. For example, image 121 may include detection region 121B surrounding bounding box 121A. The detection region may be the same size and in the same location for each image in the image stack. In some embodiments, the size and location of the detection region may be determined from the location of the bounding box in a central slice of the image stack (e.g., image 122 in this case). The detection region may be configured to include the entirety of each of the identified bounding boxes with additional padding. In some cases, the detection region may be determined by a separate neural network separate from the bounding box detection network. D. Tumor segmentation network

[0075] Referring back to FIG. 1A , the bounding box data 108 can be transmitted to a tumor segmentation controller 109, which can control and / or perform all of the functions or operations of the tumor segmentation network as described herein. The tumor segmentation network can be trained using a training dataset of at least predicted bounding box data determined during training of the bounding box detection network. A set of segmentation parameters 110 (e.g., weights) can be learned during training. In the illustrated example, the tumor segmentation network can be a neural convolutional neural network or a three-dimensional neural network configured to detect and segment tumor depictions (for example). In some cases, the tumor segmentation network does not include a neural network and may instead use clustering techniques (e.g., K-means techniques), histogram-based techniques, edge detection techniques, region growing techniques, and / or graph partitioning techniques (for example). The tumor segmentation network can be configured to segment the tumor within each detected bounding box 108.

[0076] For each medical image in the set of images 102, the bounding box 108 includes (for example) one or more portions of the image corresponding to the bounding box, or the entire image, along with an identification of the bounding box (e.g., vertex coordinates and / or edge coordinates) associated with the respective image. In some embodiments, intermediate processing (not shown) may be performed to generate a set of cropped images (e.g., referred to herein as detection regions) corresponding to only the regions of the images 102 enclosed by the bounding boxes 108. If multiple bounding boxes are defined for a given image, the tumor segmentation network may receive as input each corresponding detection region and process the detection regions separately.

[0077] The detection region can provide a focused view of the target tumor, as shown in FIG. 1B. In some cases, the detection region may be a predetermined size. In such cases, the detection region may include another set of regions adjacent to the region corresponding to the bounding box as additional padding to maintain the predetermined size of the detection region. In other cases, if the bounding box is larger than a predetermined size (e.g., 400 pixels or 200 pixels by 200 pixels), the region corresponding to the bounding box is divided into two or more windows (e.g., of a predetermined size and / or of a predetermined size or smaller) so that each window corresponds to a separate detection region. In such cases, the detection region corresponding to a single bounding box may include overlapping portions of the image.

[0078] If the bounding box spans the entire image stack (as shown in FIG. 1B), a separate detection region can be defined for each image in the image stack. In some embodiments, the detection region processing is performed by a bounding box detection network before sending the bounding box data 108 to the tumor segmentation controller 109.

[0079] The tumor segmentation controller 109 implements a tumor segmentation network configured to further identify and evaluate features (e.g., pixel intensity variation) of each detected region to identify a set of perimeters, edges, and / or contours corresponding to the tumor. The features identified by the tumor segmentation network may have similarities to and / or may be different from those identified by the bounding box detection network. Both networks can be trained to identify regions of the image corresponding to tumors, although different features may be useful for detecting relatively small structures compared to relatively large structures. In some cases, the tumor segmentation network can learn to detect the location of objects by analyzing (for example) pixel intensity, pixel color, and / or any other suitable image features. As an example, the tumor segmentation network can identify object edges by analyzing the image to detect regions with high contrast, large intensity ranges, and / or high intensity variation (e.g., as determined by comparing a region-specific metric to a predetermined threshold). The tumor segmentation network can include nodes corresponding to different receptive fields (and thus analyzing representations of different sets of pixels). Thus, the network can learn to detect and use at least several different types of features.

[0080] In some cases, the tumor segmentation network can utilize the spatial context provided by other images in an image stack to identify a set of edges and / or contours that correspond to the tumor. The image stack can include (for example) three images, with the central image being the image in which the tumor is detected.

[0081] The tumor segmentation network can further generate a two-dimensional (e.g., binary) tumor mask 110 that corresponds to the entire surface area of ​​the tumor within a given detection region using the identified edges and / or contours. The tumor mask 110 can be defined to have a value of 0 over pixels that are not identified as depicting any part of the tumor. Pixels that are identified as depicting part of the tumor can be assigned one value (e.g., for a binary mask) or another value.

[0082] In some cases, a binary tumor mask 110 is generated for each image in the image stack, with each binary tumor mask 110 corresponding to a different axial perspective of the tumor. In such cases, the post-processing controller 111 can aggregate the set of binary tumor masks 110 to construct a 3D tumor mask 110 that represents the three-dimensional positioning and shape of the tumor. E. Organ-specific segmentation network

[0083] In some cases, the neural network processing system 120 may include an organ segmentation controller 111 configured to implement an organ-specific segmentation network. The organ-specific segmentation network may include (for example) a convolutional neural network and / or a 3D neural network. Exemplary convolutional neural networks may include a VGG 16, U-Net, and / or ResNet 18 network. The organ-specific segmentation network may be configured to analyze medical images corresponding to a subject and segment one or more organs depicted in the images. In such a case, each of the one or more organ-specific segmentation networks may be configured to segment a particular type of organ (e.g., via parameters learned during training). Exemplary organs of interest may be (for example) the liver, lungs, kidneys, pancreas, etc.

[0084] In some cases, the organ-specific segmentation network can be configured to perform a series of convolutions, such as depth-wise and point-wise convolutions, as part of the segmentation process. In such cases, one or more dilations along a particular dimension can be further performed. The particular dimension may be a third dimension, a fourth dimension, etc. In some cases, the tumor segmentation network can also apply one or more filters, such as a replication filter.

[0085] In the illustrated example, the organ segmentation controller 111 can control an organ-specific segmentation network configured to detect a particular type of organ. The organ-specific segmentation network can be trained using a training dataset including training images and annotations indicating which portions within each of at least some of the training images depict a particular type of organ. The training dataset can be separate from the training datasets used by the bounding box detection network and the tumor segmentation network. The training dataset can include multiple medical images and corresponding annotations and / or segmentation boundaries (e.g., generated by the annotation device 103) for a particular organ of interest. A set of organ segmentation parameters 112 (e.g., weights) can be learned during training. In some cases, the pre-processing controller 105 can send the same set of medical images 102 to both the bounding box detection controller 106 and the organ segmentation controller 111.

[0086] The trained organ-specific segmentation network can be used to process each of the images and / or the set of preprocessed images to detect organs. The images used to detect a particular type of organ may be the same (or different) as the set of images 102 provided to the bounding box detection controller 106, and the images are then simultaneously provided to the organ segmentation controller 111. The set of images can be divided into multiple (e.g., overlapping) subsets containing one, two, or three images. For example, the subsets can be defined to have three images per subset and one image shift per subset. In some cases, the images can undergo preprocessing to align the images to a 3D image depicting a "full-body" view of the subject.

[0087] Within each image, the organ-specific segmentation network can indicate whether a given image depicts a particular type of organ and further identify the periphery of the organ's depiction. The output of the organ-specific segmentation network can include (for example) an organ mask 113 having a value of zero for pixels that do not depict a particular type of organ and a non-zero value for pixels that do depict a particular type of organ. In some cases, multiple two-dimensional organ masks can be generated corresponding to different virtual slices (e.g., viewpoints) of the organ of interest. These two-dimensional organ masks can be aggregated to generate a 3D organ mask for each organ.

[0088] The post-processing controller 114 can process the tumor mask 110 and the organ mask 113 individually and / or collectively to generate statistics and / or descriptors. For example, for each tumor, the post-processing controller 114 can identify the tumor's volume and further identify whether the tumor is located within any organ (and, if so, which type of organ). The post-processing controller 114 can further process (the two-dimensional or three-dimensional tumor mask) to calculate subject-level tumor statistics, such as the total tumor volume and / or density and / or the sum of the longest dimensions for the subject. In some cases, the sum of the longest dimensions may be the sum of the longest diameters, where the longest diameters are calculated for each tumor and summed to form the sum of the longest diameters. In some cases, the post-processing controller 114 can identify the percentage of the tumor's mass compared to the mass of the corresponding organ of interest, as another exemplary statistic.

[0089] The neural network processing system 120 can output the descriptors and / or statistics to a user device. Additionally, a representation of one or more tumor masks and / or one or more organ masks can be transmitted. For example, an image can be generated that includes a depiction of the original image with an overlay identifying the periphery of each tumor and / or organ detected for the subject. In some cases, the post-processing controller 114 can further process (e.g., send to another model and / or controller for processing) the subject-level tumor statistics to generate a score for the probability of survival using one or more treatment methods.

[0090] While the interactive system shown in FIG. 1A is directed to detecting tumors and determining whether various tumors are located in different organs, alternative embodiments may be directed to detecting other types of biological objects. For example, a first network may be trained to detect brain lesions, and other networks may be trained to detect various brain regions so that it can be determined in which brain region the lesion is located. Thus, alternative embodiments may replace at least the tumor segmentation network with a different segmentation neural network trained to segment other biological structures in medical images. IV. Prediction Network System

[0091] 2 shows an exemplary predictive neural network system 200 that can use one or more output elements (e.g., organ masks) from the neural network processing system 120 to predict a score for the probability of a subject's survival based on the effectiveness of a treatment method. Efficacy can be determined by one or more characteristics of the subject prior to administration of the treatment method (e.g., disease progression as measured in terms of tumor volume or density, etc.).

[0092] If it is desired to predict these scores, the neural network processing system 120 can utilize one or more medical images 202 and organ masks 204 in the predictive neural network system 200. The images 202 can be a subset of the same images used by the bounding box detection network and the tumor segmentation network, as discussed in Section III. In some cases, the images 202 can further include corresponding metrics such as count, volume, and / or tumor location. The organ masks 204 can further include at least one or more organ masks generated by an organ-specific segmentation neural network. In some cases, the neural network processing system 120 can further utilize a tumor mask (not shown) generated by the tumor segmentation network in the predictive neural network system 200.

[0093] In the illustrated example, the prediction network controller 206 may be configured to control and / or perform any of the operations described herein of a predictive neural network, which may be a different neural network than the bounding box detection network and tumor segmentation network described in the neural network processing system 120. The prediction network controller 206 may train the predictive neural network to predict survival or mortality associated with one or more treatment methods for subjects using images corresponding to one or more comparable subject pairs.

[0094] For example, if the first subject and the second subject have both received the same treatment method, and the first subject has a different survival time after receiving treatment compared with the second subject, the subject pair can be considered equivalent.On the other hand, if the first subject has an inconclusive first survival time, such that the first survival time is only followed for a specific period (for example, a certain period in a clinical trial), but no additional data on the first survival time is collected after the specific period, and the second subject has a second survival time that is at least the specific period after the first survival time is followed, the subject pair is not considered equivalent.Therefore, not all possible subject pairings are considered equivalent.

[0095] During training, a set of prediction parameters 208 (e.g., weights) can be determined for the predictive neural network. The training data elements can include at least one or more input images or metrics (e.g., cumulative volume of all detected biological objects) associated with each subject of a comparable subject pair, and a metric measuring each subject's survival time after treatment is administered. A score and / or rank based on each subject's survival time can also be included in the training data elements. The score and / or rank can correspond to the subject's likelihood of survival using the administered treatment. Training can utilize a loss function that maximizes the difference in scores during training between the paired subjects, such that the first subject is determined to have the best likelihood of survival using the treatment compared to the second subject.

[0096] The reference subject data 210 can be a database including at least the treatment method administered, survival time, and one or more subject-level metrics (e.g., number of tumors, tumor location, SLD, or tumor volume) for each of the multiple reference subjects, such that each of the multiple reference subjects can further include subject-level statistics, such as a rank based on the survival time of a single subject compared to the multiple reference subjects. The rank can be a value k ranging from 1 to the total number of subjects among the multiple reference subjects that predicts the relative risk of death (e.g., expressed as the probability that the subject will survive treatment or the subject's expected survival time) for each of the multiple reference subjects. The survival time of each subject can be measured from either the diagnosis of the disease or the start of the subject's treatment period. In some cases, at least some of the multiple reference subjects may have died. The reference subject data 210 can specifically group the reference subjects by the treatment method administered.

[0097] When predicting survival of a subject of interest using a particular treatment method, the predictive neural network can select one or more reference subjects from the reference subject data 210 that meet a criterion for comparability with the subject of interest to form at least one or more subject pairs, such that each pair includes the subject of interest and a different reference subject.

[0098] The prediction network can then determine a prediction score 212 for a given subject by comparing the subject of interest to each of the selected reference subjects. The prediction score 212 can be any suitable metric (e.g., a percentage or duration) indicating the probability and / or length of survival of the subject of interest. The comparison with the reference subjects can include comparing one or more features associated with each reference subject prior to receiving the treatment method to the same features associated with the subject of interest. In some cases, a ranking can be generated for one or more subject pairs such that a subject's rank value can indicate the subject's likelihood of survival. For example, a subject with the lowest rank value can be predicted to have the lowest likelihood of survival using the treatment method. The rank value can be determined from the total number, volume, or density and / or location of tumors for each subject of one or more subject pairs.

[0099] A prediction score 212 can be calculated for the subject of interest based at least on where the subject of interest falls in the rankings compared to the reference subject. It can then be predicted whether and / or to what extent a treatment method may be effective for the subject of interest. V. Exemplary High-Level Process

[0100] 3 shows a flowchart of an exemplary process 300 for using a multi-stage neural network platform to process medical images. Process 300 can be performed using one or more computing systems.

[0101] Process 300 begins at block 305, where a training dataset is accessed. The training dataset includes multiple training elements. The training elements include a set of medical images (e.g., CT images) corresponding to a subject and annotation data identifying the presence of biological objects within the set of medical images. The annotation data includes labels indicating the presence of the biological objects and, if present, the general location of the biological objects (e.g., liver, kidney, pancreas, etc.). The annotation data may be incomplete so as not to include the presence of one or more biological objects. In some cases, a medical image may correspond to two or more different sets of annotation data based on annotations from at least two or more radiologists. In such cases, the different sets of annotation data corresponding to the same image include discrepancies, such as the identification or absence of one or more additional biological objects and / or differences in annotation size and / or object perimeter of one or more biological objects. The training dataset may have been generated using one or more imaging systems and one or more annotation devices, as disclosed in Section III.

[0102] In block 310, a multi-stage neural network platform is trained using the training dataset. The multi-stage neural network platform may include a bounding box detection network and a biological structure segmentation network. In some cases, the neural network platform further includes one or more organ-specific segmentation networks.

[0103] A bounding box detection network can be trained to detect bounding boxes for regions corresponding to biological objects. In particular, training the bounding box detection network includes defining a bounding box for each region corresponding to a biological object in an image. Each of the biological objects can be further labeled to indicate that the bounding region corresponds to a given object (e.g., if multiple objects are identified across a set of images). In some cases, the label can also include the location of the biological object within the subject.

[0104] A biological structure segmentation network (similar to the tumor segmentation network described in FIG. 1A) is trained to identify the boundary and total area of ​​the depicted biological object. Training the segmentation network can include accessing an additional training dataset. The additional training dataset can include all training data elements of the initially accessed training dataset along with labeled segmentation data generated by a radiologist. The labeled segmentation data can include either a binary mask or a three-dimensional mask of the biological object. Optionally, the segmentation network is trained to further correct false positives (e.g., incorrectly labeling background regions as objects) generated by the detection network.

[0105] Training can further be performed using pixel-wise cross-entropy loss, Dice coefficient loss, or compound loss. The loss function can be based on, but is not limited to, mean squared error, median squared error, mean absolute error, and / or entropy-based error.

[0106] A validation dataset may also be accessed to evaluate the performance of the multi-stage neural network platform consistent with its training. The validation dataset may be a separate set of medical images and corresponding annotation data that is distinct from the training dataset. The training session may be terminated when target accuracies for both identification and segmentation of biological objects in the medical images of the validation dataset are reached.

[0107] In block 315, a set of medical images corresponding to the subject and / or a single imaging session is accessed. The set of medical images may depict a thoracic region, an abdominal region, and / or a "whole body" region of the subject. In some cases, a first medical image corresponding to the thoracic region, a second medical image corresponding to the abdominal region, and a third medical image corresponding to the pelvic region may be stitched together to generate a fourth medical image corresponding to the "whole body" region of the subject.

[0108] The medical images can be generated using one or more imaging systems, such as those disclosed in Section III.A. In some cases, the one or more imaging systems may be configured to generate images corresponding to different perspectives of a region of interest. In such cases, the multiple medical images can depict separate virtual slices of a particular region.

[0109] At block 320, the set of medical images is applied to a bounding box detection network. Each image is analyzed to identify one or more bounding boxes. Each bounding box can identify an image region corresponding to a target biological object. The analysis of the images can include using a first virtual slice corresponding to an upper region and / or view of the image and a second virtual slice corresponding to a lower region and / or view of the image, where the first virtual slice and the second virtual slice provide additional spatial context for determining the region corresponding to the target biological object.

[0110] In some cases, the bounding box may include a set of margins (e.g., 10px padding) surrounding the identified region corresponding to the target biological object. If more than one region corresponding to a biological object is identified in the image, the bounding box detection network may identify more than one bounding box for the image.

[0111] In block 325, one or more bounding boxes corresponding to the medical image are utilized by a segmentation network. The segmentation network can crop the medical image to generate a set of detection regions showing a zoomed-in view of each region corresponding to the bounding box. If the regions are smaller than a uniform size, the detection regions can be assigned a uniform size so that the detection regions can include additional padding along with the region corresponding to the bounding box. If the regions are larger than a uniform size, the region corresponding to the bounding box can be divided into two or more detection regions. In the case of multiple detection regions corresponding to a bounding box, the region corresponding to the bounding box can be divided into a set of sliding windows such that some of the windows include overlapping subsets of the region.

[0112] For each detection region associated with a bounding box, the biological structure segmentation network can evaluate image features of the detection region to locate the biological object and generate a first binary mask corresponding to the biological object. If multiple bounding boxes are identified for a given image, the biological structure segmentation network can identify a region within each of the bounding boxes that depicts the corresponding biological object. A binary mask may be generated for each biological object. In some cases, two or more binary masks can be generated for the biological object using images depicting different perspectives of the biological object.

[0113] In block 330, one or more binary masks corresponding to the same object can be processed (e.g., via post-processing) to generate a 3D mask. Each of the one or more binary masks and each 3D mask can correspond to a single biological object. Thus, for example, multiple 3D masks can be generated for a given subject's imaging session, each 3D mask corresponding to one of multiple biological objects.

[0114] Processing the set of binary masks may include aggregating the binary masks to form a 3D structure of the object, as described in Section III.D. Because some of the binary masks may further include overlapping regions, the segmentation network may adjust the regions of one or more binary masks to account for the overlapping regions and / or may select to avoid including one or more binary masks that may depict redundant viewpoints.

[0115] In block 335, the medical images corresponding to the one or more masks (e.g., upon access from block 315) are utilized by one or more organ-specific segmentation networks to determine the location of the biological object. Each organ-specific segmentation network can correspond to a particular organ of interest (e.g., liver, kidney, etc.) and can be trained to identify the particular organ of interest within the images. The organ-specific segmentation network can receive and process the set of images to identify the location of the corresponding organ of interest. If a corresponding organ of interest is detected, the network can further generate a mask of the corresponding organ. The generated organ masks can be binary masks and / or three-dimensional masks.

[0116] At block 340, one or more masks (e.g., one or more 3D biological object masks, one or more 2D biological object masks, and / or one or more organ masks) are analyzed to determine one or more metrics. The metrics may include characteristics of the biological objects. For example, the metrics may include object count, object location and / or type, object count for a particular location and / or type, one or more object sizes, average object size, cumulative object size, and / or the number of objects in each of one or more types of tumors.

[0117] In some cases, the metric includes one or more spatial attributes of the object, such as object volume, object length of the longest dimension, and / or object cross-sectional area. The one or more spatial attributes may be further used to generate object-level statistics for all objects detected within a given object. The object-level statistics may include (for example) a cumulative object volume for a given object, a sum of object lengths of the longest dimension for a given object (e.g., a sum of the longest diameters), and / or a cumulative cross-sectional area of ​​detected objects for a given object.

[0118] In some cases, the metric is compared to another metric associated with a medical image of the same subject collected during a previous imaging date to generate a relative metric (e.g., a percentage or absolute change). The metric can be output (e.g., transmitted to another device and / or presented to a user). The output can then be analyzed by (for example) a medical professional and / or a radiologist. In some cases, the metric is output along with a depiction of one or more masks.

[0119] The metrics can be used (e.g., in a computing system using one or more stored rules and / or via a user) to predict a subject's diagnosis and / or treatment effectiveness. For example, a subject-level statistic such as cumulative biological object volume can be used to determine disease stage (e.g., by determining a range corresponding to the cumulative volume). As another example, relative changes in biological object volume and / or count can be compared to one or more thresholds to estimate whether current and / or previous treatments have been effective.

[0120] In some cases, a metric can be used to predict the score of one or more treatment methods based on the subject's survival probability calculated by the predictive neural network. The score can be predicted using one or more spatial attributes, such as cumulative object volume and / or the sum of the object's longest dimension length. In some cases, one or more scores for the probability of survival can be generated to rank a set of subjects and / or treatments. In such cases, the score of a subject and / or treatment can be compared with one or more scores of another subject and / or another treatment to determine the ranking. The subject-specific ranking can identify at least one or more subjects with the highest probability of survival for a given treatment relative to other previous subjects who have received the given treatment. The treatment-specific ranking can identify the treatment that is most likely to be successful (e.g., survive) for a given subject compared to other treatments. In some cases, the subject-specific ranking and / or treatment-specific ranking are also returned as output. VI. Illustrative Implementation Examples VI.A. Implementation Example 1 VI.A.1. Pipeline for Automatic Tumor Identification and Segmentation

[0121] Tumor segmentation from whole-body CT scans was performed using an automated method of detection and segmentation consisting of a bounding box detection network (discussed in step 1 below) and a tumor segmentation network (discussed in steps 2-3). VI.A.1.a. Step 1: Bounding Box Detection

[0122] A bounding box detection network (referred to herein as the "detection network") with a RetinaNet architecture was used to predict whether a region of a medical image depicted a tumor, generate a bounding box identifying the general spatial location of the tumor within the region of the image, and provide a site label probability for each general spatial location depicting a tumor. Modifications were made to the published RetinaNet architecture in that all convolutions were changed to separable convolutions in training the detection network. For each medical image, an image stack containing a set of three consecutive axial CT slices (without a fixed resolution) was used as input for the detection network. The detection network was trained to detect tumor-containing regions within each slice contained within the image stack, generate a bounding box for each detected region, and attribute them to one of the following available site labels: lung, mediastinum, bone, liver, and other. FIG. 4 shows (i) an exemplary set of images showing a preprocessed whole-body scan of a subject, (ii) a bounding box identifying a tumor predicted to correspond to a mediastinal region and a bounding box identifying a tumor predicted to correspond to a lung region in an axial slice of the subject, and (iii) a bounding box identifying a tumor predicted to correspond to a liver region in another axial slice of the subject.

[0123] The detection network output (i) the proposed coordinates of a bounding box representing the tumor's general spatial location on the central axial slice, and (ii) the probability of each site label category (lung, mediastinum, bone, liver, other). The outputs were concatenated to create a bounding box for each slice of the CT scan, as shown in Figure 4. Each of three consecutive axial CT slices was 512 × 512 in size. Training was performed on 48,000 radiologist-annotated images of axial CT slices with bounding boxes around radiologist-identified RECIST target and non-target lesions and corresponding site locations from 1,202 subjects from the IMPower150 clinical trial. Hyperparameters included a batch size of 0.16, a learning rate of 0.01, and the use of the optimizer ADAM. The detection network was validated against the IMpower131 clinical trial (969 subjects). The lesion-level sensitivity to RECIST interpretation was 0.94. The voxel-level sensitivity was 0.89. VI.A.1.b. Step 2: Tumor Segmentation

[0124] A tumor segmentation network (e.g., implemented in this example as a probabilistic U-Net) was used to identify regions within each bounding box identified by the detection network that delineate tumors (e.g., regions corresponding to areas with a mask value positive and / or equal to 1). As shown in Figure 5, each of the six images corresponds to a bounding box identified by the detection network, and each outlined region identifies a tumor segmentation determined using the tumor segmentation network. A modification was made to the published probabilistic U-Net architecture in that all convolutions were replaced with separable convolutions in training the tumor segmentation network. The tumor segmentation network was configured to average 16 predictions of regions within each bounding box to mimic inter-reader variability and reduce prediction variance. Thus, each prediction corresponds to a different method or criteria used by different radiologists when annotating (or choosing not to annotate) the same lesion; the 16 average predictions were then used to generate a "consensus" by averaging the predictions for each voxel in the image and determining each voxel as part of the tumor if the average prediction was greater than 0.5, or some other threshold. Three axial slices (i.e., 256 x 256 pixels) of 0.7 x 0.7 mm size were used as inputs for the tumor segmentation network, such that each axial slice was correlated with the detected bounding box, which had undergone one or more interim preprocessing techniques (e.g., cropping).

[0125] The tumor segmentation network output a segmentation of the median axial slice, identifying the region within each bounding box that depicts the tumor. The tumor segmentation network was trained on 67,340 images using tumor masks from 1,091 subjects in IMpower150 from a radiologist and volumetric RECIST readings from 2D RECIST. Example hyperparameters included a batch size of 4, a learning rate of 0.0001, and the use of the ADAM optimizer. The exemplary network was validated against IMpower131 (969 subjects; 51,0000 256 × 256 images with 0.7 × 0.7 mm images). A Dice score of 0.82 (using the average of over 16 predictions from the network) was calculated, assuming no false positives in the validation dataset (51,000 images from IMpower131). VI.A.1.c. Step 3: Organ-specific segmentation

[0126] In step 2, the segmentation output from the tumor segmentation network was used to confirm / correct the general spatial location of the tumor proposed by the bounding box detection network in step 1. The subject's whole-body CT scan was used as input for processing by separate organ segmentation networks. In this implementation, the organ segmentation network consisted of multiple convolutional neural networks. Each individual organ segmentation network was trained to perform organ-specific segmentation and return organ masks that identified the location of the organ in the whole-body CT scan. Organ-specific segmentation was achieved by training a different organ segmentation network for each organ of interest, e.g., right lung, left lung, liver, spleen, kidney, bone, and pancreas. Each organ-specific segmentation network had a 3D U-Net architecture with batch normalization and leaky ReLU activation at each layer. Organ-specific segmentation networks for the kidney, spleen, and pancreas used publicly available datasets for training, specifically completing Kits19 for kidney (such as the dataset in Heller, N. et al., "The KiTS19 Challenge Data: 300 Kidney Tumor Cases with Clinical Context, CT Semantic Segmentations, and Surgical Outcomes" (2019), which is incorporated by reference in its entirety for all purposes), and Medical Decathlon for spleen and pancreas (as described in Simpson, A. L. et al., "A large annotated medical image dataset for the development and evaluation of segmentation algorithms" (2019), which is also incorporated by reference in its entirety for all purposes). Ground truth for the bone segmentation network was based on morphological operations.

[0127] For each organ-specific segmentation network, the input was a 256 × 256 × 256 CT volume (the concatenation of axial slices from steps 1–2) resampled to a voxel size of 2 × 2 × 2 mm. The output of each organ-specific segmentation network was an organ mask of the same size for each organ. The ground truth for each network was the corresponding 256 × 256 × 256 organ mask with the same voxel size. Hyperparameters included a batch size of 4, a learning rate of 0.0001, and the use of the optimizer ADAM. Data augmentation through a combination of rotation, translation, and zoom was used to augment the dataset for more robust segmentation and avoid overfitting. An initial version of the organ-specific segmentation network trained as described herein yielded the following results: lung: 0.951; liver: 0.964; kidney: 0.938; spleen: 0.932; pancreas: 0.815; bone: 0.917 (ground truth generated using morphological operations). VI.A.2. Time-Separated Pairwise Comparisons VI.A.2.a. Overview

[0128] The CT scans, organ-specific segmentation, and techniques described herein were further used in conjunction with automated tumor detection and segmentation to generate several other predictions and estimates to assist clinicians in determining which treatment to prescribe. Once one or more tumors and / or "whole body" tumor burden were identified using the automated pipeline, the model predicted the subject's likelihood of survival using one of a number of metrics, given each of a number of potential treatments for a given oncological indication, in terms of overall survival, progression-free survival, or other similar metrics. The model output a ranking of treatments for a given subject to identify the treatment that provided the longest survival time. Alternatively, the model output a ranking of subjects to identify those likely to experience the longest survival time with a given treatment. VI.A.2.b. Model Architecture and Training

[0129] Given two subjects A and B, it was assumed that the outcome (overall survival) was observed for at least one subject. Without loss of generality, (T A ) the observed outcome for subject A, and (T B (denoted as T) the results for subject B are censored or B >T A It was assumed that death would occur at .

[0130] The inputs to the network were CT scans and organ masks (e.g., liver, lungs, kidneys, bone, pancreas, and spleen) obtained using one or more organ-specific segmentation networks for both subjects A and B. The architecture of the organ-specific segmentation network was a dilated VGG16, ResNet18, or similar network, with separable convolution outputting a score vector with N elements (e.g., 1000) for each subject. The dilation was generally performed according to the techniques described in Carreira, J. and Zisserman, A., "Que Vadis, Action Recognition? A New Model and the Kinetics Dataset," in CVPR (2017), which is incorporated herein by reference in its entirety for all purposes. However, in this implementation, the separable convolution was performed in two steps (first a depthwise convolution, followed by a pointwise convolution). However, rather than dilating along three dimensions as in conventional convolution, the dilation was split into two steps. For depthwise convolution, dilation was performed along three dimensions, then a replication filter was applied once, and the average was calculated along the fourth dimension (number of input filters). For pointwise convolution, the average was determined over the first two dimensions, dilation was performed along the third dimension, and replication was performed in the fourth dimension. The above modifications facilitated the processing of large (by pixel / voxel count) 3D whole-body CT images using the network while achieving functional model performance.

[0131] During training, subjects A and B (S A and S B) were compared. The training procedure is a loss L = exp(S B ) / exp(S B )+exp(S A The goal was to minimize the . The training data included 42,195 pairs of matched subjects from 818 subjects in the IMpower150 clinical trial, divided by treatment group. Hyperparameter selection included a learning rate (Ir) of 0.0001, a batch size of 4, and the use of the optimizer ADAM. Example model results for pairwise comparisons show that 74% of pairwise comparisons were accurate in the test (validation) set of 143 subjects from the three treatment groups of GO29436 (IMpower150). For these results, comparisons were made only between subjects within treatment groups. VI.A.3. Results

[0132] The performance of the automated method on the training and test datasets was determined using RECIST and manual annotation of "whole body" tumor burden. RECIST interpretation was performed on both datasets as a baseline calculation of the number of identified lesions and total lesion volume for each subject.

[0133] Figure 6A shows a correlation plot comparing the number of lesions derived by RECIST interpretation (shown on the x-axis of the plot) with the number of lesions determined by the automated detection and segmentation method (shown on the y-axis) for the training dataset (IMpower150). Figure 6B shows another plot comparing the tumor burden (measured as the total volume of all identified lesions) derived by RECIST (shown on the x-axis of the plot) with the tumor burden for tumors identified by the automated method (shown on the y-axis of the plot). Both plots show a rightward slope, indicating that RECIST interpretation had the highest correlation with data from the automated method for the lower ranges of lesion number and total lesion volume. The standard deviation and standard error calculated based on the difference in lesion number predictions between the two techniques were 2.95 and 0.091, respectively. The standard deviation and standard error calculated based on the difference in total tumor volume predictions between the two techniques were 5.2 w / e + 0.01 and 2.40, respectively. Figures 7A-7B show similar correlation plots for the test dataset (IMpower131), with the mean number of lesions determined using RECIST interpretation on the x-axis and the number of lesions determined using the automated method on the y-axis. For the test dataset, the standard deviation and standard error calculated based on the difference between the two techniques' predictions of lesion count were 6.05 and 0.24, respectively. The standard deviation and standard error calculated based on the difference between the two techniques' predictions of total lesion volume were 5.22e+01, with a standard error of 2.40, respectively.

[0134] The training dataset (IMpower150) was further used to perform a full reading, which involved determining the subject's total tumor burden through manual annotation of each tumor by a radiologist, rather than annotating only a single slice as is done with RECIST reading. Figure 8 shows a plot where the y-axis corresponds to the number of lesions determined by the radiologist (e.g., for a full reading) and the x-axis corresponds to the number of lesions determined by RECIST for the set of subjects. Each point on the plot represents a subject in the training dataset for a total of 15 subjects who underwent both a full reading and a RECIST reading. Because a full reading identifies a greater amount of lesions compared to a RECIST reading, the plot shows little agreement between RECIST reading and a full reading. The standard deviation and standard error calculated based on the difference in predictions between the two techniques were 6.64 and 0.30, respectively.

[0135] Further comparisons were made between the automated method and full readings performed by radiologists to determine the total tumor burden for each subject. Figure 9 shows a correlation plot between the total lesion volume determined by full readings performed by radiologists (shown on the y-axis) and the total lesion volume determined by the automated method (shown on the x-axis), with each point representing a subject in the IMpower150 training dataset. Multiple readings were calculated for each subject from the set of training subjects, as shown in the plot. Figures 10A-10B show plots comparing the mean and median total lesion volumes determined by the automated method (shown on the x-axis, respectively) with the mean and median total lesion volumes determined by full readings for each subject (shown on the y-axis, respectively). As in Figures 8-9, each point in both plots represents a subject in the training dataset. As shown in the plots, the automated method generally identified lesions of the same or greater volume than the full readings.

[0136] Prognostic data were also collected for subjects represented in the training and test datasets, such that the number of identified lesions and the calculated total volume of lesions were used to predict the subject's probability of survival over a given period. More specifically, subjects in the training dataset were assigned to specific clusters based on various statistics of lesions detected using RECIST technology, and survival curves were calculated for each cluster to demonstrate whether the various statistics predict survival. Figures 11A-14B show Kaplan-Meier curves illustrating exemplary prognostic data for the training dataset.

[0137] Figure 11A shows the survival probability of subjects clustered based on SLD calculations for lesions identified by RECIST. Figure 11B shows the survival probability of subjects clustered based on the number of lesions identified by RECIST. The y-axis of the plot corresponds to survival probability, and the x-axis corresponds to elapsed time (e.g., measured in days). Clusters were determined such that the first quartile (Q1) corresponds to subjects with the top 25% of lesion counts and / or SLD scores, the second quartile (Q2) corresponds to subjects within the next 25%, the third quartile (Q3) corresponds to subjects within the next 25%, and the fourth quartile (Q4) corresponds to subjects within the bottom 25%. As shown in the plot, subjects within the first quartile of diameter sum SLD and subjects within the first quartile of lesion counts have a lower probability of survival compared to subjects within the fourth quartile. Therefore, spatial statistics of automatically detected lesions appear to predict survival prognosis.

[0138] Instead, Figure 11C shows a Kaplan-Meier curve illustrating subject survival probability as determined by the disclosed automated method. Figure 11C shows a plot illustrating subject survival probability over time as determined by the automated method. Regarding the clustering associated with Figure 11C, subjects were clustered based on total SLD relative to total tumor burden. Figures 12A-12B further show plots of subject survival probability based on total volume and number of identified lesions, also determined by the automated method. It is clear that high tumor burden, as measured by either a large volume of lesions or a large number of lesions, correlates with a decreased subject survival probability.

[0139] Using the identified lesion locations, we evaluated the degree to which prognosis (e.g., probability of survival) was predicted by statistics based on automated tumor detection and segmentation for subjects in the training dataset, shown in Figures 13A-14B. Specifically, Figures 13A-13B show a series of Kaplan-Meier curves illustrating subject survival. Subject groups were defined based on lung lesion volume (shown in the corresponding A plot) and lung lesion count (shown in the corresponding B plot). Notably, the survival curves differed between subject groups, suggesting that lesion volume and lesion count predict survival metrics. Figures 14A-14B show subject survival based on the spread of lesions (e.g., metastases) to the subject's lung and bone regions, respectively. Survival rates were reported as higher when lesions were not present in either the subject's lung or bone regions.

[0140] Figures 15A-19 also show Kaplan-Meier curves of exemplary prognostic data for the test dataset. Figures 15, 16, 17, and 18 correspond to the same label variables (e.g., y-axis corresponding to survival probability and x-axis corresponding to days elapsed) and methods as Figures 10, 11, 12, and 13, respectively. Figure 19 shows the survival probability of subjects based on renal metastasis in the test dataset.

[0141] It will be appreciated that the test dataset contained different images from a different set of subjects than the training dataset and was used to externally validate the results derived from the holdout portion of the training dataset. In particular, the plots shown in Figures 15A-19 illustrate subject prognosis that exhibits a greater correlation between survival and tumor location and / or tumor burden or volume when determined by automated methods compared to RECIST or full image interpretation. VI.B. Implementation Example 2 VI.B.1. Overview

[0142] This example implementation uses automated methods of bounding box detection and tumor segmentation to identify the complete 3D tumor burden on whole-body diagnostic CT scans in subjects with advanced metastatic disease (i.e., lesions spread across multiple organs). This method differs from Example Implementation 1 in that it did not use organ-specific segmentation to identify the location of the segmented tumor or generate organ masks.

[0143] The implemented method is based on a bounding box detection network implemented as RetinaNet for lesion detection and tagging, followed by a tumor segmentation network implemented as an ensemble of probabilistic UNets that allows segmentation of the detected lesions.

[0144] The presented study was developed using two multicenter clinical trials, identifying over 84,000 RECIST lesions from 2,171 patients with advanced non-small cell lung cancer across 364 clinical sites. As a result, the method accounted for inter-reader variability and heterogeneity of scan acquisition across hospital sites. Tumors identified using the automated bounding box detection and tumor segmentation techniques described in this disclosure were compared with manually identified RECIST tumors and manually segmented target lesions at the voxel level. Furthermore, fully automated estimates of baseline tumor burden were compared with radiologists' manual measurements regarding the prognostic value of tumor burden for the subject's overall survival.

[0145] The results demonstrate state-of-the-art detection and segmentation performance of RECIST target lesions in a holdout set of 969 subjects containing over 35,000 tumors. Furthermore, the results demonstrate that total tumor burden may have clinical utility as a prognostic factor for a subject's overall survival. The proposed method can be used to streamline tumor assessment in diagnostic radiology workflows and, if further developed, could potentially enable radiologists to assess response to treatment when applied serially. VI.B.2. Method

[0146] The techniques described in this disclosure were used to identify total tumor burden from whole-body CT scans. The approach involved three steps: bounding box detection, tumor segmentation, and post-processing, and the resulting end-to-end method captured the various properties of the available CT data and RECIST annotations.

[0147] The detection step utilized a bounding box detection network, implemented as RetinaNet, which used bounding boxes and lesion tags to identify both target and non-target lesions. RetinaNet uses a single-stage detection approach that provides very fast object detection. Given that whole-body CT scans often contain over 200 axial slices, efficient processing was highly advantageous.

[0148] In the segmentation step, a tumor segmentation network implemented as a set of probabilistic UNets generated an ensemble of reasonable axial lesion segmentations based solely on the 2D segmentations of the target lesions.

[0149] Tumor segmentation of metastatic cancer subjects is prone to reader subjectivity, and therefore there cannot be a single ground truth for a given lesion. Probabilistic U-Nets [8] enable memory-efficient generative segmentation that allows for sampling segmentation variants from a low-dimensional latent space. The use of probabilistic U-Nets for segmentation is further described in Kohl, S. et al., "A probabilistic U-Net for segmentation of ambiguous images," Advances in Neural Information Processing Systems (NIPS 2018), pp. 6965-6975 (2018), which is incorporated by reference in its entirety for all purposes. Therefore, probabilistic U-Nets were chosen to mimic reader-to-reader annotation variability.

[0150] This part of the model allowed for the generation of an ensemble that traded off inter-reader variability and overall agreement between radiologists' segmentations. A post-processing step combined the predicted 2D segmentations to generate a unified whole-body 3D tumor mask. Furthermore, post-processing also addressed the variability in image acquisition parameters encountered in our multisite dataset, which resulted in different information limits and varying signal-to-noise ratios across scans. Tumors detected via this automated technique were compared to tumors detected via a manual technique in which radiologists outlined marked bounding boxes around selected target and non-target lesions. VI.B.2.a. Tumor detection

[0151] In the data evaluated in this implementation, tumor location tags were highly imbalanced across organs, with lung lesions accounting for 45% and 40% of the training and test datasets, respectively, while 128 locations accounted for less than 0.5% of the tags. Focal loss was used to address class imbalance.

[0152] RetinaNet with ResNet-50-FPN was used to detect tumors axially. (See Lin, T.Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S., "Feature pyramid networks for object detection," CVPR (2017), incorporated by reference in its entirety for all purposes.) For non-maximum suppression, the maximum number of objects per image was set to 32, and the number of anchors was set to 9. Here, 32 represents an upper bound on the number of tumors that can reasonably be expected within a single axial slice. To provide spatial context around the central slice, the model was configured to receive three axial slices as input, which served as three feature channels. Due to the low prevalence of many tags, classes were simplified to lung, liver, bone, mediastinum, and other locations.

[0153] In the test setup, RetinaNet was applied sequentially to all axial slices, and the predicted bounding boxes were extended to the previous and next slices to minimize false negatives. VI.B.2.b. Tumor segmentation

[0154] Experiments were performed with β=2;5;10 and had stand-alone, or crossed or joined ensembles.

[0155] The best results were obtained using a combination of two masks with β=2 and β=10.

[0156] Varying the training loss, β, allowed us to provide different weights to the Kullback-Leibler divergence term in the loss, and thus to place different importance on spanning the latent space of segmentation variants. This parameter allowed us to generate tumor segmentation variants that mimic human reader variability, or to generate a consensus segmentation.

[0157] The training dataset was constructed using RECIST target lesion segmentations from two radiologists per scan and 3D segmentations of several scans. Images were resampled to an in-plane resolution of 0.7 × 0.7 mm, and 256 × 256 × 3 pixel patches were constructed around these lesions. The previous and next slices were used as spatial context. Larger patches than the input were employed: 180 × 180 pixels with an in-plane resolution of 0.5 × 0.5 mm. This choice was made because the data being evaluated represent advanced-stage cancers, which exhibit many large lesions.

[0158] In the test setup, a patch centered on the detected lesion (e.g., as provided by the detected bounding box) and then resampled to the input resolution of the stochastic UNet was segmented as shown in Figure 20. If the detected tumor was larger than the patch size, a sliding window was used to segment the entire detected tumor. VI.B.2.c. Systemic assessment

[0159] Acquisition protocols vary from hospital to hospital and machine to machine, even within the same institution. As a result, voxel sizes were variable in the dataset (in-plane from 0.6 to 1.45 mm and slice thickness ranging from 0.62 to 5 mm). These differences can induce signal-to-noise ratio (SNR) variations and result in tumor segmentation that can only be detected in high-resolution scans. To homogenize the information extracted from all CT scans, we applied binary closure to the tumor mask using a cubic 3 × 3 × 5 mm structuring element to account for SNR differences and retain only tumors greater than 10 mm in height. VI.B.3. Experiments and Results VI.B.3.a. Data

[0160] The dataset consisted of over 84k lesions from a total of 14,205 diagnostic computed tomography scans from two randomized clinical trials. Training and test data were split by trial. The first trial (Clinical Trial NCT02366143, described in Socinski, MA et al., "Atezolizumab for First-Line Treatment of Metastatic Nonsquamous NSCLC," N Engl J Med 378, 2288-2301 (2018)) included 1,202 available advanced nonsquamous non-small cell lung cancer subjects. This first test dataset was used for training. The second trial (Clinical Trial NCT02367794) included 969 advanced squamous non-small cell lung cancer subjects and served as a holdout set. Data were collected across 364 unique sites (238 training sets, 237 test sets) and annotated by a total of 27 different radiologists. Thus, the data provide significant subject, image acquisition, and inter-reader variability.

[0161] For each study, subjects visited an average of 6.5 times for a total of 7,861 scans in the training set and 6,344 scans in the test set. Each scan was interpreted by two radiologists according to RECIST 1.1 criteria. Tumor annotations consisted of 2D lesion segmentations of target lesions and bounding boxes of non-target lesions. In total, across all visits and radiologists, there were 48,470 annotated tumors in the training set and 35,247 in the test data. Furthermore, for each identified target and non-target tumor, we identified available lesion tags from 140 possible location labels, as detailed in Table 1. In addition to the 2D annotations, 4,342 visits (two visits per subject) resulted in volumetric segmentation of the target tumor only. Whole-body coverage was available for whole-body assessment of 1,127 subjects at screening in the training set and 914 subjects in the test set. [Table 1] VI.B.3.b. Results

[0162] Example Implementation. RetinaNet for tumor detection and tagging was implemented using PyTorch and the ADAM optimizer. ResNet-50-FPN was initialized using a pre-trained model on ImageNet. The learning rate was set to 1e-4 and the batch size was set to 16. The network was trained for 416,000 iterations.

[0163] Stochastic UNet was implemented using PyTorch and the ADAM optimizer. The learning rate was set to 1e-5 and the batch size was set to 4. Two versions were kept with training loss β = 2 and 10. The network was trained for 50 epochs.

[0164] Detection and Segmentation Performance. Average lesion and class-level sensitivity per image for detection in Table 2 and Table 1. Sensitivity was obtained with an average of 0.89 "false positives" (FP) per image. Due to incomplete RECIST annotation, these FPs may actually be unannotated lesions. As in Yan, K., et al., "MULAN: Multitask Universal Lesion Analysis Network for Joint Lesion Detection, Tagging, and Segmentation," average sensitivity values ​​were derived at 0.5, 1, 2, and 4 FP / image (88.4%). In: Frangi, AF, Schnabel, JA, Davatzikos, C., Alberola-Lopez, C., Fichtinger, G. (eds.) MICCAI 2019. LNCS, vol. 11769, pp. 194-202. Springer, Cham (2019) and Liao, F., Liang et al.: Evaluate the malignancy of pulmonary nodules using the 3D deep leaky noisy-or network.IEEE Trans.Neural Netw.Learn.Syst.(2019). [Table 2]

[0165] For segmentation, statistics included the mean voxel-level sensitivity and mean error for the estimated longest dimension of RECIST lesions in the test set.

[0166] Prediction of survival from baseline scans. Using the tumor detection and segmentation model estimated from the training data, the length along the longest dimension of all detected and segmented lesions was calculated from the baseline scan for each subject in the test dataset. With survival time as the outcome variable, the right panel of Figure 22 shows a Kaplan-Meier plot based on the empirical quartiles of baseline SLD extracted by the model (for subjects in the test set). For comparison, for the same subjects, the left panel shows a Kaplan-Meier plot based on the empirical quartiles of SLD derived by RECIST. As can be seen, the automated method closely reproduced the pretreatment survival risk profile of tumor burden compared to that generated through radiologist annotation according to RECIST criteria. VI.B.4. Interpretation

[0167] The results demonstrate the powerful performance of our multistage segmentation platform. The fully automated algorithm successfully identified and performed 3D segmentation of tumors in standard diagnostic whole-body CT scans. This methodology demonstrated superior performance in detection and segmentation compared to radiologists and, importantly, performed well on tumors in multiple different organs. These results demonstrate that this technology can be a powerful support tool for radiologists by providing an initial tumor burden assessment for examinations, which should improve accuracy, reproducibility, and speed. Furthermore, the algorithm generates metrics such as whole-body tumor volume (which typically takes too long for radiologists to assess), which could be valuable as a prognostic tool or novel endpoint for clinical trials, providing a more complete view of the target disease for use in clinical radiology practice. VII. Further Considerations

[0168] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0169] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0170] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments provides those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.

[0171] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

Claims

1. 1. A computer-implemented method comprising: accessing one or more medical images of a subject; inputting the one or more medical images into a detection network to generate one or more masks that identify a set of regions within the one or more medical images, wherein the detection network predicts that each region of the set of regions identified in the one or more masks includes a depiction of one of one or more tumors within the subject; For each region of the set of regions, processing the region of the one or more medical images using a tumor segmentation network to generate one or more tumor segmentation boundaries of the tumor present within the subject; for each tumor of the one or more tumors, determining an organ in which at least a portion of the tumor is located by using a plurality of organ-specific segmentation networks; generating an output based on the one or more tumor segmentation boundaries and the location of the organ in which at least a portion of the one or more tumors is located.

2. processing the region to generate the one or more tumor segmentation boundaries; for each of a plurality of 2D medical images, identifying a segmentation boundary of the tumor therein, the segmentation boundary being a tumor segmentation boundary of the one or more tumor segmentation boundaries; 10. The method of claim 1, comprising: defining a three-dimensional segmentation boundary based on the segmentation boundaries associated with a plurality of 2D medical images, wherein the output includes or indicates the three-dimensional segmentation boundary.

3. 2. The method of claim 1, wherein each of the one or more tumor segmentation boundaries is defined to be a segmentation boundary line of a two-dimensional cross-section of the depicted tumor, and the output includes or indicates the one or more tumor segmentation boundaries.

4. determining, for each tumor of the one or more tumors, spatial attributes based on a tumor segmentation boundary of one of the one or more tumor segmentation boundaries, wherein the spatial attributes include: tumor volume, the length of the tumor along a particular dimension or longest dimension, and / or determining, including a cross-sectional area of ​​the tumor; 10. The method of claim 1, further comprising: calculating subject-level tumor statistics for the one or more tumors based on the spatial attributes, wherein the output comprises the subject-level tumor statistics.

5. 5. The method of claim 4, wherein the one or more tumors comprise a plurality of tumors, the spatial attribute determined for each tumor of the one or more tumors comprises a length of the tumor along its longest dimension, and the subject-level tumor statistic comprises a sum of the lengths of the tumors.

6. 10. The method of claim 1, further comprising determining a percentage or absolute difference between the subject-level tumor statistic and another tumor statistic associated with the subject, the other tumor statistic being generated based on an analysis of one or more other medical images of the subject, each of the one or more other medical images being collected at a benchmark time point prior to the time at which the one or more medical images were collected, and the output includes or is based on the percentage or absolute difference.

7. comparing said percentage or absolute difference to each of one or more predetermined thresholds; 7. The method of claim 6, further comprising determining an estimate of prognosis, treatment response, or disease state based on the threshold comparison, wherein the output comprises the estimated prognosis, treatment response, or disease state.

8. The method of claim 1 , wherein the one or more medical images include one or more computed tomography (CT) images.

9. The method of claim 1 , wherein the one or more medical images include a CT image of the whole body or torso.

10. The method of claim 1 , wherein the one or more medical images include one or more MRI images.

11. The method of claim 1 , wherein the detection network is configured to use focal loss.

12. The method of claim 1 , wherein the tumor segmentation network comprises a modified U-Net that includes separable convolutions.

13. The method of claim 1 , wherein each of the plurality of organ-specific segmentation networks comprises a modified U-Net that includes separable convolutions.

14. 10. The method of claim 1, further comprising determining, for each tumor of the one or more tumors, based on location within the organ, wherein the output comprises the organ-specific count.

15. inputting the one or more medical images into a computer by a user; The method of claim 1 , further comprising presenting, by the computer, at least one visual representation of the tumor segmentation boundary.

16. The method of claim 1 , further comprising capturing the one or more medical images with a CT machine.

17. 10. The method of claim 1, further comprising providing, by a physician, a preliminary diagnosis of the presence or absence of cancer and any associated organ location, wherein the preliminary diagnosis is determined based on the output.

18. The method of claim 1 , further comprising providing, by a physician, a treatment recommendation based on the output.

19. 1. A computer-implemented method comprising: transmitting one or more medical images of a subject from a local computer to a remote computer located across a computer network, the remote computer comprising: inputting the one or more medical images into a detection network to generate one or more masks that identify a set of regions within the one or more medical images, wherein the detection network predicts that each region of the set of regions identified in the one or more masks includes a depiction of one of one or more tumors within the subject; For each region of the set of regions, processing the region of the one or more medical images using a tumor segmentation network to generate one or more tumor segmentation boundaries of the tumor present within the subject; and for each tumor of the one or more tumors, determining an organ in which at least a portion of the tumor is located by using a plurality of organ-specific segmentation networks; and receiving results based on the one or more tumor segmentation boundaries and the location of the organ in which at least a portion of the one or more tumors is located.

20. 20. The method of claim 19, further comprising capturing the one or more medical images with an MRI or CT machine.

21. 1. A computer-implemented method comprising: accessing one or more medical images of a subject; accessing a set of organ locations for a set of tumor lesions present in the one or more medical images; inputting the one or more medical images and the set of organ locations into a network associated with one of a plurality of therapeutic treatments to generate a score representing whether the subject is a good candidate for a particular therapeutic treatment compared to other therapeutic treatments; and returning the score.

22. accessing the set of organ locations for the set of tumor lesions present in the one or more medical images; inputting at least one of the one or more medical images into a detection network to generate one or more masks that identify a set of regions of the one or more medical images that are predicted to be indicative of one or more neoplastic lesions within the subject; and for each tumor in the set of tumor lesions, determining an organ in which at least a portion of the tumor is located by using multiple organ-specific segmentation networks.

23. 23. The method of claim 22, wherein the detection network is trained using a set of comparable subject pairs, where comparable subject pairs have received the therapeutic treatment and have survived for different periods of time after receiving the therapeutic treatment, and wherein said training comprises using a loss function that maximizes the difference in the scores during training between the subjects of the pairs.

24. The loss function used during training is L = -exp(S B ) / exp(S B ) + exp(S A 24. The method of claim 23, comprising:

25. 23. The method of claim 22, wherein each of the plurality of organ-specific segmentation networks comprises an inflated VGG 16 or an inflated ResNet 18 network.

26. 23. The method of claim 22, wherein each of the plurality of organ-specific segmentation networks comprises depth-wise followed by point-wise convolutions.

27. inputting the one or more medical images into a computer by a user; 23. The method of claim 22, further comprising: providing, by the computer, a recommendation as to whether the therapeutic treatment is appropriate for the subject.

28. 23. The method of claim 22, further comprising capturing the one or more medical images with an MRI or CT machine.

29. 23. The method of claim 22, further comprising: prescribing, by a physician, the therapeutic treatment responsive to the score indicating that the therapeutic treatment will be beneficial to the subject.

30. 1. A computer-implemented method comprising: transmitting one or more medical images of a subject from a local computer to a remote computer located across a computer network, the remote computer comprising: accessing a set of organ locations for a set of tumor lesions present in the one or more medical images; inputting the one or more medical images and the set of organ locations into a network associated with one of a plurality of therapeutic treatments to generate a score representing whether the subject is a good candidate for a particular therapeutic treatment compared to other therapeutic treatments; transmitting, receiving the score from the remote computer at the local computer.

31. 31. The method of claim 30, further comprising capturing the one or more medical images with a CT or MRI machine.

32. 1. A system comprising: one or more data processors; a non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.

33. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.