Diagnosis and assessment of lumen-related disease

A computer-implemented method and system using machine learning models analyze imaging data to objectively assess the severity of lumen-related diseases like CF, addressing the limitations of conventional manual analysis methods.

WO2025122741A1PCT designated stage expired Publication Date: 2025-06-12THE TRUSTEES OF INDIANA UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058658
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-05
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Conventional diagnosis and staging of lumen-related diseases such as cystic fibrosis (CF) are costly, time-consuming, and subjective, relying on manual analysis by trained experts of high-resolution 3D CT imaging datasets, which limits accessibility and accuracy.

Method used

A computer-implemented method and system that utilize machine learning models to analyze volume data from imaging modalities, counting lumens and generation splits, measuring diameters and volumes, calculating texture metrics, detecting and quantizing lesions, and determining the severity stage of lumen-related diseases based on feature metrics.

Benefits of technology

The system provides accurate, reliable, and objective preliminary assessments of disease severity, reducing interclinician and intersite discrepancies, and enabling faster, more cost-effective diagnosis and treatment planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058658_12062025_PF_FP_ABST
    Figure US2024058658_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Systems, computer-implemented methods, and computer-readable media stage lumen-related disease, such as cystic fibrosis. One or more computer processors receive volume data, such as computed tomography scan data. A number of lumens and a number of generation splits visible in the volume data are counted. Diameters of the lumens are measured. At least one volume of a network of the lumens is measured. One or more texture metrics are calculated from the volume data. Lesions are detected and quantized from the volume data. One or more trained machine learning (ML) models are used to determining a severity stage of the lumen-related disease. The one or more ML models can determine the severity stage based on feature metrics derived from the counted number of lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.
Need to check novelty before this filing date? Find Prior Art

Description

IUIC-00155 DIAGNOSIS AND ASSESSMENT OF LUMEN-RELATED DISEASE CROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Patent Application SerialNo. 63 / 607,028, filed on December 6, 2023, the disclosure of which is hereby incorporated by reference in its entirety TECHNICAL FIELD

[0002] This description relates generally to disease assessment, and more particularly todiagnosis and assessment of lumen-related disease. BACKGROUND

[0003] Lumen-related diseases may affect one or more hollow-bodied, branchinganatomical networks such as vessels or tubules. One example is cystic fibrosis (CF), a life- shortening genetic disease caused by one or more mutations in the CF transmembrane conductance regulator (CFTR) gene. The one or more mutations affect the flow of chloride ions across cell membranes, which leads to an overproduction of thick mucus in lung airways, reducing pulmonary function and resulting in cellular infiltration and persistent lung infections. CF affects more than 70,000 people worldwide, and affects one in 3,400 births in the United States. CF presently has an associated average lifetime cost of over $300,000 (U.S.) per patient, a cost that is rapidly increasing with the rising cost of healthcare generally. Although clinical symptoms and manifestations of CF can appear in diverse organs and structures, CF is nevertheless considered a pulmonary disease because mortality is mainly due to respiratory complications. SUMMARY

[0004] An example computer-implemented method stages lumen-related disease. One ormore computer processors receive volume data. A number of lumens and a number of generation splits visible in the volume data are counted. Diameters of the one or more lumens are measured. At least one volume of the one or more lumen is measured. One or more texture metrics are calculated from the volume data. Lesions on one or more walls of a hollow body forming the one or more lumen are detected and quantized from the volume data. One or more trained machine learning (ML) models are used to determine a severity stage of the lumen-IUIC-00155 related disease. The one or more ML models can determine the severity stage based on feature metrics derived from the counted number of one or more lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.

[0005] An example system includes one or more hardware processors and one or morenon-transitory computer-readable media storing instructions. When executed by the one or more hardware processors, the instructions cause receiving volume data and counting a number of one or more lumens and a number of generation splits visible in the volume data. The instructions further cause measuring diameters of the one or more lumens. The instructions further cause measuring at least one volume of the one or more lumens or calculating at least one volume of the combined one or more lumens. The instructions further cause calculation of one or more texture metrics from the volume data. The instructions further cause detection and quantization of lesions from the volume data. The instructions further cause one or more trained machine learning (ML) models to determine a severity stage of the lumen-related disease. The one or more ML models can determine the severity stage based on feature metrics derived from the counted number of lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.

[0006] An example includes one or more non-transitory computer-readable media storingprogram instructions that, when executed by one or more processors, cause the one or more processors to receive volume data and count a number of lumens and a number of generation splits visible in the volume data. The instructions further cause the one or more processors to measure diameters of the lumens. The instructions further cause the one or more processors to measure at least one volume of a network of the lumens. The instructions further cause the one or more processors to calculate one or more texture metrics from the volume data. The instructions further cause the one or more processors to detect and quantize lesions from the volume data. The instructions further cause one or more trained machine learning (ML) models to determine a severity stage of the lumen-related disease. The one or more ML models can determine the severity stage based on feature metrics derived from the counted number of lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.IUIC-00155 BRIEF DESCRIPTION OF THE DRAWINGS

[0007] FIG. 1A is a block diagram of an example system for diagnosis of lumen-relateddisease.

[0008] FIGS. 1B and 1C are block diagrams of example systems for diagnosis andassessment of lumen-related disease.

[0009] FIG. 1D is a block diagram of an example system for diagnosis of CF.

[0010] FIG. 2A is a block diagram of an example volume data input pre-processor.

[0011] FIG. 2B is a block diagram of an example lung volume data input pre-processor.

[0012] FIG. 2C is an illustration of example lung volume data as a set of computedtomography slices.

[0013] FIG. 2D is an illustration of example lung data segmented from the lung volumedata of FIG.2C.

[0014] FIG. 2E is an illustration of example trachea data segmented from the lungvolume data of FIG.2C.

[0015] FIG. 2F is an illustration of example tracheobronchial data combined from thesegmented data of FIGS.2D and 2E.

[0016] FIG. 2G is example tracheobronchial data separating the left and right lungs in thetracheobronchial data of FIG.2F.

[0017] FIG. 3A is a block diagram of an example lumen analyzer.

[0018] FIG. 3B is a block diagram of an example airway analyzer.

[0019] FIG. 3C is an example 3D rendering of an example lungs mask for example lungs.

[0020] FIG. 3D is an example 3D rendering of example detected potential airways voxelsfor the example lungs of FIG.3C.

[0021] FIG. 3E is an example 3D rendering of example airway tree or branch data for theexample lungs of FIG.3C.

[0022] FIG. 3F is an example 3D rendering of a skeletonization of the example airwaytree or branch data of FIG.3E for the example lungs of FIG.3C.

[0023] FIGS. 4A and 4B are block diagrams of example texture analyzers.

[0024] FIG. 5A is a block diagram of an example tissue or organ registrar.

[0025] FIG. 5B is a block diagram of an example tissue or lung registrar.IUIC-00155

[0026] FIGS. 6A and 6B are block diagrams of example gray-level co-occurrencematrices (GLCMs) and features calculators.

[0027] FIG. 6C is a flow diagram of an example method of GLCM calculation.

[0028] FIGS. 6D, 6E, and 6F are example GLCM texture contrast maps computed fromgray-level co-occurrence matrices.

[0029] FIGS. 6G and 6H are example GLCM texture homogeneity maps computed fromgray-level co-occurrence matrices.

[0030] FIGS. 6I and 6J are example GLCM texture energy maps computed from gray-level co-occurrence matrices.

[0031] FIG. 7A is a block diagram of an example lesion detector.

[0032] FIG. 7B is a block diagram of an example lung lesion detector.

[0033] FIG. 7C is an illustration of an example patch extraction.

[0034] FIG. 8A is a diagram of an example U-Net.

[0035] FIG. 8B is a diagram of an example encoder of a U-Net.

[0036] FIG. 8C is a flow diagram of an example functioning of an encoder of a U-Net.

[0037] FIG. 8D is a diagram of an example decoder of a U-Net.

[0038] FIG. 8E is a flow diagram of an example functioning of a decoder of a U-Net.

[0039] FIG. 9 is a block diagram of an example disease stager / grader.

[0040] FIG. 10 is a decision tree diagram showing functioning of an example diseaseseverity stager.

[0041] FIG. 11 is a diagram showing functioning of an example confidence valuecalculator.

[0042] FIG. 12A is a set of example features metric values distribution graphs prior toany adjustment or normalization of the features metric values.

[0043] FIG. 12B is the set of example features metric values distribution graphs of FIG.12A with volume features normalized for body characteristics.

[0044] FIG. 12C is the set of example features metric values distribution graphs of FIG.12B with all feature values Anscombe transformed.

[0045] FIG. 12D is the set of features metric values distribution graphs of FIG. 12C withall features z-score standardized.IUIC-00155

[0046] FIGS. 13A-13G are feature importance score graphs for different classificationmodels as computed using minimum redundancy maximum relevance.

[0047] FIGS. 14A-14G are feature importance score graphs for different classificationmodels as computed using ReliefF.

[0048] FIGS. 15A-15D are confusion matrices for different classification models.

[0049] FIG. 16A is a block diagram of an example stage grader.

[0050] FIG. 16B is an example stage grading scale.

[0051] FIG. 17 is a flow diagram of an example method for lumen-related diseasediagnosis and assessment.

[0052] FIG. 18 is a block diagram of an example processing system for executing acomputer-implemented lumen-related disease diagnosis and assessment method. DETAILED DESCRIPTION

[0053] Lumen-related diseases include any diseases affecting lumens, such as bloodvessels of the circulatory system or airways of the lungs. Conventional diagnosis and staging of lumen-related disease such as CF can require manual analysis by highly trained experts of high- resolution, calibrated, and reproducible 3D computed tomography (CT) imaging datasets. The experts can spend substantial time and effort to achieve accurate disease assessment. Conventional diagnosis and staging is therefore costly and time-consuming. Taking CF as an example, no single unified metric is used for disease staging, infecting the assessment process with subjectivity that can result in disease progression evaluations that can differ greatly between clinicians at different evaluation sites or even between clinicians at the same evaluation site. Moreover, the specialization needed to perform lumen-related disease assessments means that such assessments can only be performed by relatively few accredited clinical sites and cannot be performed by the larger number of general-practice sites, clinics, and hospitals.

[0054] As used in this description, “lumen” refers broadly to the internal space formed bya hollow body that may be filled with a gas, a liquid, or a mixture of both. Examples of a hollow body that defines a lumen are any branching anatomical network of tubes, vasculature or tubules, and can include, for example and without limitation, blood vessels, lymph vessels, interstitium channels, gastrointestinal structures, and pulmonary airways. It is envisioned herein that any hollow-bodied anatomical structure forming a lumen could be assessed and / or analyzed using the methods and systems disclosed herein, including measuring the luminal volume of that hollow-IUIC-00155 bodied anatomical structure. As used in this description, “patient” can refer to a human patient or animal patient or subject.

[0055] As used herein, “lesion” refers to damaged tissue or cells on at least one wall ofthe hollow body that forms a lumen.

[0056] Systems, computer-implemented methods, and non-transitory computer-readablemedia described herein can assist in, partially automate, and standardize diagnosis and assessment of lumen-related disease, while reducing costs associated with such diagnosis and assessment. CF and Alzheimer’s disease are two examples of lumen-related diseases to which the systems, computer-implemented methods, and non-transitory computer-readable media described herein may be employed. Diagnosis of a lumen-related disease can include determination of whether the disease is present or not. Assessments of the lumen-related disease can include staging and grading of the disease to quantify the level of severity or progression of the disease. Because this quantification, as provided by the systems, computer-implemented methods, and non-transitory computer-readable media described herein, can be based on trained machine-learning (ML) models that can be standardized and used by a variety of clinical sites, diagnoses and assessments can be freed from qualitative clinical opinion and thus can be made objective and repeatable, and thus more stable and certain, with reduced or eliminated interclinician or intersite discrepancy and with reduced or eliminated human error. Accordingly, the systems, computer-implemented methods, and non-transitory computer-readable media described herein can provide accurate, reliable, objective, timely, and repeatable preliminary assessment of the stage, severity, and extent of CF or other lumen-related disease that can be used to aid physicians and / or suggest or automate modification or initiation of treatments or therapies.

[0057] The systems, computer-implemented methods, and non-transitory computer-readable media described herein can be implemented, for example, as an extensible software application that can be used by radiologists in routine clinical environments. The extensible software application can include a framework for lung segmentation, the extraction and skeletonization of airway trees, and the generation of features or bio-markers, coupled with a deep learning model capable of generating a preliminary assessment of the stage, severity, and extent of the lumen-related disease.IUIC-00155

[0058] FIG. 1A illustrates an example system 100 for diagnosis of lumen-related disease.As examples, system 100 can be implemented using computing circuitry, as in one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or as one or more general-purpose computer processors, alone or in combination with special- purpose computing hardware such as one or more graphical processing units (GPUs), tensor processing units (TPUs), or artificial intelligence processing units (AIPUs) configured to execute the illustrated components 106, 108, 110, 112, 114 as software components, services, objects, routines, or functions. The example system 100 is configured to provide a binary disease diagnosis output 120 indicating whether or not the disease is present based on the provided volume data input 104. The system 100 includes volume data pre-processor 106, lumen analyzer 108, texture analyzer 110, lesion detector 112, and one or more disease diagnosis models 114. Volume data 104 provided to the system 100 is pre-processed by pre-processor 106 and is also provided to lesion detector 112. The pre-processed volume data from the pre-processor 106, can, for example, be of the form of one or more tissue or organ masks. The pre-processed volume data from the pre-processor 106 is provided to the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, outputs of each of which, or metric values derived therefrom, can subsequently be provided as inputs to one or more ML models trained as one or more disease diagnosis models 114. The one or more disease diagnosis models 114 process these inputs to classify a binary disease state as one of two possible states, disease or no disease. The diagnosis system 100 can thus provide as its output the disease diagnosis 120. As an example, the diagnosis 120 can be printed to a display coupled to a computing device or can be printed in a report, such as a digital or paper document. The diagnosis 120 can be used to suggest, initiate, or modify a treatment or therapy. The treatment or therapy initiation or modification can be executed or ordered either manually by a physician or as part of an automated treatment or therapy system in which system 100 is integrated.

[0059] The volume data 104 provided as input to system 100 can take a variety of forms.In some examples, the volume data 104 is a set of CT scan images representative of two- dimensional cross-sectional images (slices) of one or more three-dimensional anatomical structures. In other examples, the volume data 104 is a set of magnetic resonance imaging (MRI) slices. In still other examples, the volume data 104 is voxel format data. In still other examples, the volume data 104 is mesh geometry data. The volume data 104 may be derived, in differentIUIC-00155 examples, from imaging modalities other than CT or MRI, such as ultrasound or other echocardiography, cone beam computerized tomography (CBCT), micro computerized tomography (MCT), 3D laser scanning, structured light technique, stereophotogrammetry or 3D surface imaging systems (3dMD), 3D facial morphometry (3DFM), tuned-aperture computed tomography (TACT), positron emission tomography (PET), single photon emission computed tomography (SPECT), mammography tomosynthesis, functional MRI (fMRI), optical coherence tomography (OCT), or diffusion tensor imaging (DTI). In various examples, the volume data 104 can be a raw DICOM volume, where DICOM (Digital Imaging and Communications in Medicine) is an international standard for medical images and related information that defines formats for medical images. In other examples, the volume data 104 can be in the NIFTI (Neuroimaging Informatics Technology Initiative) file format (.nii). In addition to image data, the volume data 104 may have associated with it various metadata providing information about imaging modality, the imaging device used to acquire the imagery, resolution, date, time, patient, or other information.

[0060] The pre-processor 106 is configured to extract luminal tissue or organ(s) ofinterest from the volume data 104 and, in some examples, to separate the extracted luminal tissue or organ(s) of interest into distinct portions. Luminal tissue or organ(s) of interest can include any tissue or organs including lumen, such as lungs, arteries, veins, capillaries, or intestines. For example, where the luminal tissue of interest includes the airways of the lungs, the pre-processor 106 can be configured to extract the lungs from the volume data 104 and to separate the lungs into left and right lungs. The output of the pre-processor 106 can take the form of one or more masks that can be applied to the volume data 104 to separate the volume data 104 into distinct 3D spatial portions, or of data extracted from the volume data 104 representative of data within the one or more masks.

[0061] The lumen analyzer 108 can be configured to count, from the volume data 104, anumber of visible lumens, such as blood vessels or airways, and generation splits. The lumen analyzer 108 can further be configured to measure the diameters of the lumens, such as blood vessels or airways, visible in the volume data 104. The lumen analyzer 108 can further be configured to calculate lumen volume (i.e., luminal volume), such as lung and airway volume, based on the volume data 104. Because volume differences in airways can be used to diagnoseIUIC-00155 early stages of CF, metrics based on airways as may be derived using the lumen analyzer 108 can be useful in determining the diagnosis output 120.

[0062] The texture analyzer 110 can be configured to calculate texture metrics from thevolume data 104. For example, in the context of CF diagnosis, texture analysis techniques performed by the texture analyzer 110 can be used to quantize lung tissue density spatial differences that exist in CF patients, providing a quantitative measurement of what in conventional assessments may only be a qualitative assessment of volume image cloudiness.

[0063] The lesion detector 112 can be configured to detect and quantize, from the volumedata 104, lesions, such as lung lesions. For example, CF-related lung lesions can be extracted from CT scans using the lesion detector 112.

[0064] The one or more disease diagnosis models 114 can be one or more trained MLmodels configured to determine whether disease is present or not based on the outputs of the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112. For example, the outputs of the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, provided as inputs to the one or more disease diagnosis models 114, can collectively take the form of a feature vector comprising a plurality of numeric values. In other examples, not shown, the one or more disease diagnosis models 114 can further base the disease diagnosis 120 on data from a sweat chloride test, data from a blood test, and / or data from a genetic test, which can be provided to the diagnosis system 100. One or more of the one or more disease diagnosis models 114 can, in some examples, include support vector machine (SVM) classifiers and / or adaptive boosting (AdaBoost) classifiers configured to perform classification computational methods.

[0065] FIG. 1B illustrates an example system 101 for diagnosis and assessment oflumen-related disease. System 101 is similar to system 100, except that, rather than providing a binary disease / no-disease diagnosis output 120, the one or more disease classification and grading models 116 of system 101 are configured to provide a disease state assessment 122 as an output, which can “stage” (i.e., score the severity or progression state of) a lumen-related disease. Thus, instead of being binary, the output 122 of system 101 may be a value selected from among a set of quantized levels, such as a score from 0 to 3 representative of no disease (0), mild disease (1), moderate disease (2), and severe disease (3), respectively, or more granular scores on other scales, such as 0-10, 0-100, etc.IUIC-00155

[0066] As with system 100, system 101 can be implemented using computing circuitry,as in one or more ASICs or FPGAs, or as one or more general-purpose computer processors, alone or in combination with special-purpose computing hardware such as one or more GPUs, TPUs, or AIPUs configured to execute the illustrated components 106, 108, 110, 112, 116 as software components, services, objects, routines, or functions. The system 101 includes volume data pre-processor 106, lumen analyzer 108, texture analyzer 110, lesion detector 112, and one or more disease classification and grading models 116. Volume data 104 provided to the system 101 is pre-processed (into one or more tissue or organ masks, for example) by pre-processor 106 and is also provided to lesion detector 112. The pre-processed volume data from the pre- processor 106 is provided (in the form of one or more tissue or organ masks, for example) to the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, each of which can function as described above.

[0067] The outputs of each of the lumen analyzer 108, the texture analyzer 110, and thelesion detector 112 can subsequently be provided as inputs to one or more ML models trained as one or more disease classification and grading models 116. For example, the outputs of the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, provided as inputs to the one or more disease classification and grading models 116, can collectively take the form of a feature vector comprising a plurality of numeric values. The one or more disease classification and grading models 116 process these inputs to provide a disease state assessment 122 as a stage and / or grade value. One or more of the one or more disease classification and grading models 116 can, in some examples, include SVM classifiers and / or AdaBoost classifiers configured to perform classification computational methods. The diagnosis and assessment system 101 can thus provide as its output the disease state assessment 122. As an example, the disease state assessment 122 can be printed to a display coupled to a computing device or can be printed in a report, such as a digital or paper document. The disease state assessment 122 can be used to suggest, initiate, or modify a treatment or therapy. The treatment or therapy initiation or modification can be executed or ordered either manually by a physician or as part of an automated treatment or therapy system in which system 101 is integrated.

[0068] FIG. 1C illustrates an example system 102 for diagnosis and assessment oflumen-related disease. System 102 is similar to system 101, except that, in addition to providing disease state assessment 122, the system 102 further provides as outputs annotated volume dataIUIC-00155 124, segmented diseased lumen 126, and disease etiology 128. The one or more classification and grading models 118 of system 102 are, like the one or more ML models 116 of system 101 in FIG. 1B, configured to provide a disease state assessment 122 that stages a lumen-related disease. The output 122 of system 101 may be a value selected from a set of quantized levels, as described above.

[0069] As with systems 100 and 101, system 102 can be implemented using computingcircuitry, as in one or more ASICs or FPGAs, or as one or more general-purpose computer processors, alone or in combination with special-purpose computing hardware such as one or more GPUs, TPUs, or AIPUs configured to execute the illustrated components 106, 108, 110, 112, 118 as software components, services, objects, routines, or functions. The system 102 includes volume data pre-processor 106, lumen analyzer 108, texture analyzer 110, lesion detector 112, and one or more disease classification and grading models 118. Volume data 104 provided to the system 101 is pre-processed by pre-processor 106 and is also provided to lesion detector 112. The pre-processed volume data from the pre-processor 106 is provided to the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, each of which can function as described above.

[0070] The outputs of each of the lumen analyzer 108, the texture analyzer 110, and thelesion detector 112 can subsequently be provided as inputs to one or more ML models trained as one or more disease classification and grading models 118. For example, the outputs of the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, provided as inputs to the one or more disease classification and grading models 118, can collectively take the form of a feature vector comprising a plurality of numeric values. The one or more disease classification and grading models 118 process these inputs to provide a disease state assessment 122 as a stage and / or grade value.

[0071] One or more of the one or more disease classification and grading models 118can, in some examples, include SVM classifiers and / or AdaBoost classifiers configured to perform classification computational methods. The diagnosis and assessment system 102 can thus provide as one of its outputs the disease state assessment 122. As an example, the disease state assessment 122 can be printed to a display coupled to a computing device or can be printed in a report, such as a digital or paper document. The disease state assessment 122 can be used to suggest, initiate, or modify a treatment or therapy. The treatment or therapy initiation orIUIC-00155 modification can be executed or ordered either manually by a physician or as part of an automated treatment or therapy system in which system 101 is integrated.

[0072] The diagnosis and assessment system 102 can also provide as one of its outputsannotated volume data 124. The annotated volume data 124 can include metadata that annotates the input volume data 104, for example, to show the locations and sizes of lesions detected by the lesion detector 112, and / or to show lumen diameter, volume, branching, and / or generation metrics computed by lumen analyzer 108, and / or to show texture metrics computed by texture analyzer 110. The various metadata that make up the annotations of the annotated volume data 124 can be graphical, numerical, and / or textual metadata. The annotation metadata can be, for example, selectively displayed or un-displayed either by itself or in conjunction with the volume data 104. For example, the annotation metadata can be superimposed on the volume data 104 and can be turned on or off on a display so as to give a clinician a selectable view of the metadata or various portions thereof. In the context of CF diagnosis and assessment, for example, the metadata can highlight lung lesions on CT scans of lungs.

[0073] The diagnosis and assessment system 102 can also provide as one of its outputssegmented diseased lumen 126. The segmented diseased lumen 126 can be derived from the volume data 104 and can consist of or include 3D geometry, such as mesh geometry, or voxel data representative of the diseased lumen in the volume data 104. The segmentation can be performed by lumen analyzer 108 or, in some examples, in whole or in part by one or more of the ML models 118. For example, the system 102 can be configured such that the lumen analyzer 108 performs a rough segmentation of diseased lumen and one or more of the ML models 118 performs a more refined segmentation of the diseased lumen based on the rough segmentation performed by the lumen analyzer 108.

[0074] The diagnosis and assessment system 102 can also provide as one of its outputs atleast one estimated, predicted, or proposed disease etiology 128, or multiple such etiologies. In some examples, the one or multiple estimated, predicted, or proposed disease etiologies 128 can be accompanied by confidence metrics indicating the confidence of the estimated, predicted, or proposed disease etiologies 128. The estimated, predicted, or proposed disease etiology or etiologies 128 and their respective confidence metrics, if any, can, for example, be generated by one or more of the ML models 118 trained using disease and etiology data. In some examples, the estimated, predicted, or proposed disease etiology or etiologies 128 can be based on inputIUIC-00155 data other than volume data 104, such as genetic test data. For example, in the context of CF, the etiology output 128 can include a genetic sequence indicative of one or more deleterious mutations. In some examples, the etiology output 128 can include physiological location data indicative of a anatomical site of origin of a spreading or metastasizing disease. In some examples, etiology output 128 can include a proposed pathogen, toxicant, or environmental hazard that may be a likely cause of the disease, such as a food, drug, pollutant, bacterial or viral infection, prion, or parasite.

[0075] FIG. 1D illustrates an example system 103 for diagnosis and assessment of CF.System 103 is a species of system 102, specific to the particular lumen-related disease CF. The disease state assessment output 140 of system 103 may be a value selected from a set of quantized levels, as described above with respect to disease state assessment output 122. The input lung volume data 105 can be of a variety of forms, as described above with regard to volume data 104, but generally consists of or includes imaging data of the lungs or of lung tissue. The system 103 includes lung volume data pre-processor 130, airway analyzer 132, texture analyzer 134, lung lesion detector 136, and one or more disease classification and grading models 138. Lung volume data 105 provided to the system 103 is pre-processed by pre- processor 130 to extract the lungs from the lung volume data 105 and to separate the lungs into left and right lungs, as described above with regard to pre-processor 106, and is also provided to lung lesion detector 136. The separation of lungs into left and right lungs allows the lungs to be analyzed independently, which is useful in the context of CF, because the presence or extent of a condition in one lung does not imply the same condition or condition extent in the other lung.

[0076] The pre-processed lung volume data from the lung pre-processor 130 is provided(in the form of one or more lung masks, for example) to the airway analyzer 132, the texture analyzer 134, and the lung lesion detector 136, outputs of each of which can be subsequently provided as inputs to one or more ML models trained as one or more disease classification and grading models 138. For example, the outputs of the airway analyzer 132, the texture analyzer 134, and the lung lesion detector 136 provided as inputs to the one or more disease classification and grading models 138 can collectively take the form of a feature vector comprising a plurality of numeric values. The one or more disease classification and grading models 138 process these inputs to provide a quantified disease state assessment 140.IUIC-00155

[0077] One or more of the one or more disease classification and grading models 138can, in some examples, include SVM classifiers and / or AdaBoost classifiers configured to perform classification computational methods. As an example, the disease state assessment 140 can be printed to a display coupled to a computing device or can be printed in a report, such as a digital or paper document. The disease state assessment 140 can be used to suggest, initiate, or modify a CF treatment or therapy. The CF treatment or therapy initiation or modification can be executed or ordered either manually by a physician or as part of an automated CF treatment or therapy system in which system 103 is integrated. As with systems 100, 101, and 102, system 103 of FIG. 1D can be implemented using computing circuitry, as in one or more ASICs or FPGAs, or as one or more general-purpose computer processors, alone or in combination with special-purpose computing hardware such as one or more GPUs, TPUs, or AIPUs configured to execute the illustrated components 130, 132, 134, 136, 138 as software components, services, objects, routines, or functions.

[0078] The CF diagnosis and assessment system 103 can also provide as one of itsoutputs annotated lung volume data 142. The annotated lung volume data 142 can be as described above with regard to annotated volume data 124, and can include metadata that annotates the input lung volume data 105, for example, to show the locations and sizes of lung lesions detected by the lung lesion detector 136, and / or to show airway diameter, volume, branching, and / or generation metrics computed by airway analyzer 132, and / or to show texture metrics computed by texture analyzer 134.

[0079] The diagnosis and assessment system 103 can also provide as one of its outputssegmented diseased airways 144. The segmented diseased airways 144 can be derived from the lung volume data 105 and can consist of or include 3D geometry, such as mesh geometry, or voxel data representative of the diseased airways in the lung volume data 105. The segmentation can be performed by airway analyzer 132 or, in some examples, in whole or in part by one or more of the ML models 138. For example, the system 103 can be configured such that the airway analyzer 132 performs a rough segmentation of diseased airways and one or more of the ML models 138 performs a more refined segmentation of the diseased airways based on the rough segmentation performed by the airway analyzer 132.

[0080] The CF diagnosis and assessment system 103 can also provide as one of itsoutputs at least one estimated, predicted, or proposed disease etiology 146, or multiple suchIUIC-00155 etiologies, as described above with regard to disease etiology 128 of FIG. 1C. Disease etiology output 146 of system 103 can be specific to CF etiology and can include, for example, one or more predicted or proposed gene mutations.

[0081] In each of systems 100, 101, 102, and 103, a staging or grading score output bythe system can be used to initiate or modify a treatment or therapy, or to suggest such an initiation or modification. As examples, the treatment or therapy can be a surgery or an administration of a drug. As examples, the treatment or therapy initiation or modification can be a selection of a type of therapy or a change in a dosage of a drug. For example, in the context of CF, the system can output to a display a suggestion that a particular gene therapy drug administration regimen may be indicated by the generated diagnosis or assessment.

[0082] FIG. 2A illustrates an example pre-processor 200 that can be used to implementpre-processor 106 of systems 100, 101, or 102 in FIGS. 1A, 1B, and 1C, respectively. Pre- processor 200 is configured to output one or more tissue or organ masks 214 based on volume data 104 and on an initial seed coordinate 206. Pre-processor 200 includes tissue or organ segmenter 202, major vessel extractor 204, a refinement decider 208, a structure delineator 210, and an organ or tissue separator 212. Tissue or organ segmenter 202 is configured to implement one or more segmentation methods to extract a tissue or organ of interest from the volume data 104.

[0083] The major vessel extractor 204 in FIG. 2A is configured to implement one ormore segmentation methods to extract one or more major lumen from the volume data 104. In some examples, the major vessel extractor 204 can perform the segmentation based on an initial seed coordinate 206, which can be manually supplied by a user or automatically determined, for example, via a machine-learning method. Combining (summing) the extracted tissue or organ and major vessel results in a mask. Organ or tissue separator 212 is configured to operate on the mask to divide the segmented organ or tissue into distinct component parts. Based on it being determined 208 that refinement is needed in order to perform the separation, the structure delineator 210 can perform delineation of the structures, after which the separation can be attempted or re-attempted by the organ or tissue separator 212. For example, it can be determined 208 that refinement is needed based on the organ or tissue separator 212 failing to adequately separate the organ or tissues into the desired component parts.IUIC-00155

[0084] Segmentation methods that can be used in tissue or organ segmenter 202 andmajor vessel extractor 204 include global or adaptive thresholding, region growing, watershed segmentation, edge-based segmentation, level set methods, active contour (snake) model, machine-learning-based segmentation (using convolutional neural networks, U-Net, random forests, or support vector machines), atlas-based segmentation, fuzzy C-means clustering, graph cut segmentation, deep learning-based segmentation (using 3D convolutional neural networks and attention mechanisms), or hybrid methods that combine two or more of the above segmentation methods.

[0085] FIG. 2B illustrates an example lung pre-processor 201, which is a species of theexample pre-processor 200 of FIG. 2A that can be used to implement lung pre-processor 130 of system 103 of FIG.1D. Lung pre-processor 201 is configured to output one or more lung masks 226 based on lung volume data 105 and on an initial seed coordinate 206. Example lung volume data is shown as lung volume data 230 in FIG. 2C. Lung pre-processor 201 includes lung segmenter 216, trachea extractor 218, refinement decider 220, thoracic structure delineator 222, and lung separator 224. Lung segmenter 216 is configured to implement one or more segmentation methods to extract the lungs from the lung volume data 105. Example extracted lungs are shown as lungs 240 in FIG.2D. Segmentation methods such as those listed above can be used by lung segmenter 216 to perform the lung segmentation.

[0086] Similarly, the trachea extractor 218 is configured to implement one or moresegmentation methods to extract the trachea from the lung volume data 105. An example extracted trachea is illustrated as trachea 250 in FIG. 2E. Segmentation methods such as those listed above can be used by trachea extractor 218 to perform the trachea segmentation. In some examples, the trachea extractor 218 can perform the segmentation based on an initial seed coordinate 206, which can be manually supplied by a user or automatically determined, for example, via a machine-learning method. Combining (summing) the extracted trachea and lungs results in a tracheobronchial mask, an example of which is illustrated as tracheobronchial mask 260 in FIG. 2F. Lung separator 224 is configured to operate on the tracheobronchial mask to divide the lungs into left and right lungs, thus separating the lungs by side. An example of the lungs so separated is illustrated as separated lung data 270 in FIG. 2G. Based on it being determined 220 that refinement is needed in order to perform the left / right lung separation, the thoracic structure delineator 222 can perform delineation of the thoracic structures, after whichIUIC-00155 the lung separation can be attempted or re-attempted by the lung separator 224. For example, it can be determined 220 that refinement is needed based on the lung separator 224 failing to adequately separate the left and right lungs.

[0087] FIG. 3A illustrates an example lumen analyzer 300 that can be used to implementlumen analyzer 108 of systems 100, 101, or 102 in FIGS.1A, 1B, and 1C, respectively. Lumen analyzer 300 is configured to output volume, branch, and generation data 312 and lumen diameters and areas 314 based on volume data 104 and on the one or more tissue or organ masks 214. Lumen analyzer 300 includes potential lumen voxel detector 302, lumen tree extractor 304, lumen skeletonizer 306, lumen measurer 308, and lumen ray-caster 310. Lumen analyzer 300 is configured to implement a tree extraction method and a skeletonization method to extract and skeletonize a tree from the volume data 104 as segmented or masked by the one or more tissue or organ masks 214, and then to compute and output various metrics from the lumen tree. Potential lumen voxel detector 302 is configured to perform detection of potential lumen voxels in the volume data 104 as segmented or masked by the one or more tissue or organ masks 214. Lumen tree extractor 304 is configured to extract a lumen tree from the detected voxels. Lumen skeletonizer 306 is configured to skeletonize the lumen tree.

[0088] Skeletonization methods that can be used by lumen skeletonizer can includethinning algorithms, such as sequential thinning, as with the Zhang-Suen algorithm or the Guo- Hall algorithm, or 3D medial axis transform (MAT), using methods like Euclidean distance transform (EDT) and Voronoi diagrams; distance transform-based methods, such as distance field skeletonization; graph-based approaches, such as graph cuts or voxel-based graphs; topological methods, such as homotopy-based skeletonization; or deep learning-based methods, as with 3D convolutional neural networks, or PointNet-based methods.

[0089] Lumen measurer 308 is configured to compute one or more volumes 312 of thelumen tree and to count branches and generations 312 of the lumen tree based on the extracted lumen tree and its skeleton. Each time, along a lumen path, a hollow body is split forming two or more lumens, that split can be counted as a new generation. In the context of CF, in which the lumens of interest are the airways of the lungs, there can be an increased number of detected generations as the disease progresses. This is because generations that are usually below imaging scanner resolution can become more pronounced and detectable with disease severity asIUIC-00155 airways and bronchial walls dilate. Lumen ray-caster 310 is configured to perform ray-casting of the lumen tree to measure lumen diameters and areas 314.

[0090] FIG. 3B illustrates an example airway analyzer 301, which is a species of theexample lumen analyzer 300 of FIG. 3A that can be used to implement airway analyzer 132 of system 103 of FIG. 1D. Airway analyzer 301 is configured to output volume, branch, and generation data 326 and airway diameters and areas 328 based on lung volume data 105 and on the one or more lung masks 226. Airway analyzer 301 includes potential airway voxel detector 316, airway tree extractor 318, airway skeletonizer 320, airway measurer 322, and airway ray- caster 324. Airway analyzer 301 is configured to implement a tree extraction method and a skeletonization method to extract and skeletonize a tree from the lung volume data 105 as segmented or masked by the one or more lung masks 226, and then to compute and output various metrics from the airway tree. An example of a lungs mask of the one or more lung masks 226 is shown as lungs mask 330 of FIG.3C.

[0091] Potential airway voxel detector 316 of FIG. 3B is configured to perform detectionof potential airway voxels in the lung volume data 105 as segmented or masked by the one or more lung masks 226. An example 3D rendering of potential airway voxels detected from the example lung volume data 230 as masked by the lungs mask 330 is shown as voxel data 340 of FIG. 3D. Airway tree extractor 318 is configured to extract an airway tree from the detected voxels. An example 3D rendering of an airway tree derived from the voxel data 330 is shown as airway tree 350 of FIG.3E. Airway skeletonizer 320 is configured to skeletonize the lumen tree. An example 3D rendering of an airway skeleton skeletonized from the airway tree 350 is shown as airway skeleton 360 of FIG. 3F. The airway skeletonizer 320 can use any of the skeletonization methods described above, or a hybrid or combination of such methods.

[0092] Airway measurer 322 is configured to compute one or more volumes 326 of theairway tree and to count branches and generations 326 of the airway tree based on the extracted airway tree and its skeleton. Airway ray-caster 328 is configured to perform ray-casting of the extracted airway tree to measure airway diameters and areas 328.

[0093] FIG. 4A illustrates an example texture analyzer 400 that can be used to implementtexture analyzer 110 of systems 100, 101, or 102 in FIGS.1A, 1B, and 1C, respectively. Texture analyzer 400 is configured to generate regional texture metrics 420 and, in some examples, texture maps 418 based on volume data 104, the one or more tissue or organ masks 214, and aIUIC-00155 tissue or organ atlas 406. Regional texture metrics 420 can include contrast, correlation, energy, entropy, homogeneity, and maximum probability. The tissue or organ atlas 406 provides a common spatial anatomical reference and can be generated, for example, by applying registration to align the 3D volumetric data of corresponding tissue or organ data from control cases and then taking a spatial average of the registered tissue or organ volumes. Volume data 104 can be masked by tissue or organ masks 214 using masker 404.

[0094] Texture analyzer 400 includes tissue or organ registrar 408 and a GLCM / featurescalculator 410. The GLCM / features calculator 410 can include a GLCM obtainer 412, a Haralick feature obtainer 414, and a GLCM texture map obtainer 416. Tissue or organ registrar 408 is configured to implement a registration method to spatially register the tissue or organ volume data 104 to the common spatial anatomical reference of the tissue or organ atlas 406. The registration method can be a single-phase deformable registration or, as described below with regard to FIG. 5A, a two-phase registration method that involves an affine registration followed by a deformable registration. GLCM and Haralick features can thereafter be obtained from the atlas-registered tissue or organ data with the GLCM / features calculator 410. These features can form parts of the feature vector provided to the one or more disease classification and grading models 114, 116, or 118.

[0095] GLCMs capture spatial relationships of pixel or voxel intensities in an image orvolume. Haralick features, also called texture features or texture descriptors, are a set of statistical measures used in texture analysis to quantify the characteristics of textures in an image. Haralick features can be derived from GLCMs, and can include such features as contrast, which reflects local variations in intensity by measuring the differences in intensity between a pixel or voxel and its neighbors; correlation, which describes the linear dependency of pixel or voxel intensities; variance, which represents the amount of variation or heterogeneity; energy (also referred to as angular second moment), which measures uniformity of texture; entropy, which measures the amount of disorder or randomness in the texture; homogeneity, which measures the closeness of the distribution of elements in the GLCM to the GLCM diagonal; dissimilarity, which quantifies the average difference in intensity between pairs of pixels; and maximum probability, which provides information about the most dominant texture pattern or pair of intensity values within a specified spatial relationship. Haralick feature obtainer 414 can be configured to compute Haralick features from GLCMs produced by the GLCM obtainer 412.IUIC-00155 In some examples, only a select subset of Haralick features are computed. For example, in some examples, only contrast, energy, and homogeneity are computed and subsequently used in a feature vector used by one or more ML models for severity staging.

[0096] FIG. 4B illustrates an example texture analyzer 401, which is a species of theexample texture analyzer 400 of FIG. 4A that can be used to implement texture analyzer 134 of system 103 of FIG.1D. Texture analyzer 401 is configured to generate regional texture metrics 440 and, in some examples, lung texture maps 438 based on lung volume data 105, the one or more lung masks 226, and a lungs atlas 426. Regional texture metrics 440 can include contrast, correlation, energy, entropy, homogeneity, and maximum probability. Lung volume data 105 can be masked by lung masks 226 using masker 424.

[0097] CF detection and analysis that functions across different patients can require acommon spatial anatomical reference, as may be provided by a lungs atlas 426. Lungs present different shapes and volumes due to biological and lifestyle factors. Combined with the differences in scanner resolutions, it can be a challenge to directly compare lung scans from different patients. A volumetric lungs atlas 426 of a healthy population can be created using, for example, 100 segmented pairs of lungs of control cases, by applying registration to align the 3D volumetric data of the lung pairs and then taking a spatial average of the registered lung volumes. This lungs atlas 426 can facilitate the comparison between cases in a common spatial reference. The lungs atlas 426 can be made particular for pediatric or adult populations, or any other type of specialized population, through selection of the control cases used to create the atlas 426.

[0098] Texture analyzer 401 includes lungs registrar 428 and a GLCM / features calculator430. The GLCM / features calculator 430 can include a GLCM obtainer 432, a Haralick feature obtainer 434, and a GLCM texture map obtainer 436. Lungs registrar 301 is configured to implement a registration method to spatially register the lung volume data 105 to the common spatial anatomical reference of the lungs atlas 426. The registration method can be a single- phase deformable registration or, as described below with regard to FIG. 5B, a two-phase registration method that involves an affine registration followed by a deformable registration. GLCM and Haralick features can thereafter be obtained from the atlas-registered lungs data with the GLCM / features calculator 430. These features can form parts of the feature vector provided to the one or more disease classification and grading models 138.IUIC-00155

[0099] Haralick feature obtainer 434 can be configured to compute Haralick featuresfrom GLCMs produced by the GLCM obtainer 432. In some examples, only a select subset of Haralick features are computed. For example, in some examples, only contrast, energy, and homogeneity are computed and subsequently used in a feature vector used by one or more ML models for severity staging.

[0100] FIG. 5A illustrates an example tissue or organ registrar 500 that can be used toimplement tissue or organ registrar 408 of texture analyzer 400 of FIG. 4A. Tissue or organ registrar 500 is configured to implement a registration method to spatially register the tissue or organ volume data 104 to the common spatial anatomical reference of the tissue or organ atlas 406. As illustrated in FIG. 5A, the registration process used by the example tissue or organ registrar 500 incorporates two separate registration methods, a 3D affine registration as a course registration and a deformable registration for fine-tuning of the registration, the latter of which can use a total variation (TV) method. This two-phase registration approach exhibits improved accuracy over single-phase registration approaches, while reducing the overall time required to generate registration results. Tissue or organ registrar 500 includes 3D affine registrar 504, registration matrix applier 506, 3D deformable registrar 508, and deformation field applier 510. The two-phase registration of the lungs registrar 501 is performed separately 516 for each separately masked tissue or organ portion.

[0101] 3D affine registrar 504 of FIG. 5A is configured to perform 3D affine registrationof separately masked tissue or organ volume data (tissue or organ volume data 104 as masked by a mask of tissue or organ masks 214 using masker 404) to register the separately masked tissue or organ data to the common spatial anatomical reference of the tissue or organ atlas 406. 3D affine registrar 504 can use a minimization method, such as least squares minimization, where the sum of squared differences between corresponding points is minimized, to optimize (within a tolerance or convergence criterion) the parameters of the affine transformation to minimize a cost or error metric. Once an optimized affine registration matrix for a separately masked tissue or organ is found by the 3D affine registrar 504, the registration matrix applier 506 can apply the same affine registration matrix to the corresponding tissue or organ mask. The affine-registered tissue or organ volume data is passed on to the 3D deformable registrar 508 and the corresponding affine-registered tissue or organ mask is passed on to the deformation field applier 510.IUIC-00155

[0102] 3D deformable registrar 508 of FIG. 5A is configured to perform 3D deformableregistration of the already affine-registered tissue or organ volume data (tissue or organ volume data 104 as masked by a mask of tissue or organ masks 214 and as processed by 3D affine registrar 504) to further register the tissue or organ volume data to the common spatial anatomical reference of the tissue or organ atlas 406. 3D deformable registrar 508 can, for example, use a TV method to apply non-rigid, non-linear local deformations to fine-tune the registration. Once an optimized deformation field for a separately masked tissue or organ is found by the 3D deformable registrar 508, the deformation field applier 510 can apply the same deformation field to the corresponding affine-registered tissue or organ mask. The registered tissue or organ volume data is provided as a registered tissue or organ output 512 of the tissue or organ registrar 500 and the corresponding registered tissue or organ mask is provided as a tissue or organ mask output 514 of the tissue or organ registrar 500.

[0103] FIG. 5B illustrates an example lungs registrar 501, which is a species of theexample tissue or organ registrar 500 of FIG. 5A that can be used to implement lungs registrar 428 of texture analyzer 401 of FIG. 4B. Lungs registrar 501 is configured to implement a registration method to spatially register the lung volume data 105 to the common spatial anatomical reference of the lungs atlas 426. As illustrated in FIG. 5B, the registration process used by the example lungs registrar 501 incorporates two separate registration methods, a 3D affine registration as a course registration and a deformable registration for fine-tuning of the registration, the latter of which can use a TV method. This two-phase registration approach exhibits improved accuracy over single-phase registration approaches, while reducing the overall time required to generate registration results. Lungs registrar 501 includes 3D affine registrar 518, registration matrix applier 520, 3D deformable registrar 522, and deformation field applier 524. The two-phase registration of the lungs registrar 501 is performed separately 516 for each lung (left / right).

[0104] 3D affine registrar 518 is configured to perform 3D affine registration of left orright lung volume data (lung volume data 105 as masked by a left or right mask of lung masks 226) to register the left or right lung volume data to the common spatial anatomical reference of the lungs atlas 426. 3D affine registrar 518 can use a minimization method, such as least squares minimization, where the sum of squared differences between corresponding points is minimized, to optimize (within a tolerance or convergence criterion) the parameters of the affineIUIC-00155 transformation to minimize a cost or error metric. Once an optimized affine registration matrix for a left or right lung is found by the 3D affine registrar 518, the registration matrix applier 520 can apply the same affine registration matrix to the corresponding left or right lung mask. The affine-registered left or right lung volume data is passed on to the 3D deformable registrar 522 and the corresponding affine-registered left or right lung mask is passed on to the deformation field applier 524.

[0105] 3D deformable registrar 522 is configured to perform 3D deformable registrationof the already affine-registered left or right lung volume data (lung volume data 105 as masked by a left or right mask of lung masks 226 and as processed by 3D affine registrar 518) to further register the left or right lung volume data to the common spatial anatomical reference of the lungs atlas 426. 3D deformable registrar 522 can, for example, use a TV method to apply non- rigid, non-linear local deformations to fine-tune the registration. Once an optimized deformation field for a left or right lung is found by the 3D deformable registrar 522, the deformation field applier 524 can apply the same deformation field to the corresponding affine-registered left or right lung mask. The registered left or right lung volume data is provided as a registered lung output 526 of the lungs registrar 501 and the corresponding registered left or right lung mask is provided as a lung mask output 528 of the lungs registrar 501.

[0106] FIG. 6A illustrates an example GLCM / features calculator 600 that can be used toimplement GLCM / features calculator 410 of FIG. 4A. Gray-level co-occurrence matrices (GLCM) can be used to represent the correlation between pixels in an image. Haralick features are statistical texture descriptors that can be derived from GLCM. The GLCM / features calculator 600 is configured to perform GLCM calculations and Haralick features calculations on each slice of the registered tissue or organ data 512. For each slice 602, the GLCM calculator 604 calculates GLCM for a given direction vector d and distance k between pixels (or voxels) in the slice. As an example, the GLCM calculator 604 can compute the GLCM in accordance with method 650 described below with regard to FIG.6C. As an example, a good maximum value for k is k = 20. For each slice 602, the Haralick features calculator 606 calculates Haralick features from each GLCM. With GLCM and Haralick features calculated from each slice, texture map constructor 608 can combine the computed gray-level co-occurrence matrices to construct GLCM texture maps for different feature metrics, such as contrast, homogeneity, and energy. Texture feature calculator 610 can calculate a texture feature based on the constructed textureIUIC-00155 map. For example, texture feature calculator 610 can be configured to compute a geometric mean of a texture map to arrive at a condensed texture feature. As an example, the texture feature can be computed as: 1 N m ure1mText(1)Texture features can further be z- 610. The constructed texture map can be output as texture maps output 612. Example texture maps are shown in FIGS. 6D through 6J. The calculated texture features can be output as regional texture metrics 614, which can correspond to regional texture metrics 420 of FIG.4A or to regional texture metrics 440 of FIG.4B. FIG.6B illustrates that the example GLCM / features calculator 600 described above can also be used to implement the GLCM / features calculator 430 of FIG.4B by supplying the calculator 600 with the registered lung data 526 as its input.

[0107] The flow diagram of FIG. 6C illustrates an example method 650 of GLCMcalculation that can be used by GLCM calculator 604 to create a GLCM for parameters (d, k), where d is a direction vector and k is a distance between pixels (or voxels). A registered image or volume I is received 652 and a gray-level co-occurrence matrix M is initialized 654 as an empty matrix of size L × L, where L is the number of gray levels of the image or volume I. Then, iteratively or in parallel for each voxel v in image or volume I, the following process is performed 656. A location vector w is set 658 to the coordinates of a voxel with distance k and direction d from voxel v. A first GLCM location index a is set 660 to the gray-level value of image or volume I at voxel v. A second GLCM location index b is set 662 to the gray-level value of image or volume I at voxel w. Based on a and b not being empty, the value of the GLCM at indices a, b M(a,b) is set to the value M(a,b) + 1. After value assignments 658, 660, 662, and 664 are performed 656 for each voxel v in image or volume I, the GLCM M is returned 666.

[0108] FIG. 6D shows an example GLCM texture contrast map 670 for a left lung of aCF patient. FIG.6E shows an example GLCM texture contrast map 672 for a right lung of a CF patient. FIG. 6F shows an example GLCM texture contrast map 674 for a left lung of a controlIUIC-00155 (healthy) patient. FIG. 6G shows an example GLCM texture homogeneity map 676 for a left lung of a CF patient. FIG. 6H shows an example GLCM texture homogeneity map 678 for a right lung of a CF patient. FIG.6I shows an example GLCM texture energy map 680 for a left lung of a CF patient. FIG.6J shows an example GLCM texture energy map 682 for a right lung of a CF patient. Thirteen graphs of the respective texture map are shown in each of FIGS. 6D, 6E, 6F, 6G, 6H, 6I, and 6J. Each graph corresponds to a different direction d. Each feature metric entry (contrast entry, homogeneity entry, or energy entry) on each graph is for a value of distance k (ranging from 0 to 20) for a particular slice of the volume (ranging from 0 to 400). Contrast entry values range from 0 to 100 in the CF cases and from 0 to 50 in the control case. Homogeneity and energy entry values range from 0 to 1.

[0109] Because, in the illustrated examples, GLCMs can generate six different featuremetrics (contrast, homogeneity, energy, maximum probability, correlation, and entropy for each lung), six different texture maps can be generated for each lung, each with thirteen directions. The number of three-dimensional graphs generated in a disease assessment can thus be 2 × 6 × 13 = 156. However, for the case of CF, only three of the metrics (contrast, homogeneity, and energy) show relationship to disease severity. Thus, the number of three-dimensional graphs generated in a CF assessment example can be reduced to 2 × 3 × 13 = 78. For any of the features metrics, differences between left and right lungs of a same patient may be relatively small, as shown by comparison of left lung CF texture contrast map 670 of FIG. 6D with left lung CF texture contrast map 672 of FIG. 6E. However, comparisons of metrics of a left lung of a CF patient to a left lung of a control (healthy) patient, or a right lung of a CF patient to a right lung of a control (healthy) patient, will yield differences that are comparatively large, as shown, for example, by comparison of left lung CF texture contrast map 670 of FIG. 6D with left lung control texture contrast map 674 of FIG.6F.

[0110] The three metrics derived from GLCMs in the CF assessment example (contrast,homogeneity, and energy) can be related to changes in lung density and the presence of lesion patterns, which in turn are related to CF severity. The geometric means of each of the texture maps can be computed in order to condense the data from the maps in to a smaller, more manageable set of values. The geometric means can be computed, for example, using Equation 1, above. Table 1 shows the results of example geometric mean computations for control (healthy), mild CF, moderate CF, and severe CF disease states for each of the metrics of textureIUIC-00155 contrast, texture homogeneity, and texture energy for right lung, left lung, and combined right and left lungs. TABLE 1: Average texture metrics results (raw) Metric Control Mild Moderate Severe Texture contrast 5.8701 ± 5.9140 ± 7.7979 ± 10.9567 ±

[0111] Table 1 shows, for example, that contrast increases as disease progresses, whereashomogeneity and energy decrease with disease severity. These trends hold irrespective of lung (right / left). The raw geometric mean values of Table 1 can be transformed using the Anscombe transform and then z-scored to arrive at the values shown in Table 2, which shows the above- described trends even more clearly than in Table 1. The Anscombe transform is described in greater detail below with regard to Equation 3.IUIC-00155 TABLE 2: Average texture metrics results (Anscombe transformed and z-scored) Metric Control Mild Moderate Severe Texture contrast −0.6319 ± -0.4598 ± 0.1298 ± 1.7873 ±

[0112] FIG. 7A shows an example lesion detector 700 that can be used to implementlesion detector 112 of systems 100, 101, or 102 in FIGS. 1A, 1B, and 1C, respectively. Lesion detector 700 produces a detected lesions mask 714, indicating the locations and sizes of detected lesions in the volume data 104, and from which lesions metrics can be computed. Detection of lesions across a lumen volume, such as a lung volume, can aid determination of the extent and stage of disease severity, such as CF severity. Conventionally, the task of lesions detection is performed manually and coarsely. Two U-Nets can be used to automate the task of lung lesions detection. For example, manually annotated volumetric data sets, such as CT scans, can be used to train, validate, and test a two-U-Net model. For example, one hundred or more manually annotated CT scans can be used to train, validate, and test the two-U-Net model. Lesion detector 700 can include patches extractor 702, binary class 3D U-Net 704, multi-class 3D U-Net 706, patches merger 708, and pixel-wise multiplier 710. Volume data 104 is provided to patches extractor 702, which is configured to extract patches from the volume data. An example of patch extraction is described below with regard to FIG.7C.

[0113] Extracted patches are provided to binary-class 3D U-Net 704 and to multi-class3D U-Net 706. A U-Net is a convolutional neural network having encoders and decodersIUIC-00155 arranged in a “U” shape, as shown and described in greater detail with regard to the example of FIG. 8A. The binary-class 3D U-Net 704 is configured to classify disease or no-disease states for individual voxels of the extracted patches of the volume data 104. The multi-class 3D U-Net 706 is configured to classify different types of lesions for labeling, such as airway disease, cyst / bullae, mucus plug, parenchymal disease, and tree-in-bud, in the extracted patches of the volume data 104. Patches merger 708 is configured to merge classified patches into a volume the size of the original volume data 104. Pixel-wise multiplier 710 is configured to multiply the re-merged, lesions-classified volume with the tissue or organ mask(s) 214 to produce the detected lesions mask 714 as the output of the lesion detector 700.

[0114] FIG. 7B shows an example lung lesion detector 701 that is a species of lesiondetector 701 and can be used to implement lung lesion detector 132 of systems 103 in FIG.1D. Lung lesion detector 701 can be similar to lesion detector 700, with its U-Nets 704, 706 trained on lung lesion volume data, such as labeled CF CT volume data, and lung volume data 105 and lung mask(s) 226 provided as inputs to yield detected lung lesions mask 716 as output.

[0115] The dimensions of the CT scan or other volume scan can vary depending on thescanner specifications and the parameters chosen during the scanning process. A patch-wise approach can be implemented to make the proposed model agnostic to the dimensions of the CT scans or other volumetric data. In a patch-wise approach, the volume is divided into several sub- volumes. This type of approach has several advantages, but it can increase computational complexity. FIG. 7C shows an example patch extraction 750 as may be performed by patches extractor 702 in lesion detector 700 or 701 of FIGS. 7A and 7B, respectively. An original volume 752, which can correspond to volume 104 or 105, may be of a certain size, which, in the illustrated example, is 512 × 512 × the number of slices voxels. In patch extraction, individual smaller patches 754, 756, 758, 760 are extracted from the original volume 752, each of the individual patches being smaller than the original volume. In the illustrated example, each of the individual patches is 124 × 124 × 116 voxels.

[0116] FIG. 8A shows an example U-Net 800, so-called because the structure of dataflow in the network resembles the letter “U.” Going down the U, data of an input patch 802 is processed by a series of encoders 804, 806, 808 until a bridge 810 is reached. Going up the U, the data is processed by a series of decoders 812, 814, 816, which are followed by a classification layer 818, which produces an output 820. Skip connections couple encodersIUIC-00155 directly to decoders. In the illustrated example, first encoder 804 is coupled by a skip connection to third decoder 816, second encoder 806 is coupled by a skip connection to second decoder 814, and third encoder 808 is coupled by a skip connection to first decoder 812. FIG. 8B shows the general connection structure of an example U-Net encoder 830, having an input 832 from a previous layer, an output 834 for a skip connection, and an output 836 for the next layer. FIG. 8C shows the internal processing structure of a U-Net encoder as a flow diagram 838. An input 840 to the U-Net encoder is processed with first 3D convolution 842, first batch normalization 844, first rectification (ReLU) 846, second 3D convolution 848, second batch normalization 850, second rectification (ReLU) 852, and max pooling 854 to provide the U-Net encoder outputs 856. For the U-Net encoder 830, the number of filters for the first convolutional layer 872 is 25+D, and for the second convolutional layer 878 is 26+D.

[0117] FIG. 8D shows the general connection structure of an example U-Net decoder858, having an input 860 from a skip connection, an input 862 from a previous layer, and an output 864 for the next layer. FIG. 8E shows the internal processing structure of a U-Net decoder as a flow diagram 866. The two inputs 868 to the U-Net decoder are combined with a concatenation 870, which is then processed with a first 3D convolution 872, first batch normalization 874, first rectification (ReLU) 876, second 3D convolution 878, second batch normalization 880, second rectification (ReLU) 882, and up convolution 884 to provide the U-Net decoder output 886. For the U-Net decoder 858, the number of filters for both convolutional layers 872, 878 is 29-D. In a CF lesion detection context, the two-U-Net architecture of lesion detectors 700, 701 can classify healthy and disease voxels with an accuracy of 99.85% and an intersection over union score of 0.6643. Table 3 shows the breakdown of intersection over union scores for this architecture for the different lung lesion types. TABLE 3 Lung lesion Intersection over union 5871023IUIC-00155

[0118] Class imbalance is a concern in classification problems. This imbalance ispresent in CF-related lung lesions. Generalized Dice loss can be used to mitigate the class imbalance problem. As an example, generalized Dice loss can be defined as: 2∑KwM k∑ Y TLoss = 1 −k=1 m=1km kmKw Y2 + T2 (2)where K is the number of of Y, and wk is a class-specific weighting factor that controls the contribution each class makes to the loss. The use of generalized Dice loss to mitigate class imbalance reduces the influence of large regions while improving the performance in segmenting smaller regions. Table 4 shows example lesion data include lesion volume percentage data that can be computed from the detected lung lesions mask. TABLE 4 Lung lesion CT scans with the given Volume occupied by the lesion (%) lesions (%)

[0119] FIG. 9 shows an example disease stager / grader 900 that employs seven differentML models to each perform a classification task. In different examples, the example disease stager / grader 900 can be used to implement the disease diagnosis model(s) 114 in system 100 FIG. 1A, the disease classification and grading models 116 or 118 of systems 101 and 102 in FIGS. 1B and 1C, or the disease classification and grading model(s) 138 in system 103 of FIG. 1D. Each ML model in the example disease stager / grader 900 can output a classification decision, choosing between one of several decision options (classes), and can, in some examples, also output a confidence value associated with its classification decision. The classification and confidence outputs of the ML models can be provided to a staging and confidence calculator 910 to compute a stage of disease severity. Although the illustrated example uses seven ML modelsIUIC-00155 916, 920, 924, 928, 932, 936, 940 and combines their output, other examples can use fewer or more ML models, such as a single ML model.

[0120] The staging and confidence calculator 910 is configured to classify CF of apatient according to one of four severity stages: normal (no detectable disease), mild, moderate, and severe. An ensemble of ML models, as in the example disease stager / grader 900, can be used to assess the severity of CF from the features extracted from the above-described subsystems. The assessment can include the predicted severity stages along with a confidence value and a stage grading. The stage grading can be performed by the grader 912 of the disease stager / grader 900. As described in greater detail below with regard to FIGS. 16A and 16B, the stage grading performed by the grader 912 further refines the granularity of the output stage values.

[0121] Each ML model 916, 920, 924, 928, 932, 936, 940 in the example diseasestager / grader 900 of FIG. 9 makes a classification decision by deciding on one of two, one of three, or one of four different decision options (classes), depending on the model, on the basis of the features vector 902 input to the models. An example features vector is described in greater detail below with regard to Tables 8, 9, 10, and 11. The features of the features vector 902 can, for example, be based on the outputs of the lumen analyzer 108, the texture analyzer 110, and the lesion detector 112, in the cases of systems 100, 101, or 102, or on the outputs of the airway analyzer 132, the texture analyzer 134, and the lung lesion detector 136, in the case of system 103. In the illustrated example of FIG. 9, ML models 916, 920, 924, and 928 are two-class (binary) classification models 904; ML models 932 and 936 are 3-classes classification models 906, and ML model 940 is a 4-classes classification model 908. The classes decided between by each model are described in greater detail below with regard to Tables 6 and 7.

[0122] Each ML model 916, 920, 924, 928, 932, 936, 940 in the example diseasestager / grader 900 of FIG. 9 can be considered a weak learner with a specific classification purpose. The ensemble combination of weak learners, however, can form a robust, strong learner. Each feature’s importance varies by classification task. Features should thus, for example, not be weighted the same in a model whose goal is to detect between normal and disease states as they would be weighted in a model whose goal is to discriminate between a mild and a severe disease stage. The features employed by each model can be selected, for example, using the results obtained by the minimum redundancy maximum relevance (MRMR)IUIC-00155 and / or the ReliefF algorithm. Hyperparameters of each ML model can be obtained using Bayesian optimization. The results of the seven ML models 916, 920, 924, 928, 932, 936, 940 can be used to generate the stage classification 946.

[0123] In the example disease stager / grader 900 of FIG. 9, features selector 914 isconfigured to select features from features vector 902 to provide to first binary classification model 916; features selector 918 is configured to select features from features vector 902 to provide to second binary classification model 920; features selector 922 is configured to select features from features vector 902 to provide to third binary classification model 924; features selector 926 is configured to select features from features vector 902 to provide to fourth binary classification model 928; features selector 930 is configured to select features from features vector 902 to provide to first 3-classes classification model 932; features selector 934 is configured to select features from features vector 902 to provide to second 3-classes classification model 936; and features selector 934 is configured to select features from features vector 902 to provide to 4-classes classification model 940. For each features selector 914, 918, 922, 926, 930, 934, 938, the features to be selected can be manually pre-programmed into each respective features selector, or in some examples, one or more of the features selectors can be programmed to execute an algorithm, such as MRMR or ReliefF, to determine which features to select.

[0124] In examples, the outputs from all seven of the ML models 916, 920, 924, 928,932, 936, 940 can be provided to severity stager 942, which can be configured to output a severity stage (such as one of normal, mild, moderate or severe) for the disease based on the outputs of the seven ML models 916, 920, 924, 928, 932, 936, 940. A confidence value calculator 944 can generate a confidence value for the stage output, indicative of the confidence placed in the severity determination. This stage and confidence 946 can be supplied as outputs of the disease stager / grader 900 and of the system 100, 101, 102, or 103. In the case of diagnosis system 100, in which the output is binary (disease / no disease), any stage output other than normal can result in a “disease” diagnosis.

[0125] FIG. 10 illustrates an example severity stager 1000 that can be used to implementseverity stager 942 of disease stager / grader 904 in FIG. 9. The functioning of severity stager 1000 is illustrated as a decision tree, which outputs a disease stage as one of normal, mild, moderate, or severe based on outputs 1002, 1010, 1018 of some or all of the first, fourth, second,IUIC-00155 and third binary classification models 916, 928, 920, 924, the first and second 3-class classification models 932, 936, and the 4-class classification model 940. In the decision tree of severity stager 1000, the first binary classification model output 1002, which is a binary normal / not normal (no disease / disease) classification is queried. Based on the output 1002 indicating 1004 normal, a normal stage 1006 is reported as the output. Based on the output 1002 indicating 1004 not normal, the fourth binary classification model output 1010, which classifies as severe / not severe, is queried. Based on the output 1010 indicating 1012 severe disease, a severe state 1014 is reported as the output of the severity stager 1000. Based, however, on the output 1010 indicating 1012 that the disease stage is not severe, the outputs 1018 of the second and third binary classification models 920, 924, the first and second 3-class classification models 932, 936, and the 4-class classification model 940 are queried and these outputs 1018 are counted as votes. Based on the number of votes in these outputs 1018 for mild disease stage being greater than the number of votes in these outputs 1018 for moderate disease stage, the disease is staged as mild 1022 and the severity stager 1000 reports a mild stage as its output. Otherwise, the severity stager 1000 reports moderate disease 1024 as its output. The reported severity stage D output by severity stager 1000 can form part of the stage and confidence output 946 of disease stager / grader 900 of FIG.9.

[0126] FIG. 11 illustrates an example confidence value calculator 1100 that can be usedto implement confidence value calculator 944 in system 900 of FIG. 9. Based on the stage decision D (normal, mild, moderate, severe) 1102, as may be arrived at by severity stager 942 or 1000, indicating normal, the confidence value 1104 reported by the first binary classification model 916 is reported as the confidence value 1106. Based on the stage decision D 1102 indicating the severe disease stage, the confidence value 1108 reported by the fourth binary classification model 928 is reported as the confidence value 1106. Based on the stage decision D 1102 indicating a mild or moderate disease stage, the confidence value 1106 is computed as the average 1110 of the confidence values 1112 in the stage decision D reported by the second and third binary classification models 920, 924, the first and second 3-class classification models 932, 936, and the 4-class classification model 940. The reported confidence value 1106 output by confidence value calculator 1100 can form part of the stage and confidence output 946 in disease stager / grader 900 of FIG.9.IUIC-00155

[0127] A number of corrections and transforms can be made to the features metrics of thefeatures vector 902 prior to classification to improve classification. Volumetric measures are related to body characteristics, such as height, weight, age, and body mass index. Volumetric measures do not, however, substantially vary by CF disease stage or severity. Accordingly, features metrics of the features vector 902 may be corrected for such body characteristics prior to classification by the models 904, 906, 908. Additionally, the Anscombe transformation can be used to normalize the features values. Anscombe transformation of a value can be defined as: A= 2√value +3 (3)The features values can be standardized using z-scores to eliminate the differences in ranges and units. FIGS. 12A through 12D illustrate example features normalization for various features of the features vector 902. FIG.12A shows example raw features metric values 1200, prior to any normalization. FIG.12B shows example features metric values 1202 following normalization of volume features for body characteristics. Accordingly, only the five top-row graphs of Lung Volume, Right Lung Volume, Left Lung Volume, Lung Ratio, and Airways Volume are changed between FIGS. 12A and 12B. FIG. 12C shows example features metric values 1204 following subsequent application of an Anscombe transform. FIG. 12D shows example feature metric values 1206 following subsequent z-score standardization.

[0128] Using an excessive number of features as inputs to the classification models 916,920, 924, 928, 932, 936, 940 can have a negative impact on the performance of the classification models. Different features metrics can have a different importance level to the different classification models 916, 920, 924, 928, 932, 936, 940. This can be shown with two different feature ranking and selection algorithms, MRMR and ReliefF. The MRMR algorithm can be used to find a subset of features as inputs to the classification models that maximizes the relevance of the input features and minimizes the redundancy between predictors. An example implementation of the MRMR algorithm can be found in Ding and Peng, Minimum redundancy feature selection from microarray gene expression data, JOURNAL OF BIOINFORMATICS AND COMPUTATIONALBIOLOGY, Vol. 3, No. 2, pp. 185-205 (2005). The ReliefF algorithm can be used to estimate the relevance of the input features based on their relationship with the responseIUIC-00155 variable and between nearest-neighbor instance pairs. An advantage of the ReliefF algorithm is its ability to estimate input feature relevance without assuming conditional independence between different features. An example description of the ReliefF algorithm can be found in Kononenko, Simec, and Robnik-Sikonja, Overcoming the myopia of inductive learning algorithms with ReliefF, APPLIED INTELLIGENCE, Vol.7, pp.39-55 (1997).

[0129] The bar graphs of FIGS. 13A through 13G show the relative importance of thefeatures metrics to the respective classification models 916, 920, 924, 928, 932, 936, 940 as determined using MRMR. Graph 1300 of FIG. 13A shows MRMR importance scores of the different features metrics to the first binary classification model 916. Graph 1302 of FIG. 13B shows MRMR importance scores of the different features metrics to the second binary classification model 920. Graph 1304 of FIG. 13C shows MRMR importance scores of the different features metrics to the third binary classification model 924. Graph 1306 of FIG. 13D shows MRMR importance scores of the different features metrics to the fourth binary classification model 928. Graph 1308 of FIG. 13E shows MRMR importance scores of the different features metrics to the first 3-classes classification model 932. Graph 1310 of FIG.13F shows MRMR importance scores of the different features metrics to the second 3-classes classification model 936. Graph 1312 of FIG. 13G shows MRMR importance scores of the different features metrics to the 4-classes classification model 940.

[0130] The bar graphs of FIGS. 14A through 14G show the relative importance of thefeatures metrics to the respective classification models 916, 920, 924, 928, 932, 936, 940 as determined using ReliefF. Graph 1400 of FIG. 14A shows ReliefF importance scores of the different features metrics to the first binary classification model 916. Graph 1402 of FIG. 14B shows ReliefF importance scores of the different features metrics to the second binary classification model 920. Graph 1404 of FIG. 14C shows ReliefF importance scores of the different features metrics to the third binary classification model 924. Graph 1406 of FIG. 14D shows ReliefF importance scores of the different features metrics to the fourth binary classification model 928. Graph 1408 of FIG. 14E shows ReliefF importance scores of the different features metrics to the first 3-classes classification model 932. Graph 1410 of FIG.14F shows ReliefF importance scores of the different features metrics to the second 3-classes classification model 936. Graph 1412 of FIG. 14G shows ReliefF importance scores of the different features metrics to the 4-classes classification model 940.IUIC-00155

[0131] The four binary classification models 916, 920, 924, 928 can use support vectormachines (SVMs) to decide in favor of one of the stages against all the other stages. For example, first binary classification model 916 can be trained to make the classification decision normal / not normal, second binary classification model 920 can be trained to make the classification decision mild / not mild, third binary classification model 924 can be trained to make the classification decision moderate / not moderate, and fourth binary classification model 928 can be trained to make the classification decision severe / not severe. Table 5 shows example hyperparameters and labels associated with the example binary classification models 916, 920, 924, 928 shown in the example disease stager / grader 900 of FIG.9. TABLE 5: BINARY CLASSIFICATION MODELS HYPERPARAMETERS AND LABELS Model Algorithm Class Regularization Kernel Kernel v labels cost type coefficients

[0132] The 3-classes classification models can combine two stages, and can make theclassification decision between, as one option, the combination of the two stages and, as the other two options, the other two stages. For example, the first 3-classes classification model 932 can be trained to make the classification decision between the three options of normal + mild, moderate, and severe; and the second 3-classes classification model 936 can be trained to make the classification decision between the three options of normal, mild, and moderate + severe. The 4-classes classification model 940 can be trained to make the classification decision between each of the four classification options (normal / mild / moderate / severe) distinctly. Thus, normal isIUIC-00155 used as the outlier for mild in the first 3-classes classification model 932, and in the second 3-classes classification model 936, severe is used as the outlier for moderate, to facilitate the classification task. Table 6 shows example hyperparameters and labels associated with the example 3-classes and 4-classes classification models 932, 936, 940 shown in the example disease stager / grader 900 of FIG.9. TABLE 6: 3- AND 4-CLASS CLASSIFICATION MODELS HYPERPARAMETERS AND LABELS Model Algorithm Class labels NumberLearning Maximum of weak rate number of

[0133] Table 7 shows performance metrics per classification model. Rather than usingan ensemble of classification models, as in the example disease stager / grader 900 of FIG. 9, some examples (not shown) can use only one model, a 4-classes model. This is the equivalent of using only the 4-classes classification model 940 as the single model to stage the lumen-related disease and would mean that the accuracy of just using the one model would have an accuracy of about 81%, as shown in the second row from the bottom of Table 7. That is, it would produce accuracy equivalent to using the 4-classes classification model 940 only. However, using the ensemble of classification models can improve accuracy to near 87% (and 83% on the F1-score). The use of an ensemble of models thus improves accuracy of prediction over using a single model. The area under the receiver operating characteristic curve (area under the ROC curve, or AUC) provides an aggregate measure of performance across all possible classification thresholds, that is, the probability that the model ranks a random positive example more highly than a random negative example. Thus, a model with an AUC of 1 makes correct predictionsIUIC-00155 100% of the time, whereas a model with an AUC of 0 makes incorrect predictions 100% of the time. TABLE 7: PERFORMANCE METRICS PER MODEL Model Accuracy Recall Specificity Precision F1-score AUCFirst binary classification 1.000 1.000 1.000 1.0000 1.0000 1.000 0 1 0 2 2 2

[0134] FIGS. 15A through 15D show example confusion matrices for the classificationmodels 916, 920, 924, 928, 932, 936, 940 and the ensemble model that comprises the seven classification models used together. FIG. 15A shows example confusion matrices 1502, 1504, 1506, 1508 for the first binary classification model 916, the second binary classification model 920, the third binary classification model 924, and the fourth binary classification model 928, respectively. FIG. 15B shows example confusion matrices 1510, 1512 for the first 3-classes classification model 932 and the second 3-classes classification model 936, respectively. FIG. 15C shows an example confusion matrix 1514 for the 4-classes classification model 940. FIG. 15D shows an example confusion matrix 1516 for the ensemble model that comprises the seven classification models used together.

[0135] FIG. 16A shows an example stage grader 1600 that can be used to implementgrader 912 in the disease stager / grader 900 of FIG.9. Stage grader 1600 is configured to further stratify the staging provided by severity stager 942 and the classification models that precede it. For example, the severity grade can be used to determine if the predicted stage corresponds to (I) a case at the predicted stage's early phases; (II) a case falling within the average for the predictedIUIC-00155 stage; or (III) a case near to transition to a higher-severity stage. Stage grader 1600 can thus refine the staging, for example, in accordance with the example stage grading scale 1650 shown in FIG. 16B. In the example stage grading scale 1650, the normal, mild, moderate, and severe stages are further divided as normal, mild I, mild II, mild III, moderate I, moderate II, moderate III, severe I, and severe II grades. Stage grader 1600 can include tertiles for D selector 1606, per-feature grade calculator 1610, and stage grade calculator 1614, the latter of which outputs stage grading 1616. The features tertiles 1604 can be obtained using the features extracted from a training set. Each severity stage has its own tertile vector. The tertiles can be selected by the tertiles for D selector 1606 The per-feature grade calculator 161, which can correspond to per- feature grader 948 in FIG.9, can assign a feature grade for each feature by determining in which tertile the feature value falls. The stage grade calculator 1614 is configured to calculate the stage grade by computing an average of the per-feature grades of relevant features as calculated by the per-feature grade calculator 1610.

[0136] Tables 8 through 11 show example features vectors for a variety of patients ofdifferent disease severities. A features vector 902 for a single patient can include, as an example, an age, a lungs and trachea volume (in liters), an age-adjusted lungs and trachea volume (in liters per age), a right lung volume (in liters), an age-adjusted right lung volume (in liters per age), a left lung volume (in liters), an age-adjusted left lung volume (in liters per age), a lungs volume ratio (right / left), an airways volume (in liters), an age-adjusted airways volume (in liters per age), an airways volume percentage, branches, generations, lesions percentage, and GLCM metrics (“geometrics”) for the left lung, the right lung, and the two lungs combined: left lung contrast, left lung homogeneity, left lung energy, right lung contrast, right lung homogeneity, right lung energy, combined lungs contrast, combined lungs homogeneity, and combined lungs energy. The features vector 902 can also include a label, which is present in classifier model training data, and which is to be determined as the staging value for test data.

[0137] FIG. 17 illustrates an example method 1700 of lumen-related disease diagnosisand severity assessment. Any of the above-described systems, such as systems 100, 101, 102, and 103 of FIGS. 1A through 1D, can be used to perform the method of FIG. 17. In method 1700, volume data is provided or received 1702. The volume data can be of any of the medical imaging modality types described above, such as CT scan slices or MRI slices. As examples, the receiving 1702 can be by way of receiving a transmission of data from an external device, suchIUIC-00155 as a different computer system on a network or a scanner such as a CT scanner or MRI scanner, or can be by retrieval of the volume data from a local memory. The number of visible lumens, such as the hollow space in blood vessels or airways, and generation splits visible in the volume data are counted 1704. To aid in the counting 1704, the lumens in the volume data can be segmented and skeletonized. The diameters of the lumens, such as that defined within blood vessels or airways, visible in the volume data are measured 1706. The volume of the network of the lumens (e.g., the one or more lumen that make up a whole or a part of an organ of interest) is measured 1708 based on the volume data.

[0138] Texture metrics are calculated 1710 from the volume data. Lesions are detectedin the volume data and quantized 1712. Then, using one or more ML models, the severity stage of the lumen-related disease is determined 1714 based on one or more feature metrics derived from (a) the counted number of lumens and generation splits, (b) the measured diameters and volume, (c) the calculated texture metrics, and (d) the detected and quantized lesions. The severity stage can be further refined in granularity by the computation of grading values by the ML models or using computations derived from the outputs of the ML models. The severity stage can be displayed 1716. In some examples, disease grading, annotated volume data, segmented diseased tissue based on the differences of measured lumens, and / or disease etiology can also be displayed. In some examples, a treatment or therapy can be modified or initiated 1718 based on the displayed severity stage and / or grading.

[0139] The above-described systems, such as systems 100, 101, 102, and 103 of FIGS.1A through 1D, can operate using software that encodes instructions on one or more non- transitory computer-readable media. The computer-readable media can be read by general- purpose or special-purposes processors. In some examples, a diagnosis and / or assessment system, such as one of systems 100, 101, 102, or 103, can be defined by software instructions that are executed on one or more general-purpose or special-purposes processors. In some examples, the diagnosis and / or assessment system is provided as a cloud service that is run in the cloud. In such examples, input data, such as volume data, is uploaded to the cloud service, where it is processed in the cloud to generate the results, including a disease severity score.

[0140] FIG. 18 shows an example processing system 1800 capable of performing thecomputer-implemented method steps described herein or to execute the computer code stored on the non-transitory computer-readable media described herein. As examples, system 1800 isIUIC-00155 capable of generating any of systems 100, 101, 102, or 103 from software code. As another example, system 1800 is capable of executing method 1700, or portions thereof, as software code. The illustrated example processing system 1800 includes computer system 1802 having one or more processing units 1804, which can be one or more central processing units (CPUs), GPUs, TPUs, AIPUs, or other kinds of processing units, and having memory 1806. The one or more processing units 1804 can be coupled to a network adapter 1808 and / or to an input / output (I / O) interface 1810.

[0141] The network adapter 1808 can be communicatively coupled to a computernetwork 1812, such as an internet, intranet, or cloud via a computer network connection, which can be a wireless or wired connection. The I / O interface 1810 can be communicatively coupled to one or more external devices 1814, such as input devices, such as a keyboard, mouse, or 3D pointing device, and / or can be coupled to one or more displays 1816, such as one or more computer monitors and / or one or more 3D displays, such as a head-mounted display (HMD). Memory 1806 can include a hierarchy of memories of different speeds and storage capacities. For example, memory 1806 can include comparatively fast random access memory (RAM) 1818, an even faster but smaller-capacity cache memory 1820, and / or a relatively slower but larger- capacity system storage 1822, which can include, for example, one or more hard disk drives (HDDs) and / or one or more solid state drives (SSDs).

[0142] System storage 1822 can be configured to store a software application that, whenexecuted by the one or more processing units 1804, is configured to execute systems 100, 101, 102, or 103 and / or method 1700, as examples. As an example, one or more of the one or more processing units 1804 can execute the above-described one or more ML models to generate assessment scores, such as staging and / or grading, and / or other useful outputs, based on volume data supplied from one or more of the one or more external devices 1814 and / or via the network 1812. The assessment scores and / or the other useful outputs can be output to one or more of the one or more displays 1816.IUIC-00155 TABLE 8 No. Label Age Lungs & trachea Right lunge 147648885229627283287382552234134944IUIC-00155 TABLE 9 No. Left lung LungsAirways ratio 751566188832632983594979246757114869IUIC-00155 TABLE 10 No. Branches Generation Lesions % GLCM Metrics - GeometricsLeft lung 726973298511867497374636277812127783IUIC-00155 TABLE 11 No. GLCM Metrics - Geometrics Right lung Combined left & right lungs 156428512934346745632388561613389492

[0143] The systems, methods, and computer-readable media described herein canprovide technical improvements over conventional methods for diagnosis and assessment ofIUIC-00155 disease, such as CF. Conventional methods for diagnosis and assessment of CF can include, as examples, a sweat chloride test, a genetic test, and / or a clinical evaluation of CF from an accredited center. Lung lesions are the pulmonary manifestations of CF. In CF-affected patients, the lungs can be the most severely affected organs and are the primary cause of mortality. CT imaging of the lungs can be used to perform clinical evaluation of CF, allowing physicians to assess the severity and extent of the disease. The analysis of CT datasets of CF patients can require physicians to manually analyze and annotate different lung lesions visible on CT scan images, in many instances labeling each CT scan slice individually. Visual features of CT scan slice images that can be annotated in such a process can include atelectasis, bronchiectasis, bronchial wall thickening, mucus plugging, parenchymal opacities, and tree-in- bud features.

[0144] In addition to being time-consuming, this manual analysis and annotation toassess the stage and extent of CF is at least partially qualitative and is not perfectly repeatable, and thus can produce assessment results that can vary between experts and clinical sites. Clinically accepted scoring methods for CF staging include the Brasfield score, the Brody score, the severe advanced lung disease (SALD) four-category scoring system, and the Perth-Rotterdam annotated grid morphometric analysis for CF (PRAGMA-CF). In performing their assessments, these methods disadvantageously either evaluate a volume as a whole or evaluate only a relatively small subset of the slices of available volume data, resulting in either too coarse or insufficiently granular analysis. For example, the Brasfield score results from a sum of subjective scores made by a clinician based on a subjective assessment of an image or set of images of the lungs. As another example, PRAGMA-CF selects 10 equidistant slices from the CT volume, and for each selected slice, overlays a grid on the slice and annotates each grid cell that is covered by at least 50% of the lung field, then counts the number of cells for each annotation, computes the proportion of each annotation from the total number of annotated cells, and analyzes the proportions of the annotations to determine the stage and severity of CF. PRAGMA-CF thus omits from consideration lesions not in the selected slices and fails to consider when multiple lesions are present in a single grid cell.

[0145] In the context of CF, machine learning and deep learning techniques can be usedto predict survivorship of the patient, automate the Brasfield scoring process, and classify lung tissue, pulmonary opacities, and pleural effusions. A non-categorical method as developed byIUIC-00155 Wieying Kuo and colleagues can determine the number of abnormal airway diameters and count the number of visible airways per segmental generation. As described herein, machine learning and deep learning techniques can be used to provide rapid, accurate, and reproducible assessment of CF stage and severity using an automated classification system. The systems, computer- implemented methods, and non-transitory computer-readable media described herein can annotate and evaluate volume data, such as lung volume data, at the pixel level rather than at a grid level, providing advantageous improvements over assessment methods such as PRAGMA-CF.

[0146] The systems, computer-implemented methods, and non-transitory computer-readable media described herein can accurately and repeatably quantify the severity of lumen- related disease and its progression, providing an objective analysis that is not physician- dependent or scoring-technique-dependent. The systems, computer-implemented methods, and non-transitory computer-readable media described herein furthermore can be used by non-expert clinicians, for example, in rural hospitals that lack access to specialists, with a reduced amount of training, thus providing better and lower-cost diagnostic and assessment service to a broader range of patients, and helping to close healthcare availability and affordability gaps that exist between accredited assessment centers and other clinical sites.

[0147] The systems, computer-implemented methods, and non-transitory computer-readable media described herein furthermore can be used with pediatric patients or animal patients or subjects who cannot accurately or adequately convey verbally to a clinician their subjective experience of symptoms. Because the systems, methods, and computer-readable media described herein can use an ensemble model, they can show an increase in performance compared to using a single model to stage CF. Additionally, because the systems, computer- implemented methods, and non-transitory computer-readable media described herein can provide a confidence value on the prediction, they can help physicians to determine if further evaluation is required. Further training the classification models on clinical data, such as blood tests or and genetic tests, can improve the assessment results over using volume data alone.

[0148] The systems, computer-implemented methods, and non-transitory computer-readable media described herein can also be used or adapted to assess and stage a variety of lumen-related diseases, including CF, black lung, and COVID-19 lung. In some examples, the lesions detection portions of the described systems, methods, and computer-readable media canIUIC-00155 be adapted to detect other patterns indicative of lumen inflammation or injury. The systems, computer-implemented methods, and non-transitory computer-readable media described herein can also be used to extract blood vessels in the brain and count the number of branching points. A marker that can be used to detect the onset of Alzheimer’s is the number blood vessels and their branches and their perfusion throughout the brain. Accordingly, the systems, computer- implemented methods, and non-transitory computer-readable media described herein may also be used to diagnose and stage Alzheimer's disease and related diseases.

[0149] Modifications are possible in the described examples, and other examples arepossible within the scope of the claims. In this description, “based on” means “based at least in part on.”

Claims

IUIC-00155 CLAIMS What is claimed is:

1. A computer-implemented method of lumen-related disease staging comprising: one or more computer processors receiving volume data; counting, with the one or more computer processors, a number of lumens and a number of generation splits visible in the volume data; measuring, with the one or more computer processors, diameters of the lumens; measuring, with the one or more computer processors, at least one volume of a network of the lumens; calculating, with the one or more computer processors, one or more texture metrics from the volume data; detecting and quantizing, with the one or more computer processors, lesions from the volume data; and using one or more trained machine learning (ML) models, determining a severity stage of the lumen-related disease.

2. The method of claim 1, wherein the counting the number of lumens and the number of generation splits comprises extracting a lumen tree from the volume data and skeletonizing the lumen tree.

3. The method of claim 1, wherein the measuring the diameters of the lumens and the at least one volume of the network of the lumens comprises extracting a lumen tree from the volume data and ray-casting the extracted lumen tree.

4. The method of claim 1, wherein the calculating the one or more texture metrics from the volume data comprises computing gray-level co-occurrence matrices (GLCMs) from the volume data.IUIC-00155 5. The method of claim 4, wherein the calculating the one or more texture metrics from the volume data further comprises computing one or more Haralick features from the computed GLCMs.

6. The method of claim 5, wherein the computed one or more Haralick features are contrast, energy, and homogeneity.

7. The method of claim 1, wherein the one or more ML models determine the severity stage based on feature metrics derived from the counted number of lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.

8. The method of claim 7, wherein: the one or more ML models comprise four binary classification models, two 3-classes classification models, and one 3-class classification model, the classification outputs of the ML models are combined to determine the severity stage, and the severity stage is one of normal, mild, moderate, or severe.

9. The method of claim 8, further comprising computing a severity grade that provides additional granularity to the severity stage, wherein the severity stage is based on the determined severity stage and on features tertiles of the feature metrics.

10. A system comprising: one or more hardware processors; and one or more non-transitory computer-readable media storing instructions which, when executed by the one or more hardware processors cause: receiving volume data; counting a number of lumens and a number of generation splits visible in the volume data; measuring diameters of the lumens;IUIC-00155 measuring at least one volume of a network of the lumens; calculating one or more texture metrics from the volume data; detecting and quantizing lesions from the volume data; and using one or more trained machine learning (ML) models, determining a severity stage of the lumen-related disease.

11. The system of claim 10, wherein the counting the number of lumens and the number of generation splits comprises extracting a lumen tree from the volume data and skeletonizing the lumen tree.

12. The system of claim 10, wherein the measuring the diameters of the lumens and the at least one volume of the network of the lumens comprises extracting a lumen tree from the volume data and ray-casting the extracted lumen tree.

13. The system of claim 10, wherein the calculating the one or more texture metrics from the volume data comprises computing gray-level co-occurrence matrices (GLCMs) from the volume data.

14. The system of claim 13, wherein the calculating the one or more texture metrics from the volume data further comprises computing one or more Haralick features from the computed GLCMs.

15. The system of claim 14, wherein the computed one or more Haralick features are contrast, energy, and homogeneity.

16. The system of claim 10, wherein the one or more ML models determine the severity stage based on feature metrics derived from the counted number of lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.IUIC-00155 17. The system of claim 16, wherein: the one or more ML models comprise four binary classification models, two 3-classes classification models, and one 3-class classification model, the classification outputs of the ML models are combined to determine the severity stage, and the severity stage is one of normal, mild, moderate, or severe.

18. The system of claim 17, further comprising computing a severity grade that provides additional granularity to the severity stage, wherein the severity stage is based on the determined severity stage and on features tertiles of the feature metrics.

19. One or more non-transitory computer-readable media storing program instructions that, when executed by one or more processors, cause the one or more processors to: receive volume data; count a number of lumens and a number of generation splits visible in the volume data; measure diameters of the lumens; measure at least one volume of a network of the lumens; calculate one or more texture metrics from the volume data; detect and quantize lesions from the volume data; and use one or more trained machine learning (ML) models to determine a severity stage of the lumen-related disease.

20. The one or more computer-readable media of claim 19, wherein the instructions cause the one or more processors to determine, using the ML models, the severity stage based on feature metrics derived from the counted number of lumens and generation splits, the measured diameters and volume, the one or more texture metrics, and the detected and quantized lesions.

Citation Information

Patent Citations

  • Colorectal cancer diagnostic composition, and method for detecting diagnostic marker

    US20190018015A1

  • Trimethylamine-containing compounds for diagnosis and prediction of disease

    US20190234917A1

  • Creating a vascular tree model

    US20200077968A1

  • Non-invasive assessment and therapy guidance for coronary artery disease in diffuse and tandem lesions

    US20210085397A1

  • System for computing quantitative biomarkers of texture features in tomographic images

    US9092691B1