Multi-modal method and system for disease diagnosis
By performing metagenomic assembly of tumor tissue and blood samples, combining machine learning models and multiomic data, the problem of insufficient sensitivity in early cancer detection is solved, and efficient and accurate diagnosis of micro tumors and uncertain lung nodules is achieved.
Patent Information
- Application Number
- CN202380079080.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-29
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has insufficient sensitivity in early cancer detection, especially the difficulty in diagnosing micro-tumors and uncertain lung nodules. Traditional methods rely on human biomarker detection and have limitations, making it difficult to achieve efficient early screening and accurate evaluation.
By metagenomic assembly of tumor tissue and blood samples, a reference database is generated, combined with machine learning models, and using microbiomics and proteomics data, a multiomic and multispecies testing method is constructed, combined with radiological images and electronic medical record information, early prediction and diagnosis of cancer is achieved.
Improves the diagnostic sensitivity and specificity of early cancers, especially the identification of tumor mass and uncertain lung nodules with diameters less than 3 cm, reduces the need for invasive examinations and provides more efficient screening compliance and accuracy.
Smart Images

Figure CN120239873A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 412,369, filed on September 30, 2022, which is incorporated herein by reference. BACKGROUND OF THE INVENTION
[0003] Until recently, cancer has been widely considered a sterile tissue, and thus methods for diagnosing cancer have relied on detecting human biomarkers of low to high complexity. However, due to tumor size limiting the molecular abundance in the circulation, such as circulating tumor DNA (ctDNA), the sensitivity of these methods for detecting early disease is challenging. Targeting more biomarkers, deeper sequencing, and / or more frequent testing have been used to improve early detection, but they cannot circumvent the biologically imposed constraints that may hinder clinical utility.
[0004] The prior art in the relevant field may include: US2018 / 0223338, US2018 / 0258495, WO 2019 / 191649, WO 2019 / 079635, WO 2022 / 140386, or WO 2022 / 212283. SUMMARY OF THE INVENTION
[0005] Lung cancer is a typical example of this difficulty, being the leading cause of cancer-related deaths worldwide, and approximately 23% of patients in the United States are diagnosed at the local stage. Although lung cancer screening by low-dose computed tomography (LDCT) has improved early diagnosis and reduced mortality in high-risk individuals, low patient compliance limits its benefits. Minimally invasive liquid biopsies can improve screening compliance, but their sensitivity in early disease is lower than that of LDCT, which can reliably detect nodules as small as 4 millimeters in diameter. For example, in a validation cohort comparing stage I disease with primary healthy controls, commercially available cell-free DNA (cfDNA) methylation assays reported approximately 25% sensitivity at 99.3% specificity (PMID: 33506766); ctDNA-based assays reported approximately 30% average sensitivity at 98% specificity (PMID: 32269342); fragmentomics assays reported approximately 50% sensitivity at 80% specificity (PMID: 34417454); and integrated multi-omics tests reported approximately 40% sensitivity at 96.3% specificity (PMID: 29348365). In contrast, LDCT sensitivity ranges between 59 - 100%. These differences highlight the need for different approaches to early lung cancer detection.
[0006] Another diagnostic challenge comes from indeterminate pulmonary nodules (IPNs) that are difficult to manage, which are non-calcified masses 6-30 millimeters in diameter with a suspicious malignant status. It is reported that approximately 30% of chest CTs show incidental pulmonary nodules, and >1.5 million nodules are detected annually in the United States. Patients with IPNs have a higher risk of lung cancer than the general population, but most IPNs are benign, complicating the assessment. In addition, 25.8% of transthoracic needle biopsies cause clinical complications. For this reason, the standard of care (SOC) for IPN management is conservative monitoring of nodule size and / or PET-CT to rule out malignancy. However, PET-CT assessment is complicated by diabetes hyperglycemia, inconsistent standardized uptake values (SUVs), and false positives due to infectious etiologies. Liquid biopsies can contribute to the diagnostic adjudication of IPNs but will need to have high sensitivity for small nodule sizes (≤3 cm) to improve the PET-CT SOC. To the inventors' knowledge, only one non-PET-CT test is available for IPN malignancy determination, with an AUROC of 0.76 (97% sensitivity, 44% specificity; PMID: 29496499) in the validation cohort, demonstrating that the need for innovation has not been met.
[0007] The inventors and others have previously characterized intracellular, cancer-type specific communities of the tumor microbiome, whose genomes are detectable in the circulation and contain orthogonal biomarkers of human-derived molecules (PMID: 32214244; 32467386; 36179670). However, this work identified a small fraction of the total metagenomic content, either by enriching for microbe-specific amplicons or by matching to capture trace DNA or RNA fragments in publicly available reference databases. Even after multiple rounds of sensitive host depletion and quality filtering, 98.8% of the approximately 4.4 billion non-human DNA or RNA reads in The Cancer Genome Atlas (TCGA) failed to map to any known organism in the RefSeq30 (version 200) multi-domain database (PMID: 36179670), indicating that a large amount of cancer-related microbial diversity remains unexplored.
[0008] Starting with the scientific question of unmappable microbial reads, in some embodiments, the methods and / or systems described elsewhere herein can generate a tumor-centric reference database by metagenomic assembly of 5187 whole-genome sequenced tumor tissue and / or blood-derived cancer samples. In some cases, the metagenomic assembly can include de novo metagenomic assembly. In some cases, the tumor-centric reference data can include up to about 1562 metagenomic bins. In some cases, the tumor-centric reference data can include at least about 1562 metagenomic bins. In some embodiments, these bins increase the median mapping rate of non-human DNA in TCGA by at least about 891-fold compared to publicly available reference genomes, while reducing the median mapping rate of reagent-based contaminants by at least about 7.6-fold. As explored elsewhere herein, in some embodiments, the methods and / or systems of the present invention utilize these cancer-derived metagenomic bins to create two different multi-omics, multi-species tests for lung cancer (e.g., early) detection, which combine bin-derived information with other data features (e.g., plasma proteins and clinical risk scores described elsewhere herein) through predictive (e.g., machine learning) modeling. In some cases, the first model identifies lung cancer in a subject (e.g., a subject in an otherwise healthy population), and the second model can determine the malignant status of, for example, the detected lung cancer or nodule IPN in the subject. In some cases, each of the first model and the second model shows strong predictive performance in a validation subgroup of stage I disease, with a model predictive performance of at least about 85% accuracy, specificity, sensitivity, precision, AUPR, AUROC, or any combination thereof. Thus, by examining the unexplored microbial diversity, the present disclosure demonstrates and describes the practical utility of plasma-derived metagenomics in diagnosing early cancer, and additionally demonstrates the use of metagenomic bins as a pan-cancer (i.e., multiple cancer types, not limited to lung cancer) database from which cancer-associated microbial biomarkers can be identified.
[0009] Aspects of the disclosure provided herein describe a method for determining a disease in a subject, the method comprising: receiving a biological sample, electronic medical record information, and one or more radiological images of the subject; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and determining the disease of the subject as an output of a prediction model when providing the one or more nucleic acid molecule sequencing reads, the electronic medical record information, and data derived from the one or more radiological images of the subject as inputs to the prediction model. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as an input to the prediction model. In some embodiments, the method further comprises identifying one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers from the biological sample of the subject are provided to the prediction model. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof. In some embodiments, the one or more radiological images comprise an x-ray image, a computed tomography (CT) image, a low-dose computed tomography image, a magnetic resonance imaging (MRI) image, an ultrasound image, a positron emission tomography image, a fluoroscopy image, an angiography image, or any combination thereof. In some embodiments, the cancer comprises a tumor mass having a diameter of less than 3 centimeters. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, stomach adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the method further comprises calculating one or more features of the one or more radiological images, wherein the one or more features of the one or more radiological images are provided as input to the prediction model. In some embodiments, the one or more features comprise a Brock cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof. In some embodiments, the method further comprises mapping or aligning the one or more nucleic acid sequencing reads to a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a microbial genomic database. In some embodiments, the microbial genomic database comprises a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a healthy state. In some embodiments, the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the healthy state comprises cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state. In some embodiments, the microbial genomic database comprises the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the prediction model comprises a machine learning model. In some embodiments, the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the prediction model is trained using leave-one-out validation. In some embodiments, the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads. In some embodiments, decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, sequencing comprises shotgun metagenomic sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more features. In some embodiments, the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject. In some embodiments, the mapping or alignment is done using Deblur, Bowtie2, Kraken, or any combination thereof.
[0010] Another aspect of the disclosure provided herein describes a method comprising: receiving a biological sample, electronic medical record information, data derived from one or more radiological images, and a corresponding disease of one or more subjects; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and identifying one or more features corresponding to the disease of the one or more subjects in the one or more nucleic acid molecule sequencing reads, the electronic medical record information, and the data derived from the one or more radiological images. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as input to the prediction model. In some embodiments, identifying comprises aligning the one or more sequencing reads against a genomic database. In some embodiments, the method further comprises training a prediction model with the one or more features and the corresponding disease of the nucleic acid molecule sequencing reads, the electronic medical record information, and the data derived from the one or more radiological images of the one or more subjects. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the method further comprises identifying one or more features of one or more protein biomarkers of the biological sample of the subject. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof. In some embodiments, the one or more radiological images comprise x-ray images, computed tomography (CT) images, low-dose computed tomography images, magnetic resonance imaging (MRI) images, ultrasound images, positron emission tomography images, fluoroscopy images, angiography images, or any combination thereof. In some embodiments, the cancer comprises a tumor mass with a diameter less than 3 centimeters. In some embodiments, sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, stomach adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the one or more radiological image features comprise Brock cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof. In some embodiments, the method further comprises mapping or aligning the one or more nucleic acid sequencing reads against a genomic database to determine one or more human, non-human, or combination features of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a microbial genomic database. In some embodiments, the microbial genomic database comprises de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a healthy state. In some embodiments, the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the healthy state comprises cancer, pre-cancerous state, non-malignant disease state, or disease-free healthy state. In some embodiments, the microbial genomic database comprises the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the prediction model comprises a machine learning model.In some embodiments, the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the prediction model is trained using leave-one-out validation. In some embodiments, the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads. In some embodiments, decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more features. In some embodiments, the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject. In some embodiments, the mapping or alignment is done using Deblur, Bowtie2, Kraken, or any combination thereof.
[0011] Another aspect of the disclosure provided herein describes a computer system configured to determine a disease of a subject, the computer system comprising: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, wherein the software comprises executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) receive one or more sequencing reads, electronic medical record information, and one or more images of a biological sample of the subject; and (ii) determine the disease of the subject as an output of a prediction model when providing the one or more nucleic acid molecule sequencing reads, the electronic medical record information, and data derived from one or more radiological images of the subject as inputs to the prediction model. In some embodiments, the one or more nucleic acid sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as an input to the prediction model. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the biological sample comprises a tissue biopsy, a liquid biopsy, or a combination thereof. In some embodiments, the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers from the biological sample of the subject are provided to the prediction model. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, or a combination thereof. In some embodiments, the prediction model is trained with one or more features of the nucleic acid molecule sequencing reads, the electronic medical record information, and the data derived from the one or more radiological images of the one or more subjects and the corresponding diseases. In some embodiments, the executable instructions comprise identifying one or more features of one or more protein biomarkers of the biological sample of the subject. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the one or more radiological images comprise x-ray images, computed tomography (CT) images, low-dose computed tomography images, magnetic resonance imaging (MRI) images, ultrasound images, positron emission tomography images, fluoroscopy images, angiography images, or any combination thereof.In some embodiments, the cancer comprises a tumor mass with a diameter less than 3 cm. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads. In some embodiments, the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise sequencing reads of the following: mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, low-grade glioma of the brain, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, renal chromophobe cell carcinoma, renal clear cell carcinoma of the kidney, renal papillary cell carcinoma of the kidney, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the one or more radiological image features comprise a Brock cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof. In some embodiments, the executable instructions further comprise mapping or aligning the one or more nucleic acid sequencing reads against a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a microbial genomic database. In some embodiments, the microbial genomic database comprises de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a healthy state. In some embodiments, the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.In some embodiments, the health state includes cancer, pre-cancerous state, non-malignant disease state, or disease-free healthy state. In some embodiments, the microbial genome database includes the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof. In some embodiments, the genome database includes a human genome database. In some embodiments, the prediction model includes a machine learning model. In some embodiments, the prediction model includes a neural network, a convolutional neural network, logistic regression, random forest, support vector machine, or any combination thereof. In some embodiments, the machine learning model includes a machine learning classifier. In some embodiments, the machine learning model includes a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the prediction model is trained using leave-one-out validation. In some embodiments, the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the executable instructions further include decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads. In some embodiments, decontamination includes computer simulation decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the executable instructions further include determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include non-microbial taxonomic abundances, mammalian genome coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more features. In some embodiments, the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject. In some embodiments, the mapping or alignment is done using Deblur, Bowtie2, Kraken, or any combination thereof.
[0012] Another aspect of the disclosure provided herein describes a method for determining a disease in a subject, the method comprising: receiving a biological sample from the subject; sequencing one or more nucleic acid molecules of the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and determining the disease of the subject as an output of the prediction model when providing the one or more nucleic acid molecule sequencing reads of the subject to the prediction model, wherein the prediction model is trained with one or more nucleic acid molecule sequencing reads and corresponding diseases of one or more liquid biological samples and one or more tissue biological samples of one or more subjects. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided to the prediction model as an input. In some embodiments, the disease comprises cancer, non-cancerous diseases, or a combination thereof. In some embodiments, the method further comprises identifying one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers from the biological sample of the subject are provided to the prediction model. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass with a diameter less than 3 centimeters. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.In some embodiments, the cancer includes: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, stomach adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the method further includes mapping or aligning the one or more nucleic acid sequencing reads relative to a genomic database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads provided as input to the prediction model. In some embodiments, the genomic database includes a microbial genomic database. In some embodiments, the microbial genomic database includes a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representative of a healthy state. In some embodiments, the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample includes plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the healthy state includes cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state. In some embodiments, the microbial genomic database includes the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof. In some embodiments, the genomic database includes a human genomic database. In some embodiments, the prediction model includes a machine learning model. In some embodiments, the prediction model includes a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof. In some embodiments, the machine learning model includes a machine learning classifier. In some embodiments, the machine learning model includes a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the prediction model is trained using leave-one-out validation. In some embodiments, the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV.In some embodiments, the method further comprises decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads, wherein the one or more decontaminated nucleic acid molecules are provided as input to the prediction model. In some embodiments, decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more characteristics of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more characteristics. In some embodiments, the prediction model is configured to distinguish cancer and non-cancerous diseases in the subject. In some embodiments, mapping or alignment is done with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
[0013] Another aspect of the disclosure provided herein describes a method for identifying one or more non-human genomic features, the method comprising: receiving one or more liquid biological samples, one or more tissue biological samples, and a corresponding disease of one or more subjects; sequencing one or more nucleic acid molecules of the one or more liquid biological samples and the one or more tissue biological samples, thereby generating one or more sequencing reads; and identifying from the one or more sequencing reads one or more non-human genomic features corresponding to the disease of the one or more subjects. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as input to the prediction model. In some embodiments, identifying comprises aligning or mapping the one or more sequencing reads against a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a microbial genomic database. In some embodiments, the microbial genomic database comprises a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a healthy state. In some embodiments, the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the healthy state comprises cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state. In some embodiments, the microbial genomic database comprises a RefSeq database, a LifeNet database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof. In some embodiments, the method further comprises training a prediction model with the one or more non-human genomic features and the corresponding disease of the one or more subjects. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the method further comprises identifying one or more features of one or more protein biomarkers of the one or more liquid biological samples, the one or more tissue biological samples, or a combination thereof.In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass with a diameter less than 3 cm. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules include mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biological sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer includes lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer includes: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, low-grade glioma of the brain, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, renal chromophobe cell carcinoma, renal clear cell carcinoma of the kidney, renal papillary cell carcinoma of the kidney, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid cancer, uterine carcinosarcoma, endometrial carcinoma of the uterine body, uveal melanoma, or any combination thereof. In some embodiments, the genomic database includes a human genomic database. In some embodiments, the prediction model includes a machine learning model. In some embodiments, the prediction model includes a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof. In some embodiments, the machine learning model includes a machine learning classifier.In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the prediction model is trained using leave-one-out validation. In some embodiments, the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads. In some embodiments, decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more characteristics of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more characteristics. In some embodiments, the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject. In some embodiments, mapping or alignment is done using Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
[0014] Another aspect of the disclosure provided herein describes a computer system configured to determine a disease of a subject, the computer system comprising: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, wherein the software comprises executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) receive one or more sequencing reads of a biological sample of the subject; and (ii) determine the disease of the subject as an output of a prediction model when providing the one or more nucleic acid molecule sequencing reads of the subject to the prediction model, wherein the prediction model is trained with one or more nucleic acid molecule sequencing reads and corresponding diseases of one or more liquid biological samples and one or more tissue biological samples of one or more subjects. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided to the prediction model as an input. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers from the biological sample of the subject are provided to the prediction model. In some embodiments, the executable instructions comprise identifying one or more characteristics of the one or more protein biomarkers of the biological sample of the subject. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass with a diameter of less than 3 centimeters. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads. In some embodiments, the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of the one or more nucleic acid molecules.In some embodiments, the one or more nucleic acid molecule sequencing reads comprise sequencing reads of mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biological sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine corpus endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the executable instructions further comprise mapping or aligning the one or more nucleic acid sequencing reads against a genomic database to determine one or more human, non-human, or combined characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a microbial genomic database. In some embodiments, the microbial genomic database comprises a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a healthy state. In some embodiments, the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the healthy state comprises cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state. In some embodiments, the microbial genomic database comprises the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the prediction model comprises a machine learning model. In some embodiments, the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.In some embodiments, the machine learning model includes a machine learning classifier. In some embodiments, the machine learning model includes a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the prediction model is trained using leave-one-out validation. In some embodiments, the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the executable instructions further include decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads. In some embodiments, decontamination includes in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the executable instructions further include determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more features. In some embodiments, the prediction model is configured to distinguish the cancer of the subject from non-cancerous diseases. In some embodiments, the mapping or alignment is done using Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
[0015] In some embodiments, aspects of the disclosure described herein describe a method for determining a disease of a subject, the method comprising: (a) receiving a biological sample, electronic medical record information, and radiology data of the subject; (b) sequencing a plurality of non-human nucleic acid molecules of the biological sample, thereby generating a plurality of microbial sequencing reads; and (c) processing the plurality of microbial sequencing reads, the electronic medical record information, and the radiology data with a trained prediction model, thereby determining the disease of the subject with at least about 80% accuracy, wherein the trained prediction model is trained with a plurality of microbial abundances and corresponding cancer types, wherein the trained prediction model includes a first prediction model and a second prediction model, and wherein the first prediction model processes the plurality of microbial sequencing reads, the electronic medical record information, and the radiology data, and wherein the second trained prediction model processes the output of the first prediction model. In some embodiments, the biological sample comprises a liquid biopsy. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the sequencing comprises shotgun sequencing. In some embodiments, the method further comprises receiving the concentration of one or more plasma proteins. In some embodiments, the trained prediction model comprises one or more machine learning models. In some embodiments, the method further comprises aligning the plurality of nucleic acid molecule sequencing reads of the biological sample to a human reference genome to identify a plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, the method further comprises aligning the plurality of non-human nucleic acid molecule sequencing reads to a microbial genome database to identify the plurality of microbial sequencing reads. In some embodiments, the database comprises a de novo metagenomic assembly including genomic contigs. In some embodiments, the genomic contigs comprise one or more metagenomic bins. In some embodiments, aligning the plurality of non-human nucleic acid molecule sequencing reads to the de novo metagenomic assembly yields the aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, the trained prediction model is configured to process the aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads of the subject. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing of the plurality of nucleic acid molecules of the biological sample. In some embodiments, the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of the plurality of nucleic acid molecules of the biological sample. In some embodiments, the method further comprises determining one or more features of the radiology data, wherein the one or more features of the radiology data are processed by the trained prediction model.In some embodiments, the one or more features of the radiological data include a Brock cancer probability score, a cancer lesion diameter, a cancer lesion spiculation sign, a cancer lesion solidity, or any combination thereof. In some embodiments, the disease includes cancer. In some embodiments, the cancer includes a tumor mass with a diameter of less than about 3 centimeters or less than about 8 millimeters. In some embodiments, the trained prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, determining the disease of the subject includes differentiating the subject's cancer from a non-cancerous disease.
[0016] In some embodiments, aspects of the disclosure described herein describe a system configured to determine a disease of a subject, the system comprising: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, wherein the software includes executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) receive a plurality of non-human nucleic acid molecule sequencing reads, electronic medical record information, and radiology data of a biological sample of the subject; and (ii) process a plurality of microbial nucleic acid molecule sequencing reads, the electronic medical record information, and the radiology data of the subject in the plurality of non-human nucleic acid molecule sequencing reads with a trained prediction model, thereby determining the disease of the subject with an accuracy of at least about 80%, wherein the trained prediction model is trained with a plurality of microbial abundances and corresponding cancer types, wherein the trained prediction model includes a first prediction model and a second prediction model, and wherein the first prediction model processes the plurality of microbial nucleic acid molecule sequencing reads, the electronic medical record information, and the radiology data, and wherein the second prediction model processes the output of the first prediction model. In some embodiments, the biological sample includes a liquid biopsy. In some embodiments, the liquid biopsy includes plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the plurality of microbial nucleic acid molecule sequencing reads are generated by shotgun sequencing. In some embodiments, the executable instructions include receiving the concentration of one or more plasma proteins. In some embodiments, the trained prediction model includes one or more machine learning models. In some embodiments, the executable instructions cause the one or more processors to align the plurality of nucleic acid molecule sequencing reads of the biological sample with a human reference genome library to identify the plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, the executable instructions cause the one or more processors to align the plurality of non-human nucleic acid molecule sequencing reads with a microbial genome database to identify the plurality of microbial sequencing reads. In some embodiments, the database includes a de novo metagenomic assembly including genomic contigs. In some embodiments, the genomic contigs include one or more metagenomic bins. In some embodiments, the executable instructions cause the one or more processors to align the plurality of non-human nucleic acid molecule sequencing reads with the de novo metagenomic assembly to generate the aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, the trained prediction model is configured to process the aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads of the subject.In some embodiments, the executable instructions cause the one or more processors to receive a plurality of amplicon-based 16S rRNA sequencing reads of the plurality of nucleic acid molecules of the biological sample. In some embodiments, the plurality of amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of the plurality of nucleic acid molecules of the biological sample. In some embodiments, the executable instructions cause the one or more processors to determine one or more features of the radiology data, wherein the one or more features of the radiology data are processed by the trained prediction model. In some embodiments, the one or more features of the radiology data comprise a Brock cancer probability score, a cancer lesion diameter, a cancer lesion spiculation, a cancer lesion solidity, or any combination thereof. In some embodiments, the disease comprises cancer. In some embodiments, the cancer comprises a tumor mass having a diameter of up to about 3 centimeters or up to about 8 millimeters. In some embodiments, the trained prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof. In some embodiments, determining the disease of the subject comprises differentiating the subject's cancer from non-cancerous diseases.
[0017] Incorporation by Reference
[0018] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings in which:
[0020] Figure 1 A flowchart illustrating data types and data flow structures for machine learning training for stacking, as described in some embodiments herein.
[0021] Figure 2 A flowchart illustrating metagenomic data isolated and / or identified from a human subject, as described in some embodiments herein.
[0022] Figure 3 A flowchart illustrating proteomic data derived from a human subject, as described in some embodiments herein.
[0023] Figure 4Shows a flowchart of data types derived from human subject samples and human subject medical records for training and / or developing a diagnostic classifier as described in some embodiments herein.
[0024] Figure 5 Shows a flowchart of clinical protein metagenomic features for generating a lung cancer classifier as described in some embodiments herein.
[0025] Figure 6 Shows a computer system configured to implement the methods of the present disclosure as described in some embodiments herein.
[0026] Figure 7 A-7K shows experimental data and graphs of metagenomic features as described in some embodiments herein, which widely distinguish untreated lung cancer and healthy samples across different lung cancer tissue types and stages using stacked machine learning.
[0027] Figure 8 A-8H shows experimental data and graphs of using metagenomic plasma biomarkers to distinguish lung cancer from lung diseases and developing metagenomic bins to enhance the discrimination performance as described in some embodiments herein.
[0028] Figure 9 A-9M shows experimental data and graphs of the performance of a metagenomic bin-based pan-cancer classifier as described in some embodiments herein, which diagnoses the presence and type of cancer in blood and plasma of two independent cohorts.
[0029] Figure 10 A-10I shows experimental data, graphs, and / or work flowcharts of the development and validation of a clinical protein metagenomic classifier for determining the malignancy of lung nodules as described in some embodiments herein.
[0030] Figure 11 A-11C shows experimental data and graphs of the decontamination and total genomic coverage of plasma-derived microbiomes in a sample cohort as described in some embodiments herein.
[0031] Figure 12 A-12D shows experimental data and graphs of batch correction analysis of the microbiomes of Cancer Genome Atlas (TCGA) biological samples using metagenomic bin features as described in some embodiments herein.
[0032] Figure 13 A-13E shows experimental data and graphs of the performance of a TCGA classifier for cancer discrimination using metagenomic bins between cancer types as described in some embodiments herein.
[0033] Figure 14 A-14B shows the experimental data and graphs of the alignment analysis as described in some embodiments herein, which validates the TCGA tissue-based classifier performance by comparing machine learning models built with scrambled metadata or shuffled samples.
[0034] Figure 15 A-15C shows the experimental data and graphs of the control analysis as described in some embodiments herein, which validates the performance of the TCGA blood-based classifier and further performs machine learning using blood samples from low-stage cancers or comparing primary tumors from low and high clinical stages.
[0035] Figure 16 A-16F shows the experimental data and graphs of the TCGA classifier performance for primary tumor cancer type discrimination using a subset of the raw data with metagenomic bin abundances as described in some embodiments herein.
[0036] Figure 17 A-17F shows the experimental data and graphs of the control analysis of the raw data as described in some embodiments herein to validate the performance of the TCGA primary tumor classifier.
[0037] Figure 18 A-18D shows the experimental data and graphs of the TCGA classifier performance as described in some embodiments herein, which uses a subset of the raw data with metagenomic bins and control analysis for discrimination between primary tumors and adjacent normal tissues.
[0038] Figure 19 A-19H shows the experimental data and graphs of the TCGA blood classifier performance as described in some embodiments herein, which uses a subset, raw, metagenomic bin abundance for cancer discrimination and control analysis.
[0039] Figure 20 A-20E shows the experimental data and graphs of the metagenomic bin alpha diversity across TCGA primary tumors as described in some embodiments herein.
[0040] Figure 21 A-21E shows the experimental data and graphs of the classical metagenomic analysis of cancer type specificity in TCGA primary tumor samples when using raw metagenomic bin abundances as described in some embodiments herein.
[0041] Figure 22 A-22G shows the experimental data and graphs of the differential abundance of metagenomic bins in TCGA primary tumor cancer types as described in some embodiments herein.
[0042] Figure 23A-23E shows experimental data and graphs of the metagenomic bin alpha diversity across TCGA blood samples as described in some embodiments herein.
[0043] Figure 24 A-24E shows experimental data and graphs of a classical metagenomic analysis specific to cancer types in TCGA blood samples when using raw metagenomic bin abundances as described in some embodiments herein.
[0044] Figure 25 A-25F shows experimental data and graphs of the differential abundances of metagenomic bins in the blood samples of TCGA and their accompanying cancer types as described in some embodiments herein.
[0045] Figure 26 A-26C shows experimental data and graphs of the diagnostic performance of metagenomics, proteins, and amplicons in the Comprehensive Oncology Diagnostic Identification for Early Cancer Screening (CODICES) cohort as described in some embodiments herein.
[0046] Figure 27 Shows experimental data and graphs of the diagnostic performance of metagenomic bins in the CODICES cohort as described in some embodiments herein.
[0047] Figure 28 Shows a flow chart of data types derived from human subject samples and human subject medical records for training and / or developing diagnostic classifiers as described in some embodiments herein.
[0048] Figure 29 A-29N shows experimental data and graphs of the diagnostic performance of genomic and DNA fragmentomics (nucleotide frequency) analysis obtained from publicly available cell-free DNA datasets as described in some embodiments herein.
[0049] Figure 30 A-30F shows experimental data and graphs of the diagnostic performance of lung cancer relative to healthy controls based on metagenomics, plasma proteins, and DNA fragmentomics (nucleotide frequency) analysis derived from the CODICES cohort as described in some embodiments herein.
[0050] Figure 31 Shows experimental data and graphs of the performance of lung disease relative to lung cancer based on metagenomics, plasma proteins, and DNA fragmentomics (nucleotide frequency) analysis derived from the CODICES cohort as described in some embodiments herein.
[0051] Figure 32A-32D shows an improvement in the whole-genome and RNA sequencing mapping rates of metagenomic bins compared to the publicly available RefSeq206 genome database.
[0052] Figures 33A-33I show the colorectal and lung cancer classifier performance of machine learning models generated from fecal microbiome data, where microbial abundances are derived from alignments of sequencing reads to the metagenomic assemblies (‘bins’) of the present invention. Detailed Description
[0053] In some embodiments, the present disclosure describes methods for determining, differentiating, and / or distinguishing a malignant disease (e.g., a malignant lung neoplasm, alternatively referred to as a tumor) from a benign disease by analyzing and / or assaying a combination of two or more data types of a patient sample. In some cases, the present disclosure describes methods and / or systems for determining and / or diagnosing a disease. In some cases, the disease can include cancer. In some cases, the disease can include a non-cancerous disease. In some cases, the methods and / or systems described elsewhere herein can differentiate and / or distinguish between a cancerous and a non-cancerous disease of a subject. In some cases, the cancer can include a tumor mass with a diameter of less than about 3 cm or less than about 8 mm. In some embodiments, the data types of a patient sample can include one or more analyte types of the patient sample. The one or more analyte types can include the presence and / or abundance of one or more non-human source nucleic acid molecules (e.g., microorganisms, bacteria, viruses, and / or fungi), a liquid-based (e.g., blood-derived) patient sample, non-microbial human-source analytes (e.g., proteins, human genomic nucleic acid molecules), or a combination thereof. In some cases, a liquid biopsy can include plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some cases, the data types can include patient history and / or diagnostic data types. In some embodiments, the present specification provides a method and / or system configured to and / or capable of detecting and / or determining the presence of early clinical stage cancer (e.g., stage I and stage II lung cancer) in an asymptomatic subject at high risk of lung cancer due to the subject's smoking history. For example, such subjects may include those who meet the annual low-dose computed tomography screening recommendations of the United States Preventive Services Task Force: adults aged 50-80 years, with a 20-pack-year smoking history, and currently smoking or having quit smoking within the past 15 years. As described elsewhere herein, the method can include training one or more prediction models (e.g., machine learning algorithms) on a data set (e.g., a feature table), where the data set includes the aforementioned data types, and thereby identifying disease (e.g., cancer)-related patterns in these data types. Specifically, as described elsewhere herein, the method and / or system can train one or more prediction models with an analyte type of microbial abundance (relative or absolute) obtained from sequencing only one or more microbial nucleic acid molecules (e.g., determining microbial read counts by next-generation sequencing) or in combination with a measured plasma protein concentration. In some cases, the sequencing can include shotgun sequencing. In some embodiments, the training data set can include microbial abundance, measured plasma proteins, human genomic information (e.g., DNA fragmentation patterns or chromosomal copy number variations), a cancer risk score calculated based on information obtained from a patient medical chart, or any combination thereof.In some embodiments, the methods and / or systems described elsewhere herein can train a predictive model (e.g., a multimodal model) with one or more data types to produce a trained diagnostic model (e.g., a trained multimodal diagnostic model).
[0054] In some cases, the methods and / or systems described elsewhere herein achieve at least about 70% or at least about 0.7 of a performance metric (e.g., accuracy, sensitivity, specificity, NPV, PPV, AUROC, and / or AUPR) of a predictive model when determining, diagnosing, and / or detecting cancer in a subject based on one or more data types. In some cases, when training one or more predictive models described elsewhere herein, using nucleic acids of an analyte type of non-human origin can improve the performance of the trained model in differentiating between a first disease state and a second disease state in the lung (e.g., the presence of a lung cancer nodule and a non-cancerous lung nodule). In some cases, non-cancerous lung nodules may be caused by etiologies including sarcoidosis, interstitial pulmonary fibrosis, bronchiectasis, pneumonia, chronic obstructive pulmonary disease, and / or hamartoma. Traditionally, such nodule-bearing conditions can be distinguished from true lung malignancies by invasive intrathoracic biopsy followed by histopathological examination. The methods and / or systems described elsewhere herein provide methods and / or systems that are superior to typical pathology reports when examining, for example, intrathoracic tissue biopsies because the methods and / or systems of the present disclosure do not rely on subjective measurements of observed tissue architecture, cellular atypia, or any other subjective measurements traditionally used to diagnose cancer. In some embodiments, the methods and / or systems can alternatively rely on the measured quantitative presence and / or abundance of one or more data types. In some cases, the methods and / or systems can demonstrate an increase in the performance metric by narrowing the features only on a microbial source rather than on a modified human (i.e., cancerous) source, which is often modified at a very low frequency in the context of a 'normal' human source. Additionally, the methods and / or systems of the present disclosure can utilize and / or be performed on blood-derived samples, which are minimally invasive samples compared to, for example, invasive biopsies, and thus pose little risk to the patient, can be repeated longitudinally, and are low cost.
[0055] In some embodiments, the methods and / or systems described elsewhere herein utilize, analyze, and / or train a predictive model with metagenomic assemblies generated from non-human sequencing reads obtained from tumor whole-genome sequencing data. In some cases, the metagenomic assemblies can include de novo metagenomic assemblies. In some cases, by using the non-human components of the whole-genome sequencing data, the methods can determine the microbial components of a tumor type and use these components as a reference database for subsequent NGS sequencing read alignment. As described elsewhere herein, the metagenomic assemblies utilized by the methods and / or systems are derived in part from untreated biopsy tumor samples in The Cancer Genome Atlas (TCGA) and thus represent distinct metagenomes from more than 30 different cancer types and / or (sub)types. In some embodiments, as described elsewhere herein, cell-free DNA sequencing reads from a subject (e.g., a patient) sample are computationally aligned against one or more tumor-derived metagenomes to identify the presence and / or abundance of tumor-associated microbial signatures. These sample-associated microbial abundances can then be combined with other data types (e.g., plasma protein concentrations and / or clinical data features) to form a multimodal feature set, which is input into a trained diagnostic model to generate a likelihood of a cancer score. Healthcare providers and / or clinicians can use this score to determine whether a subsequent, more invasive biopsy, such as a chest biopsy, is needed or whether continued non-invasive monitoring of a lung nodule via radiomics imaging, such as low-dose computed tomography and / or further liquid biopsy-based assays and / or measurements, as described elsewhere herein, is needed.
[0056] In some embodiments, metagenomic data can be combined with other data types to form a multimodal input data set for training a predictive model (e.g., one or more machine learning algorithms). In some embodiments, a trained diagnostic model is generated using multimodal data derived from one or more patients and / or subjects with a known health history. For example, in the case of lung cancer, metagenomic data from known healthy subject 101, subject 102 with lung cancer, and subject 103 with non-cancerous lung conditions (e.g., sarcoidosis, interstitial pulmonary fibrosis, bronchiectasis, pneumonia, chronic obstructive pulmonary disease, etc.) can be combined with proteomic data (104-106) and clinical data (107-109) from the same subjects and used as a training data set for a collective ("stacked") predictive model (e.g., a stacked machine learning algorithm) 110 capable of learning from data containing both numerical and categorical data types. In some embodiments, the predictions from stacked predictive models from different model types (e.g., logistic regression, random forest, and support vector machine) can be combined to produce a single final predictive model 111, as Figure 1As shown. The results of training this stacked prediction model can include a diagnostic classifier that can be used to distinguish one or more cancerous subjects from healthy subjects and / or one or more cancerous subjects from non-cancer (i.e., lung disease) subjects based on the analysis of one or more data types of a subject sample that was not previously used to train the prediction model.
[0057] In some embodiments, metagenomic data (101-103, Figure 1 ; and / or 112, Figure 2) can be derived from biological samples of one or more training and / or test subjects. In some embodiments, the biological samples of the one or more training and / or test subjects may contain human 113 and non-human 115 nucleic acid molecules. In some cases, one or more sequences of the human 113 and / or non-human 115 nucleic acid molecules can be determined by sequencing. In some cases, sequencing may include next-generation sequencing. In some cases, sequencing may include shotgun sequencing. In some cases, sequencing of one or more sequences of the human 113 and / or non-human 115 nucleic acid molecules can generate and / or produce one or more sequencing reads of the human 113 and / or non-human 115 nucleic acid molecules. In some cases, as described elsewhere herein, one or more genomic analyses and / or assays can be performed and / or applied to the human 113 sequencing reads to generate a set of human genomic features for subsequent prediction model learning, analysis, and / or training 117. For example, human genome 114 analysis may include detecting cancer-related DNA mutations, copy number variations, DNA fragmentation profiles, DNA end analysis (e.g., analyzing nucleotide frequencies at the ends of DNA fragments), chromosomal instability profiles, and epigenetic profiles (e.g., DNA methylation patterns), or any combination thereof. In some embodiments, DNA fragmentation (e.g., DNA end analysis) analysis can provide nucleotide frequencies of one or more nucleotide bases. In some cases, the nucleic acid molecule sequence read length of 100 base pairs (bp) may include from about 1 to about 30 bases for analyzing nucleotide frequencies. In some cases, the nucleic acid molecule sequence read length of 100 bp may include from about 1 to about 25 bases for analyzing nucleotide frequencies. In some cases, nucleotide frequency includes the occurrence of a given nucleotide in the analyzed base segment. In some cases, at least 3 nucleotide bases of the nucleic acid molecule sequence are analyzed. In some embodiments, one or more analyses and / or assays can be performed on the non-human component 115 of the metagenomic dataset to generate a feature table for prediction model learning, training, and / or analysis. In some cases, non-human (e.g., microbial) genome analysis and / or assay 116 may include determining the taxonomic abundance of microorganisms in the sample; inferring the abundance of biochemical pathways represented by the identified microorganisms; aligning the sequencing reads to a metagenomic assembly ('bin') to obtain bin abundance; determining amplicon sequence variants present in a targeted amplicon sequencing dataset; or any combination thereof. In some cases, multiple nucleic acid molecules of a biological sample (e.g., liquid biopsy) can be aligned to a human reference genome to identify multiple non-human nucleic acid molecule sequencing reads. In some cases, the non-human nucleic acid molecule sequencing reads can be aligned to a microbial genome database to identify multiple microbial sequencing reads. In some cases, the microbial genome database may include a de novo metagenomic assembly of genomic contigs.In some cases, the genomic contigs can include one or more metagenomic bins. In some embodiments, the plurality of non-human nucleic acid molecule sequencing reads can be aligned relative to de novo metagenomic assembly to generate aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads. In some cases, as described elsewhere herein, one or more predictive models can process the aligned bin abundances of one or more subjects to determine a disease of the subject, as described elsewhere herein. In some cases, a targeted amplicon sequencing dataset can include an amplicon-based 16S rRNA sequencing read dataset of a plurality of nucleic acid molecules of a biological sample. In some cases, the amplicon-based 16S rRNA sequencing reads can include sequencing reads of the V6 region of a plurality of nucleic acid molecules of a biological sample of a subject as described elsewhere herein.
[0058] In some embodiments, proteomic data (104 - 106, Figure 1 ; 118, Figure 3 ) derived from biological samples of one or more training and / or test subjects can include serum or plasma proteins 119 of blood origin. In some embodiments, in plasma, the proteins 120 can include carcinoembryonic antigen (CEA), osteopontin (OPN), cancer antigen 125 (CA125), cancer antigen 19-9 (CA 19-9), cancer antigen 15-3 (CA 15-3), interleukin-8 (IL-8), prolactin (PRL), and cytokeratin 19 fragment (CYRA 21-1). The concentrations of these plasma proteins (usually in ng / mL) can be determined by immunoassays, as detailed herein, and used as input 117 (along with other data types (e.g., Figure 1 )) to train a predictive model (e.g., a machine learning algorithm) and / or be analyzed by a previously trained predictive model.
[0059] In some embodiments, data 121 from a human subject can include a combination of proteomic data 122 and non-genomic data 131, e.g., including various data features extracted from the subject's medical history, such as Figure 4 as shown. In some embodiments, the data features can include the subject's smoking history 123, a clinical score 124 for assessing cancer risk, clinical findings and family health history 125, and data 126 derived from medical imaging. These data types can be analyzed separately - in the absence of metagenomic data ( Figure 4 ) or can be combined with metagenomic data features (132, 133) as shown in Figure 28 to train a predictive model and / or be analyzed and / or input 117 into a previously trained predictive model to determine a diagnostic output, as described elsewhere herein.
[0060] In some embodiments, a trained prediction model (e.g., a trained machine learning classifier) that can distinguish malignant cancerous (e.g., lung cancer) nodules from benign cancer nodules (e.g., benign lung cancer nodules) (130, Figure 5 ) can be determined and / or obtained using the prediction model input feature table 117, which contains metagenomic bin abundances 127 derived from cell-free DNA sequencing data, protein concentrations 128 of plasma proteins CEA and OPN, and / or clinical data 129 obtained from the subject's medical record sheet. In some cases, the clinical data can include radiological data. In some cases, the radiological data can include radiological images. In some cases, the radiological images can be generated by X-ray, computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), or any combination thereof. In some cases, one or more features of the radiological data can be determined and utilized to diagnose and / or determine the subject's disease and / or to train one or more prediction models only in combination with other data types, as described elsewhere herein. In some cases, one or more features of the radiological data can be analyzed, determined, and / or identified. In some cases, the one or more features of the radiological data can include a Brock cancer probability score, cancer lesion diameter, cancer lesion spiculation, cancer lesion solidity, or any combination thereof. In some embodiments, the clinical data features can include a Mayo lung cancer probability score, a Brock lung cancer probability score, the subject's smoking status, or any combination thereof. In some cases, the clinical data features can include features determined from medical imaging, e.g., lung tumor / nodule size, tumor solidity, evidence of spiculation, lung nodule location, presence of emphysema, or any combination thereof.
[0061] Prediction Model
[0062] The methods and / or systems of the present disclosure may utilize and / or access external capabilities of artificial intelligence, predictive models, and / or machine learning techniques to identify one or more microbial signatures of an enriched (e.g., hybridization-enriched) biological sample of one or more subjects. In some cases, the microbial signatures determined from the hybridization-enriched biological samples of the subjects may predict cancer and / or non-cancerous diseases of one or more subjects. In some cases, the signatures may be used to train one or more predictive models (e.g., one or more machine learning algorithms), as described in other parts of this document. These signatures may be used to predict diseases, such as cancer, non-cancerous diseases, disorders, or any combination thereof. Using such trained predictive models, healthcare providers (e.g., physicians) may make informed, accurate risk-based decisions, thereby improving the quality of care and monitoring provided to patients regarding cancer, non-cancerous diseases, disorders, or any combination thereof. In some cases, the one or more predictive models may include a first predictive model and a second predictive model, the first predictive model being configured to process and / or receive one or more data types as input, as described in other parts of this document, the second predictive model being configured to receive and / or process the output of the first predictive model, wherein the output of the second predictive model determines and / or diagnoses the disease of the subject. In some cases, at least three predictive models may provide output to a single predictive model (e.g., a logistic regression model), which may then output a determination, detection, and / or diagnosis of the disease of the subject. In some cases, the at least three predictive models may be trained with one or more data types, as described in other parts of this document. In some cases, a single predictive model that receives input from at least three predictive models may weight the outputs of the at least three predictive models, wherein each model has a different weight.
[0063] The methods and / or systems described in other parts of this disclosure can analyze the presence and / or abundance of microorganisms (e.g., the abundance of microorganisms of a particular genus and / or taxon) in a biological sample enriched by hybridization probes, where the hybridization probes can bind non-specifically to microbial nucleic acids, as described in other parts. The presence and / or abundance of microorganisms can be used to determine one or more microbial signatures and / or non-microbial signatures that can predict cancer and / or non-cancerous diseases in one or more subjects. In some cases, the methods and / or systems described in other parts of this disclosure can train a prediction model with one or more microbial signatures and / or non-microbial signatures indicative of cancer and / or non-cancerous diseases in a subject. In some cases, the trained prediction model can be used to generate the likelihood (e.g., prediction) of cancer and / or non-cancerous diseases in one or more subjects different from the one or more subjects used to train the prediction model. The trained prediction model can comprise an artificial intelligence-based model, such as a machine learning-based classifier, configured to process one or more microbial nucleic acid molecule sequencing reads obtained from a hybrid enrichment biological sample to generate the likelihood that a subject has a disease or disorder. The presence and / or abundance of microorganisms in a hybrid enrichment biological sample from one or more groups of patients, such as cancer patients, patients with non-cancerous diseases, patients without disease and without cancer, cancer patients receiving treatment for cancer, patients receiving treatment for non-cancerous diseases, or any combination thereof, can be used to train the model. In some cases, a prediction model can be trained to provide treatment predictions for cancer in one or more patients not part of the training dataset of the prediction model. When provided with an input of the presence and abundance of one or more microorganisms in a hybrid enrichment biological sample of a patient, such a prediction model can output treatment recommendations for one or more patients not part of the training dataset.
[0064] The prediction model can include one or more prediction models. The model can include one or more machine learning algorithms. The machine learning algorithms can include support vector machines (SVMs), naive Bayes classification, random forests, neural networks (such as deep neural networks (DNNs)), recurrent neural networks (RNNs), deep RNNs, long short-term memory (LSTM) recurrent neural networks (RNNs), gated recurrent units (GRUs), gradient boosting machines, random forests, or other supervised learning algorithms or unsupervised machine learning, statistics, linear regression, k-nearest neighbors, k-means, decision trees, logistic regression, or any combination thereof. The model can be used for taxonomy or regression. The model can provide an estimated output of a set of one or more models that include multiple prediction models and utilize techniques such as gradient boosting, for example, in the construction of gradient boosting decision trees. The model can be trained using one or more training data sets that include one or more microbial features, patient data (e.g., patient history, patient's family history), patient vital signs (e.g., blood pressure, pulse, body temperature, blood oxygen saturation), or any combination thereof.
[0065] The prediction model can include any number of machine learning algorithms. In some embodiments, the random forest machine learning algorithm can be an ensemble of bagged decision trees. The ensemble can be at least about 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 500, 1000 or more bagged decision trees. The ensemble can be at most about 1000, 500, 250, 200, 180, 160, 140, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2 or fewer bagged decision trees. The ensemble can be about 1 to 1000, 1 to 500, 1 to 200, 1 to 100 or 1 to 10 bagged decision trees.
[0066] In some embodiments, the machine learning algorithm can have various parameters. The various parameters can be, for example, learning rate, mini-batch size, number of training epochs, momentum, learning weight decay, or number of neural network layers, etc.
[0067] In some embodiments, the learning rate can be about 0.00001 to 0.1.
[0068] In some embodiments, the mini-batch size can be about 16 to 128.
[0069] In some embodiments, the neural network can include neural network layers. The neural network can have at least about 2 to 1000 or more neural network layers.
[0070] In some embodiments, the number of training times can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 500, 1000, 10000 times or more.
[0071] In some embodiments, the momentum can be at least about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 or greater. In some embodiments, the momentum can be at most about 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1 or less.
[0072] In some embodiments, the learning weight decay can be at least about 0.00001, 0.0001, 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1 or greater. In some embodiments, the learning weight decay can be at most about 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, 0.0001, 0.00001 or less.
[0073] In some embodiments, the machine learning algorithm can use a loss function. The loss function can be, for example, regression loss, mean absolute error, mean bias error, hinge loss, Adam optimizer, and / or cross entropy, etc.
[0074] In some embodiments, the parameters of the machine learning algorithm can be adjusted with the aid of manual and / or computer systems.
[0075] In some embodiments, the machine learning algorithm can prioritize certain features. The machine learning algorithm can prioritize features that may be more relevant to detecting cancer, non-cancerous diseases, disorders, or any combination thereof. If a feature is classified more frequently than another feature when determining cancer, non-cancerous diseases, and / or disorders, then the feature may be more relevant to detecting cancer, non-cancerous diseases, and / or disorders. In some cases, features can be prioritized using a weighting system. In some cases, features can be prioritized based on the frequency and / or quantity of feature occurrences according to probability statistics. The machine learning algorithm can prioritize certain features with the aid of manual and / or computer systems.
[0076] In some cases, the prediction model can prioritize certain features to reduce computational costs, save processing power, save processing time, increase reliability, and / or reduce the use of random access memory, among other things.
[0077] The training dataset can be generated, for example, from one or more groups of patients diagnosed with common cancers, non-cancerous diseases, or medical conditions. The training dataset can include one or more microbial features in the form of the presence and / or abundance of microorganisms in an enriched biological sample (e.g., a hybridization-enriched biological sample) of one or more subjects. The features can include a cancer diagnosis of one or more subjects corresponding to the microbial features. In some cases, the features can include patient information such as patient age, patient medical history, other medical conditions, current or past medications, clinical risk scores, and time since last observation. For example, the set of features collected from a given patient at a given time point can be used together as a signature that can indicate the health status or condition of the patient at the given time point.
[0078] The label can include clinical outcomes such as the presence, absence, diagnosis, and / or prognosis of cancer, non-cancerous diseases, medical conditions, or combinations thereof in a subject (e.g., a patient). The clinical outcomes can include treatment efficacy (e.g., whether the subject responds positively or negatively to cancer- and / or disease-based treatments).
[0079] The input features can be constructed by binning the data or alternatively using one-hot encoding. The input can also include feature values or vectors derived from the aforementioned inputs, such as cross-correlation.
[0080] The training dataset can be constructed from the presence and / or abundance features of one or more microorganisms in a hybridization-enriched biological sample, or a combination of the presence and / or abundance features of one or more microorganisms and one or more somatic nucleic acid molecules in an enriched biological sample, where the features indicate cancer, non-cancerous diseases, medical conditions, or any combination thereof.
[0081] The model can process input features to generate output values that include one or more taxonomies, one or more predictions, or a combination thereof. For example, such taxonomies or predictions can include a binary taxonomy of the presence or absence of cancer; the presence of a non-cancerous disease; the presence of a disorder; or any combination of taxonomies for a subject. In some cases, the one or more prediction models (e.g., machine learning algorithms) can classify a subject among: a set of classification labels (e.g., 'no cancer, non-cancerous disease, and / or disorder', 'apparent cancer, non-cancerous disease, and / or disorder', and 'probable cancer, non-cancerous disease, and / or disorder'); the likelihood of developing a specific cancer, non-cancerous disease, and / or disorder (e.g., relative likelihood or probability); a score indicating the presence of cancer, non-cancerous disease, and / or disorder; 'risk factors' for the likelihood of patient death, confidence intervals for any numerical prediction; or any combination thereof. Various machine learning techniques can be cascaded such that the output of a machine learning technique can also be used as an input feature for a subsequent layer or segment of the model.
[0082] To train the model (e.g., by determining the weights and correlations of the model) to generate real-time taxonomies or predictions, the model can be trained using the training datasets and / or one or more training features described elsewhere in this document. Such datasets and / or features may be large enough to generate statistically significant classifications or predictions. For example, the dataset can include a database of data that includes the presence and / or abundance of fungi, viruses, archaea, microorganisms, bacteria, or any combination thereof in biological samples of one or more subjects.
[0083] The dataset can be divided into subsets (e.g., discrete or overlapping), such as a training dataset, a development dataset, and a test dataset. For example, the dataset can be divided into a training dataset that includes 80% of the dataset, a development dataset that includes 10% of the dataset, and a test dataset that includes 10% of the dataset. The training dataset can include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The development dataset can include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The test dataset can include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, leave-one-out cross-validation can be employed. The training set (e.g., the training dataset) can be selected by randomly sampling the dataset corresponding to one or more patient groups to ensure sampling independence. Alternatively, the training set (e.g., the training dataset) can be selected by proportionally sampling the dataset corresponding to one or more patient groups to ensure sampling independence.
[0084] To improve the accuracy of model prediction and reduce overfitting of the model, the dataset can be augmented to increase the number of samples in the training set. For example, data augmentation can include rearranging the order of observations in the training records. To accommodate datasets with depleted observations, methods for inferring depleted data can be used, such as forward filling, backward filling, linear interpolation, and / or multi-task Gaussian processes. The dataset can be filtered and / or batch corrected to eliminate or mitigate confounding factors. For example, within a database, a subset of patients can be excluded.
[0085] The model can include one or more neural networks, such as neural networks, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), and / or deep RNNs. The recurrent neural network can include units that can be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, as described in other parts of this document, the model can include an algorithmic architecture that includes a neural network with a set of input features (e.g., microbial features, patient and / or subject vital measurements, patient and / or subject medical history, patient and / or subject demographics, or any combination thereof). Neural network techniques, such as dropout or regularization, can be used during model training to prevent overfitting. The neural network can include multiple sub-networks, each of which is configured to generate a classification or prediction of a different type of output information, and the classifications or predictions can be combined to form the overall output of the neural network. The machine learning model can alternatively utilize statistical or correlation algorithms, including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, any combination thereof, and / or their ensemble and gradient boosting variants.
[0086] When the model generates a classification or prediction of cancer, non-cancerous diseases, conditions, or a combination thereof, a notification (e.g., an alert or alarm) can be generated and transmitted to healthcare providers, such as doctors, nurses, and / or other members of the patient's treatment team within a hospital. The notification can be transmitted via an automated phone call, a short message service (SMS) or multimedia message service (MMS) message, an email, and / or an alert within a dashboard. The notification can include output information, such as a prediction of cancer, non-cancerous diseases, and / or conditions; the likelihood of the predicted cancer, non-cancerous diseases, and / or conditions; the time before the expected onset of cancer, non-cancerous diseases, and / or conditions; the confidence interval of the likelihood or time of cancer, non-cancerous diseases, and / or conditions, recommended treatment options, or any combination of such information. In some cases, the time until the expected onset of cancer can include a time of at least 1 year, at least 2 years, or at least 3 years from the detection of a pre-cancerous lesion.
[0087] To verify the performance of a model, different performance metrics can be generated. For example, the area under the receiver operating characteristic curve (AUROC) can be used to determine the diagnostic, prognostic, screening, or any combination thereof capabilities of the model. For example, the model can use an adjustable taxonomic threshold such that specificity and sensitivity are adjustable, and the receiver operating characteristic curve (ROC) can be used to identify different operating points corresponding to different specificity and sensitivity values.
[0088] In some cases, such as when the dataset is not large enough, cross-validation can be performed to evaluate the robustness of the model in different training and test datasets.
[0089] To calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), area under the precision-recall curve (AUPR), AUROC, or any combination thereof, the following definitions can be used. A "false positive" can refer to a result that incorrectly or prematurely generates a positive result or consequence (e.g., before the actual onset of cancer, non-cancerous disease, and / or disorder or without any onset). A "true positive" can refer to a result where a positive result or consequence has been correctly generated when the patient has cancer, non-cancerous disease, and / or disorder (e.g., the patient shows symptoms of cancer, non-cancerous disease, and / or disorder, or the patient's record indicates cancer, non-cancerous disease, and / or disorder). A "false negative" can refer to a result where a negative result and / or consequence has been generated, but the patient has cancer, non-cancerous disease, and / or disorder (e.g., the patient shows symptoms of cancer, non-cancerous disease, and / or disorder, or the patient's record indicates cancer, non-cancerous disease, and / or disorder). A "true negative" can refer to a result where a negative result or consequence has been generated (e.g., before the actual onset of cancer, non-cancerous disease, and / or disorder or without any onset).
[0090] The model can be trained until certain predefined conditions of accuracy or performance are met, such as having a minimum expected value corresponding to a diagnostic accuracy measurement. For example, the diagnostic accuracy measurement can correspond to a prediction of the likelihood of occurrence of cancer, non-cancerous disease, and / or disorder in a subject. As another example, the diagnostic accuracy measurement can correspond to a prediction of the likelihood of deterioration and / or recurrence of cancer, non-cancerous disease, and / or disorder in a subject who has previously received treatment. Examples of diagnostic accuracy measurements can include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, AUPR, and AUROC corresponding to the diagnostic accuracy of detecting or predicting cancer, non-cancerous disease, and / or disorder.
[0091] For example, such a predetermined condition may be that the sensitivity for predicting cancer, non-cancerous diseases, and / or disorders includes a value of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0092] As another example, such a predetermined condition may be that the specificity for predicting cancer, non-cancerous diseases, and / or disorders includes a value of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0093] As another example, such a predetermined condition may be that the positive predictive value (PPV) for predicting cancer, non-cancerous diseases, and / or disorders includes a value of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0094] As another example, such a predetermined condition may be that the negative predictive value (NPV) for predicting cancer, non-cancerous diseases, and / or disorders includes a value of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0095] As another example, such a predetermined condition may be that the area under the curve (AUC) (AUROC) of the receiver operating characteristic (ROC) curve for predicting cancer, non-cancerous diseases, and / or disorders includes a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0096] As another example, such a predetermined condition can be that the area under the precision-recall curve (AUPR) for predicting cancer, non-cancerous diseases, and / or disorders includes a value of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0097] In some embodiments, the trained model can be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0098] In some embodiments, the trained model can be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0099] In some embodiments, the trained model can be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with a positive predictive value (PPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0100] In some embodiments, the trained model can be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with a negative predictive value (NPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0101] In some embodiments, the trained model can be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with an area under the curve (AUC) of the receiver operating characteristic (ROC) curve (AUROC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0102] In some embodiments, the trained model can be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with an area under the precision-recall curve (AUPR) of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0103] The training dataset can be collected from training subjects (e.g., humans). Each training dataset of the training dataset can include the corresponding diagnostic status of the given training data of the subject, the diagnostic status indicating that the subject has been diagnosed with a disease or has not been diagnosed with a disease (e.g., cancer, non-cancerous diseases, and / or disorders).
[0104] In some embodiments, the model is a neural network or a convolutional neural network. See Vincent et al., 2010, "Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion", Journal of Machine Learning Research 11, pp. 3371-3408; Larochelle et al., 2009, "Exploring strategies for training deep neural networks", Journal of Machine Learning Research 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of the references being hereby incorporated by reference.
[0105] In some embodiments, independent component analysis (ICA) is used to reduce the dimensionality of the data, as described in Lee, T.-W. (1998): Independent Component Analysis: Theory and Applications, Boston, Mass.: Kluwer Academic Publishers, ISBN 0-7923-8261-7; and A.; Karhunen, J.; Oja, E. (2001): Independent Component Analysis, New York: Wiley, ISBN 978-0-471-40540-5, each of the references being hereby incorporated by reference in its entirety.
[0106] In some embodiments, principal component analysis (PCA) is used to reduce the dimensionality of data, such as Jolliffe, I.T. (2002). Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag. Doi: 10.1007 / b98835. ISBN 978-0-387-95442-4, which is hereby incorporated by reference in its entirety.
[0107] SVM is described in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines”, Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, “Statistical Learning Theory”, Wiley, New York; Mount, 2001, “Bioinformatics: sequence and genome analysis”, Cold Spring Harbor Laboratory Press, Cold Spring Harbor; Duda, “Pattern Classification”, 2nd Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, “The Elements of Statistical Learning”, Springer, New York; and Furey et al., 2000, “Bioinformatics” 16, 906-914, each of the references being hereby incorporated by reference in its entirety. When used for classification, an SVM separates a given binary-labeled data set from a hyperplane that is maximally distant from the labeled data. For cases where there may not be a linear separation, an SVM can operate in conjunction with a “kernel” technique that automatically implements a non-linear mapping of the feature space. The hyperplane found by the SVM in the feature space corresponds to a non-linear decision boundary in the input space.
[0108] A decision tree is described generally by: Duda, 2001, *Pattern Classification*, John Wiley & Sons, New York, pages 395-396, which reference is hereby incorporated by reference. Tree-based methods partition the feature space into a set of rectangles and then fit a model (such as a constant) in each rectangle. In some embodiments, the decision tree is a random forest regression. One particular algorithm that can be used is Classification and Regression Trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, *Pattern Classification*, John Wiley & Sons, New York, pages 396-408 and 411-412, which reference is hereby incorporated by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, *The Elements of Statistical Learning*, Springer-Verlag, New York, Chapter 9, which reference is hereby incorporated by reference in its entirety. Random forest is described in Breiman, 1999, “Random Forests—Random Features”, *Technical Report* 567, Statistics Department, U.C. Berkeley, September 1999, which document is hereby incorporated by reference in its entirety.
[0109] Clustering (e.g., unsupervised clustering model algorithms and supervised clustering model algorithms) is described in Duda and Hart, "Pattern Classification and Scene Analysis", 1973, John Wiley & Sons, New York, pages 211-256 (hereinafter referred to as "Duda 1973"), which is hereby incorporated by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as the problem of finding natural groupings in a dataset. To identify natural groupings, two problems are solved. First, a way to measure the similarity (or dissimilarity) between two samples is determined. This metric (similarity measure) is used to ensure that samples within one cluster are more similar to each other than samples in other clusters. Second, a mechanism for using the similarity measure to partition the data into clusters is determined. Similarity measures are discussed in Section 6.7 of Duda 1973, where it is stated that one way to begin a clustering investigation is to define a distance function and compute a matrix of distances between all pairs of samples in the training set. If the distance is a good similarity measure, the distances between reference entities within the same cluster will be significantly less than the distances between reference entities in different clusters. However, as stated on page 215 of Duda 1973, clustering does not require the use of a distance measure. For example, a non-metric similarity function s(x, x') can be used to compare two vectors x and x'. Generally, s(x, x') is a symmetric function with a larger value when x and x' are "similar" to some extent. Page 218 of Duda 1973 provides an example of an asymmetric similarity function s(x, x'). Once a method for measuring "similarity" or "dissimilarity" between points in a dataset has been selected, clustering requires a criterion function for measuring the clustering quality of any partition of the data. The partition of the dataset that extremizes the criterion function is used to cluster the data. See page 217 of Duda 1973. Criterion functions are discussed in Section 6.8 of Duda 1973. Recently, Duda et al., "Pattern Classification", 2nd Edition, John Wiley & Sons, New York, pages 537-563, has been published which describes clustering in detail.More information about clustering techniques can be found in Kaufman and Rousseeuw, 1990, *Finding Groups in Data: An Introduction to Cluster Analysis*, Wiley, New York, N.Y.; Everitt, 1993, *Cluster analysis* (3rd ed.), Wiley, New York, N.Y.; and Backer, 1995, *Computer-Assisted Reasoning in Cluster Analysis*, Prentice Hall, Upper Saddle River, N.J., each of the references being hereby incorporated by reference. Specific exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor algorithm, farthest neighbor algorithm, average linkage algorithm, centroid algorithm, or sum of squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering. In some embodiments, the clustering includes unsupervised clustering, where no preconceived notion of what clusters should be formed is imposed when clustering a training set.
[0110] Regression models, such as regression models for multinomial logistic models, are described in Agresti, *An Introduction to Categorical Data Analysis*, 1996, John Wiley & Sons, New York, Chapter 8, which reference is hereby incorporated by reference in its entirety. In some embodiments, the model utilizes regression models disclosed in Hastie et al., 2001, *The Elements of Statistical Learning*, Springer-Verlag, New York, which reference is hereby incorporated by reference in its entirety. In some embodiments, gradient boosting models are used for, e.g., the taxonomy algorithms described herein; these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019). "Gradient Boosting". *Hands-On Machine Learning with R*. Chapman and Hall. pp. 221-245. ISBN 978-1-138-49568-5., which reference is hereby incorporated by reference in its entirety. In some embodiments, ensemble modeling techniques are used; these ensemble modeling techniques are described in the embodiments of the classification models herein and in Zhou Zhihua (2012). *Ensemble Methods: Foundations and Algorithms*. Chapman and Hall / CRC. ISBN 978-1-439-83003-1, which reference is hereby incorporated by reference in its entirety.
[0111] In some embodiments, the machine learning analysis is performed by a device executing one or more programs (e.g., one or more programs stored in non-persistent memory or persistent memory) that include instructions for performing data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., a processing core) and a memory (e.g., one or more programs stored in non-persistent memory or persistent memory) that includes instructions for performing data analysis. In some embodiments, the prediction model may be stored on a computer memory and executed by one or more processors of the system, as described elsewhere herein.
[0112] System
[0113] In some embodiments, the present disclosure describes a computer system that can be programmed to implement one or more of the methods of the present disclosure, as described elsewhere herein. Figure 6A computer system 600 is shown, which can be programmed or otherwise configured to predict cancer, non-cancerous diseases, or a combination thereof; train a prediction model; generate a recommended treatment plan; or any combination of methods as described elsewhere herein. The computer system 600 can be the user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0114] In some embodiments, the computer system 600 can include one or more central processing units (CPUs, also referred to herein as "processors" and "computer processors") 606, which can be a single-core or multi-core processor or multiple processors for parallel processing. In some embodiments, the computer system 600 can include a memory and / or memory location 604 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 602 (e.g., hard disk), a communication interface 608 for communicating with one or more other systems (e.g., network adapter), and / or peripheral devices 610, such as caches, other memories, data storage devices, and / or electronic display adapters. The memory 604, storage unit 602, interface 608, and peripheral devices 610 can communicate with the CPU 606 via a communication bus (solid lines), such as a motherboard. The storage unit 602 can be a data storage unit (or data repository) for storing data. The computer system 600 can be operably coupled to a computer network ("network") 612 via the communication interface 608. The network 612 can be the Internet, an intranet, and / or an extranet or an intranet and / or extranet communicating with the Internet. In some cases, the network 612 can include a telecommunications network and / or a data network. The network 612 can include one or more computer servers, which can implement distributed computing such as cloud computing. In some cases, the network 612 can implement a peer-to-peer network via the computer system 600, which can enable devices to be coupled to the computer system 600 to act as clients or servers.
[0115] The CPU 606 can execute a series of machine-readable instructions, such as one or more steps of the methods described elsewhere herein, and the series of machine-readable instructions can be embodied in a program or software. The instructions can be stored in a memory location such as the memory 604. The instructions can be directed to the CPU 606, which can then program or otherwise configure the CPU 606 to implement the methods of the present disclosure described elsewhere herein. Examples of operations performed by the CPU 606 can include fetching, decoding, executing, and / or writing back.
[0116] The CPU 606 can be part of a circuit such as an integrated circuit. One or more other components of the system 600 can be included in the circuit. In some cases, the circuit can include an application specific integrated circuit (ASIC).
[0117] The storage unit 602 can store files such as drivers, libraries, and / or saved programs (e.g., including one or more steps of the methods described in other parts of this document). The storage unit 602 can store user data, such as classifications and / or predictions provided by one or more models, training user data, test user data, user preferences, and user programs, or any combination thereof. In some cases, the computer system 600 can include one or more additional data storage units that are external to the computer system 600, such as located on a remote server that communicates with the computer system 600 via an intranet or the Internet.
[0118] The computer system 600 can communicate with one or more remote computer systems via the network 612. For example, the computer system 600 can communicate with a user's remote computer system. Examples of remote computer systems can include personal computers (e.g., portable PCs), tablet computers or tablet PCs (e.g., iPad, Galaxy Tab), telephones, smartphones (e.g., iPhone, Android - enabled devices, ) or personal digital assistants. A user can access the computer system 600 via the network 612.
[0119] The methods described herein can be implemented by machine (e.g., one or more processors) executable code stored on an electronic storage location of the computer system 600 (e.g., on the memory 604 or the electronic storage unit 602). The machine executable code or machine readable code can be provided in the form of software. During use, the code can be executed by one or more processors 606. In some cases, the code can be retrieved from the storage unit 602 and stored on the memory 604 for ready access by the processor 606. In some scenarios, the electronic storage unit 602 can be excluded and machine executable instructions can be stored on the memory 604.
[0120] The code can be pre - compiled and configured for use with a machine having a processor adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language, and the programming language can be selected to enable the execution of the code in a pre - compiled or as - compiled manner.
[0121] In some embodiments, a system as described elsewhere herein may include: one or more processors; and a non-transitory computer-readable storage medium containing software, where the software includes executable instructions that, as a result of execution, cause the one or more processors of the computer system to: receive one or more nucleic acid molecule sequencing reads of a biological sample of a subject, where the subject has a disease, and where the one or more nucleic acid molecule sequencing reads are obtained from one or more nucleic acid molecules enriched by one or more probes exposed to the biological sample of the subject; map the one or more nucleic acid molecule sequencing reads to a genomic database, thereby identifying one or more non-human sequencing reads among the one or more nucleic acid molecule sequencing reads; and identify one or more microbial signatures of the one or more non-human sequencing reads to classify the disease of the subject.
[0122] Aspects of the systems and methods described elsewhere herein, such as computer system 600, may be embodied in programming. Aspects of the technology may be considered a “product” or “article of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine-readable medium. The machine executable code may be stored on an electronic storage unit such as a memory (e.g., read only memory, random access memory, flash memory) or a hard disk. A “storage” type medium may include any or all of the tangible memories of a computer, processor, etc., or associated modules thereof, such as various semiconductor memories, tape drives, hard disk drives, etc., that may provide non-transitory storage for software programming at any time. All or part of the software may sometimes be communicated via the Internet or various other telecommunications networks. Such communication, for example, may enable loading of the software from one computer or processor to another, such as from a management server or host computer to the computer platform of an application server. Thus, another type of medium that may carry software elements includes light waves, radio waves, and electromagnetic waves, such as used across physical interfaces between local devices via wired and optical landlines as well as via various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry software. As used herein, unless limited to non-transitory tangible “storage” media, terms such as computer or machine “readable media” refer to any medium that participates in providing instructions to a processor for execution.
[0123] Thus, machine-readable media such as computer-executable code can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical discs or magnetic disks, such as any storage device in any computer, such as may be used to implement databases shown in the accompanying drawings. Volatile storage media can contain dynamic memory, such as the main memory of such a computer platform. Tangible transmission media can include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier transmission media can take the form of electrical signals or electromagnetic signals, or acoustic or light waves, such as in the form of signals generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punched card tapes, any other physical storage media with hole patterns, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carrier-transmitted data or instructions, cables or links that transmit such carriers, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0124] Computer system 600 may include or communicate with an electronic display 614 that includes a user interface (UI) 616 for providing, for example, a display for visualization of prediction results, and the UI may include one or more interfaces and / or panels for training a prediction model, managing and / or manipulating subject and / or patient data, or any combination thereof. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0125] ***
[0126] While the preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0127] Although the steps of the methods as described and claimed herein show and / or describe each method or group of operations according to embodiments, those of ordinary skill in the art will recognize many variations based on the teachings described herein. The steps may be completed in a different order. Steps may be added or omitted. Some steps may contain sub-steps. Many steps may be repeated multiple times as long as beneficial. Two or more steps may be performed simultaneously.
[0128] One or more steps of each method or group of operations may be performed by circuitry as described herein, e.g., one or more processors or logic circuitry such as programmable array logic or field programmable gate arrays. The circuitry may be programmed to provide one or more steps of each method or group of operations, and the program may include program instructions stored on a computer-readable memory or programming steps of the logic circuitry such as programmable array logic or, e.g., field programmable gate arrays.
[0129] Definitions
[0130] Unless otherwise defined, all special terms, symbols, and other technical and scientific terms used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not be construed as representing a substantial difference from the commonly understood meaning in the art.
[0131] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an immutable limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as the individual numerical values within that range. For example, a range description such as 1 to 6 should be considered to have specifically disclosed sub-ranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as the individual numbers within that range, e.g., 1, 2, 3, 4, 5, and 6. This applies regardless of the width of the range.
[0132] As used in the specification and claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "sample" includes a plurality of samples, including mixtures thereof.
[0133] The terms "determine", "measure", "evaluate", "assess", "assay", and "analyze" are generally used interchangeably herein to refer to various forms of measurement. The terms include determining the presence of an element (e.g., detecting). These terms can include quantitative, qualitative, or both quantitative and qualitative determinations. An assessment can be relative or absolute. "Detecting presence" can include determining the amount of something present in addition to determining whether something is present or absent depending on the context.
[0134] The terms "subject", "individual", or "patient" are generally used interchangeably herein. A "subject" can be a biological entity containing expressed genetic material. The biological entity can be a plant, animal, or microorganism, including, for example, bacteria, viruses, fungi, and protozoa. A subject can be a tissue, cell, and progeny thereof of a biological entity obtained in vivo or cultured in vitro. The subject can be a mammal. The mammal can be a human. A subject can be diagnosed or suspected of being at high risk of a disease. In some cases, a subject is not necessarily diagnosed or suspected of being at high risk of a disease.
[0135] As used herein, the term "in vivo" is used to describe an event that occurs within the body of a subject.
[0136] As used herein, the term "ex vivo" is used to describe an event that occurs outside the body of a subject. Ex vivo assays are not performed on the subject. Instead, ex vivo assays are performed on a sample isolated from the subject. An example of an ex vivo assay performed on a sample is an "in vitro" assay.
[0137] As used herein, the term "in vitro" is used to describe an event that occurs in a container used to hold laboratory reagents such that it is separated from the biological source from which the material was obtained. In vitro assays can encompass cell-based assays in which live or dead cells are used. In vitro assays can also encompass cell-free assays that do not use intact cells.
[0138] As used herein, the terms "metagenome", "metagenomes", and "metagenomic" are used to refer to the sum total of all genomic information represented in a sample without regard to the species of origin and thus include human and non-human genomic information.
[0139] As used herein, the term "metagenome assembly" refers to the process of reconstructing microbial genomes from metagenomic sequencing data, and "metagenome assembly" is used to refer to the product of the metagenome assembly process.
[0140] As used herein, the term "contig" is used to refer to a non-redundant nucleic acid genomic sequence formed by joining one or more smaller sequences based on sequence overlap. In some cases, the length of a contig can be from about several kilobases (kb) to about several hundred kb.
[0141] As used herein, the term "bin" is used to refer to a collection of one or more contigs grouped together due to the likelihood of originating from the same parental genome.
[0142] As used herein, the term "bin abundance" or "bin abundances" is used to refer to the number of nucleic acid sequencing reads aligned to one or more bins.
[0143] As used herein, the terms "nodule", "neoplasm", and "tumor" are generally used interchangeably herein and are intended to denote an abnormal growth of tissue, regardless of its malignant state.
[0144] As used herein, the term "about" with respect to a number means that number plus or minus 10% of that number. The term "about" with respect to a range means that range minus 10% of its lowest value and plus 10% of its highest value.
[0145] As used herein, the use of absolute or sequential terms, e.g., "will", "will not", "should", "should not", "must", "cannot", "first", "initially", "next", "subsequently", "before", "after", "last", and "finally" are not intended to limit the scope of the embodiments disclosed herein, but are exemplary.
[0146] Any system, method, software, composition, and platform described herein are modular and not limited to sequential steps. Thus, terms such as "first" and "second" do not necessarily imply a priority, order of importance, or order of acts.
[0147] As used herein, the terms "treatment" or "treating" are used to refer to a drug or other intervention regimen used to obtain a beneficial or desired result in a recipient. Beneficial or desired results include, but are not limited to, therapeutic benefits and / or prophylactic benefits. A therapeutic benefit may refer to eradicating or ameliorating symptoms or symptoms of an underlying condition being treated. Additionally, a therapeutic benefit may be achieved by eradicating or ameliorating one or more of the physiological symptoms associated with the underlying condition such that an improvement is observed in the subject, although the subject may still have the underlying condition. Prophylactic effects include delaying, preventing or eliminating the onset of a disease or medical condition, delaying or eliminating the onset of symptoms of a disease or medical condition, slowing, halting or reversing the progression of a disease or medical condition or any combination thereof. For prophylactic benefits, a subject at risk of developing a particular disease or reporting one or more physiological symptoms of a disease may receive treatment even if a diagnosis of such disease may not yet have been made.
[0148] The subsection headings used herein are for purposes of arrangement only and should not be construed as limiting the subject matter described.
[0149] Examples
[0150] Example 1: Early Detection of Lung Cancer Using Blood and Tissue Metagenomics
[0151] The cell-free tumor DNA fraction of most patients with stage I lung cancer is less than 0.1%, precluding sensitive assessment of host-based diagnosis that relies on separating normal from tumor-derived molecules. As described in Example 1, an alternative and orthogonal cell-free metagenome-centric approach was investigated that circumvents the limitations of host-centric approaches and provides state-of-the-art detection of lung cancers as small as 8 mm in diameter. De novo metagenomic assembly was performed on 5187 whole-genome sequenced blood and tissue samples, 1562 pan-cancer bins were identified, and their diagnostic capabilities and cancer type specificity were demonstrated in 17,079 samples across four independent cohorts. Plasma sequencing based on amplicons (16S rRNA) was also found to approximate the diagnostic performance of shotgun-sequenced genomes, suggesting cost-effective opportunities while validating the method. Additionally, since this strategy is independent of and complementary to host-centric approaches, a clinical protein metagenomic diagnostic test was developed and validated that outperformed PET-CT in diagnostic performance and clinical risk scoring in a blinded validation cohort of 106 stage I cancers and lung diseases of diverse etiologies (mean ± SD = 1.91 ± 0.84 cm). Overall, the findings described herein firmly establish the utility of cell-free metagenomics as a generalizable and sensitive strategy for early cancer detection, warranting application in additional cancer types.
[0152] Existing liquid biopsy methods for lung cancer have size-dependent sensitivity, which precludes optimal detection of early disease. Orthogonal, cancer-associated microbial biomarkers can overcome some of these limitations. The study described in Example 1: (i) describes a metagenome-centric lung cancer diagnostic test, (ii) performs metagenomic assembly in cancer blood and tissue, (iii) evaluates amplicon-based plasma sequencing for cancer detection, and (iv) demonstrates the superiority of this method relative to PET-CT in the context of stage I disease.
[0153] Until recently, cancer was primarily considered a sterile disease of the human genome, and thus methods for diagnosing its presence and / or type relied on detecting human-centric biomarkers. Although a variety of diagnostic strategies have been developed, ranging from low-complexity PCR-based mutanomes to high-complexity assays of millions of DNA fragments, the size of cancer lesions inherently limits the availability and abundance of cancer-derived molecules in the circulation and the sensitivity of the accompanying tests. These challenges can be partially alleviated by measuring more biomarkers, deeper sequencing, and / or more frequent testing (e.g., monthly), but ultimately there are hard biological limits on how much cancer-derived material is present in a vial of blood taken during early disease, which has precluded widespread diagnosis of these tumors.
[0154] Among cancers that are difficult to detect, lung cancer is the leading cause of cancer-related death worldwide and is often detected in advanced disease. Despite recent advances in liquid biopsy, blood tests that have proven capable of detecting stage I lung cancer lesions, particularly those ≤3 cm in diameter (i.e., clinical stage T1), have been difficult. For example, in this context, a recent validation cohort of a circulating tumor DNA (ctDNA)-based test, Lung-CLiP, achieved an area under the receiver operating characteristic curve (AUROC) of only 0.69 in stage I cancer versus risk-matched controls. Another recent fragmentomics diagnosis showed an AUROC of 0.76 for stage I cancer in its discovery cohort, although performance in its validation cohort was not reported, other than approximately 50% sensitivity at 80% specificity. Apart from PET-CT scans, the only blood test clinically available for determining early malignancy of small nodules (≤3 cm) is an integrated clinical proteomics panel, with an AUROC of 0.76 in the validation cohort. These results underscore the need to approach early lung cancer detection in different ways.
[0155] An intracellular cancer type-specific community of fungi, bacteria, and viruses within tumors, whose genomic content can be detected in plasma-derived cell-free DNA (15, 16), describes a method, as described in Example 1 and elsewhere herein, for addressing early lung cancer detection through microbial biomarkers that are independent of and complementary to host-centric approaches. First, the ability of a parsimonious signature containing 300 plasma-derived, cancer-associated, fungal, and bacterial biomarkers to diagnose early disease in a low-risk clinical setting of 808 never-treated patients was evaluated, showing state-of-the-art performance and validating the amplicon-based plasma sequencing method. For the high-risk patient setting, metagenomic bins of cancer-associated microbial DNA were developed using 5187 whole-genome sequenced blood and tissue samples. Compared to 300 known cancer-associated microbial biomarkers, these metagenomic bins provided better diagnostic performance for untreated lung disease samples of 222 different etiologies; furthermore, by analyzing their abundances in 17,079 samples from four independent cohorts, the metagenomic bins were found to have broad cancer type-specificity and could be used for early lung cancer identification. Finally, using a subset of 454 patients from a cohort with matched clinical and imaging metadata, a clinical proteomic metagenomic diagnosis was established and validated, providing state-of-the-art diagnostic performance in a blinded validation cohort of 106 patients with nodules as small as 6 mm. Through this study, as described in Example 1, the utility and generalizability of plasma-derived metagenomics for early cancer detection were demonstrated.
[0156] Subjects
[0157] All subject samples were obtained retrospectively from biobank cohorts at New York University, University of California, San Francisco, Miami Cancer Institute, University of North Carolina at Chapel Hill, and two commercial sources. Sample collection was performed according to protocols approved by the institutional review boards of the respective institutions; all subjects signed written informed consent to provide samples for the study. Subjects with lung cancer or lung nodules included those with a suspicious nodule detected by chest imaging (low-dose computed tomography or chest X-ray) who underwent diagnostic biopsy or surgical resection. Histopathological diagnosis was performed on all biopsied / resected lung neoplasms, and lung cancer subtypes and stages were assigned to malignant neoplasms based on histopathology. The disease severity of subjects with sarcoidosis or COPD was scored according to the Scadding staging system and the GOLD criteria, respectively. Patients with a prior history of cancer or recent (within the last 2 months) use of antibiotics were excluded from the study.
[0158] 454 samples (44%, 342 lung cancers, 112 non-cancers) were initially collected at the NYU Lung Cancer Biomarker Center from subjects with lung nodules who underwent diagnostic biopsies or tissue resections during the period from 2006 to 2021. Plasma samples from 89 subjects were obtained from SubPopulations and InteRmediate Outcome Measures. In the COPD study, SPIROMICS; all 89 subjects were current smokers, 47 were diagnosed with GOLD stage II (mild) COPD, and the remaining 42 were GOLD stage III (severe) COPD. Samples from 25 subjects with sarcoidosis were obtained from the UCSF Sarcoidosis Research Project; 12 / 25 were Scadding stage 2 sarcoidosis, and 13 / 25 were Scadding stage 4 sarcoidosis. Among the 402 cancer samples, the two major histological subtypes of lung cancer - lung adenocarcinoma (LUAD) and squamous cell carcinoma (LUSC) accounted for 57.2% (230) and 29.8% of the cancer samples, respectively. The number (percentage) of cancers represented by pathological stage was as follows: stage I, 131 (32.59%); stage II, 110 (27.36%); stage III, 98 (24.38%); stage IV, 57 (14.18%); unknown, 6 (1.49%). The smoking status (current / former / never) was known for 708 (68.73%) of the cohort subjects. Shotgun metagenomic sequencing was performed on all CODICES cohort samples (1030), and targeted 16S rRNA gene sequencing was performed on a subset of 335 samples (142 lung cancers, 97 lung diseases, and 96 healthy). All cancer plasma samples were obtained from untreated patients, and the age (±5 years) and sex of all disease state samples were matched to plasma samples from healthy donors from two commercial sources.
[0159] Sample Processing and Sequencing Library Preparation
[0160] Total circulating DNA was extracted from 400 μL of plasma from each sample using the QIAamp Circulating Nucleic Acid Kit (QIAGEN 55114) according to the manufacturer's instructions and purified with Agencourt AMPure XP beads (Beckman Coulter). Sequencing libraries were prepared from cfDNA using the KAPA HyperPlus Kit (Roche Diagnostics) and unique dual-index primers. The final libraries were analyzed using the Agilent 4200 TapeStation system (High Sensitivity DNA Kit) and quantified by qPCR using Illumina's NEBNext Library Quant Kit (New England Biolabs). Paired-end 2 × 150-bp sequencing was performed on a NovaSeq 6000 instrument, S4 flow cell (Illumina).
[0161] 16S sequencing
[0162] The V6 region of the 16S ribosomal RNA was targeted using primers 967F (0.3 μM, 5'-CNACGCGAAGAACCTTANC-3') and 1064R (0.3 μM, 5'-CGACRRCCATGCANCACCT-3'), and initial amplification was performed by PCR using 2X KAPA HiFi HotStart ReadyMix (Kapa Biosystems, Boston, MA, USA) and 10 ng of input DNA. The cycling program was set as follows: 5 minutes at 95 °C, 20 cycles of 20 seconds at 98 °C, 15 seconds at 52 °C, 15 seconds at 72 °C, and a final extension of 5 minutes at 72 °C. These PCR products were subjected to two subsequent 10-cycle PCRs as described by Glenn et al. (39) to introduce universal adapters and unique sample identification dual indices suitable for the Illumina sequencing platform. The PCR conditions were the same as before, except that the annealing temperature reached 60 °C and the primer concentration was increased to 0.5 μM. All PCR products were purified using AMPure XP beads (Beckman Coulter, #A63881) following the manufacturer's instructions. The final library was quantified using Qubit DNA Assay (Thermo Fisher Scientific) and by qPCR using Illumina's NEBNext Library Quant Kit (New England Biolabs), and its quality was examined on an Agilent 4200 TapeStation system. Sequencing analysis was performed on an Illumina MiSeq instrument using kit v2 (Illumina, #MS-102-2002).
[0163] Shotgun metagenomic sequencing analysis
[0164] The reads were quality-filtered using fastp to remove adapter sequences and low-quality reads. Reads that aligned to the human genome were separated from non-human reads by aligning them against the complete human reference genome T2T-CHM13 (v2.0) using Bowtie2 with the sensitive parameter set. Reads that did not align to the human genome were removed using vsearch. The non-human reads were aligned against the RefSeq database version 206 using Bowtie2 with the sensitive parameter set. The abundances of each microorganism were tallied from the alignments and used for downstream machine learning analysis.
[0165] Plasma proteome analysis
[0166] The Bio-Plex 200 platform (Bio-Rad, Hercules, CA) was used to evaluate the levels of target proteins in human plasma samples. Briefly, plasma samples were centrifuged, diluted 1:2, and subjected to a Milliplex bead-based immunoassay (Millipore Sigma, Burlington, MA) according to the manufacturer's protocol. The Milliplex HCCBP1MAG-58K (Millipore Sigma, Burlington, MA) panel was used to detect the following protein analytes: cancer antigen 15-3 (CA15-3), cancer antigen 19-9 (CA19-9), carcinoembryonic antigen (CEA), cancer antigen 125 (CA 125), interleukin-8 (IL-8), prolactin (PRL), cytokeratin 19 fragment (CYFRA 21-1), and osteopontin (OPN). The concentration of each protein analyte was determined using 5-parameter logistic curve fitting available in Bio-Plex Manager 6.2 software (Bio-Rad, Hercules, CA) and protein standards provided by the Milliplex assay.
[0167] TCGA and Plasma Cell-Free Microbial DNA Metagenomic Assembly, Binning, and Analysis
[0168] First, whole-genome sequencing (WGS) samples were co-assembled by cancer and sample type preprocessing (i.e., trimming, quality control, and human read filtering). Co-assembly by sample and cancer type was motivated by past findings that the distribution of microbes varies by cancer type and that most individual samples do not have sufficient read depth to assemble microbial genomes. Co-assembly was performed by metaSPAdes (v.3.13.1) with 1TB RAM limit allocated across ten threads and k-mer sizes of 21, 33, 55, 77, 99, and 127. Contigs greater than 1500 in length were filtered from each co-assembly and separated into prokaryotic or eukaryotic origin using EukRep (v.0.6.6). Binning of each set of prokaryotic or eukaryotic contigs was performed using the default parameters of the jgi_summarize_bam_contig_depths function of MetaBAT2 (v.2.12.1) on the contig abundance profiles independently estimated for each sample using Vamb (v.3.0.3). The abundance profile of each sample was estimated by mapping reads against the binned contigs using the Salmon (v.0.13.1) quant function with the --meta flag enabled. Quality metrics of the resulting prokaryotic exact bins were calculated using CheckM (v.1.0.13). All prokaryotic bins were filtered based on CheckM statistical completeness greater than 10% and contamination less than 5%.
[0169] Microbial diversity analysis
[0170] α and β diversities were calculated using Qiime2 on the microbial abundance table of Rep206 species, filtered to overlap with the WIS dataset. For α diversity, samples were first thinned to the minimum number that retained the upper quartile of the samples. The statistical differences between groups were calculated using the Mann-Whitney-Wilcoxon test two-sided test. The β diversity distances and ordinations were calculated using the robust Aitchison PCA metric of DEICODE. The PERMANOVA test was used to determine the statistical differences between groups based on β diversity.
[0171] 16S sequencing analysis
[0172] The raw 16S reads were quality filtered using fastp to remove adapter sequences and low-quality reads. Unique 16S sequences were identified using Deblur within Qiime 2 and distinguished from sequencing noise. As a second method, vsearch was also used to cluster the sequences of 90% OTUs. Taxonomic identification of Deblur sub-operational taxonomic units (sOTUs) and 90% OTUs was performed using the trained Greengenes 13_8 99% OTU classifier and the classify-sklearn method in Qiime 2. The resulting feature table was used for downstream machine learning analysis.
[0173] PICRUSt2 was also run on the 90% OTU sequences. First, the 90% OTU table was filtered to a minimum frequency of 10 sequences and a minimum prevalence of 10 samples. Then PICRUSt2 was run on the filtered table, where the output data types were KEGG orthologs, enzymes, and pathways. These data tables were used for downstream machine learning analysis.
[0174] Genome coverage analysis
[0175] The genome coverage of each genome in Rep206 was calculated. Briefly, for each microbial genome, the total number of uniquely covered bases in the genome was calculated by aggregating the alignments of all plasma samples in the dataset. The same process was applied to all sequencing blank samples. Then the proportion of covered bases in the genome to the total bases was calculated for each Rep206 genome in plasma samples and blank samples.
[0176] Stacked machine learning methods
[0177] Different ML architectures can provide different performances on the same feature set. The observation is appropriate when attempting to cascade multi-omics, multi-species (host and microbe) data together, especially when cascading categorical data (e.g., smoker status) and quantitative data (e.g., metagenomic abundances) of the same sample. Generally, class logistic regression classifiers perform better on categorical data and protein abundances, but perform poorly on metagenomic data. This suggests that a large amount of optimization of the hyperparameters of a given single ML model is required to achieve a trade-off in model performance across multiple data types, or to identify different mechanisms through multiple models that can be integrated.
[0178] Inspired by recent multi-omics models (PMID: 34875674) and modern convolutional neural networks (PMID: 31631918) used to predict breast cancer therapy response, stacked ensemble modeling was implemented to integrate predictions from multiple models, each of which would ideally 'pick up' different aspects or distributions in the data. Combinations of two or more of the following models were explored: random forest, gradient boosting machine, logistic regression, elastic net, support vector machine (linear, radial, and radial Σ). Internal testing (data not shown) revealed that the best performance was usually achieved using the following three tandem models: random forest, elastic net, and gradient boosting machine. Therefore, these three basic algorithms were used for ensemble modeling, followed by a meta-learner containing logistic regression to weight the probability outputs of the basic algorithms while calculating the final prediction based on the weights. To ensure that the basic algorithms were consistent and provided predictions for the same held-out folds simultaneously, a standardized 10-fold cross-validation index was predefined and input into the ensemble modeling. The caretEnsemble (https: / / github.com / zachmayer / caretEnsemble) R package was used to perform model tuning, CV-fold synchronization, hyperparameter tuning of the basic algorithms, and weight tuning of the meta-learner.
[0179] For the base algorithms, the number of hyperparameter combinations is modified using the "tuneLength" parameter of caretEnsemble in the caretList() function and is set to equal 20. Since this value represents the maximum number of iterations each hyperparameter in the model attempts to obtain, and the model may have several hyperparameters, the total number of hyperparameters evaluated is the combined and accumulated value. For example, the random forest base algorithm has a single tunable hyperparameter,'mtry', which means that up to 20 mtry values will be tried, as follows: {2,3,4,7,10,16,25,39,60,91,140,215,328,503,770,1178,1802,2757,4219,6454}. Clearly, the gradient boosting machine has two tunable hyperparameters ('n.trees' and 'interaction.depth'), which means that up to 20*20 values for each hyperparameter will be tried, including combinations, or up to 400 values in total. The ranges of these hyperparameters are as follows: n.trees, 50 to 1000 in steps of 50; interaction.depth, 1 to 20 in steps of 1. Similarly, the elastic net model has two tunable hyperparameters ('α', 'λ'), resulting in 20*20 combinations. The possible α parameters generated automatically include: {0.1,0.147368421052632,0.194736842105263,0.242105263157895,0.289473684210526,0.336842105263158,0.384210526315789,0.431578947368421,0.478947368421053,0.526315789473684,0.573684210526316,0.621052631578947,0.668421052631579,0.71578947368421,0.763157894736842,0.810526315789474,0.857894736842105,0.905263157894737,0.952631578947368,1}.The possible λ parameters generated automatically include: {0.00662229466399123, 0.00824606200985824, 0.0102679723752198, 0.0127856492677635, 0.0159206531946638, 0.0198243509450728, 0.0246852240035688, 0.0307379689652752, 0.0382748293462366, 0.0476597059206647, 0.0593457268717406, 0.0738971260921729, 0.0920164859802709, 0.114578660090195, 0.142673013517157, 0.177656020501926, 0.221216758814626, 0.275458463170505, 0.343000075305506, 0.427102693834312}. In the case where the number of features is less than the allowed hyperparameters, those corresponding hyperparameters are excluded from the grid search. Unless the feature set is subset (e.g., only PET-CT SUV), the hyperparameter search for the basic algorithm occurs step by step in up to 820 different combinations (in each model).
[0180] For the model comparison with the smallest feature set, such as only CEA or OPN or PET-CT SUV, a stacked ML architecture is used and the CV-folds are matched to ensure comparable performance values. The only difference is that for a single feature, it is not possible to include the elastic net in the stacked ML because it cannot regularize a single feature; as an alternative in these cases, the elastic net for simple logistic regression is swapped and used with the basic algorithms and meta-learners of random forest and gradient boosting machine.
[0181] To (i) ensure an approximately equal representation of samples in the sequencing runs and diagnostics within a separate CV-fold, (ii) ensure the matching of CV-folds during the basic algorithm learning, and (iii) provide reproducibility and comparability between feature sets, the per CV-fold index is calculated after subsetting but before the stacked ML and saved in the metadata for each comparison type. For example, for the comparison of lung cancer with health using metagenomic data ( Figure 7C), the following operations were performed: (a) Starting from the metadata of the entire CODICES cohort, the samples were a subset of 808 healthy or cancer - affected subjects; (b) Using the data subset, based on stratified random sampling of the sequence runs and diagnoses made in the metadata, a predefined 10 - fold CV index was calculated to ensure equal representation, and these indices were saved as new metadata columns; (c) During stacked ML, predefined CV - folds were used for the base algorithms and the meta - learner. Additionally, since these CV - folds were determined on the basis of the sample metadata, this means that all feature sets were trained and evaluated on the same CV - folds. Thus, the observed performance of different feature sets ( Figure 7 C) can be directly compared to each other, as well as to the AUROC confidence intervals calculated by using the performance with separate CV - folds.
[0182] For metagenomic data normalization, before stacked ML, the unsupervised batch - corrected metagenomic count data (see above) was transformed using centered log - ratio transformation. This was done using the corresponding subset before stacked ML, which means, for example, that it would occur on 808 healthy samples and lung cancer samples during the corresponding stacked ML ( Figure 7 C), rather than on the entire CODICES cohort (n = 1030).
[0183] Since protein concentrations have already been corrected on a plate - by - plate basis using multi - dimensional normalization (see above) (PMID: 27570895), no further transformation was performed on them. If included, the normalized CEA and OPN values were simply concatenated as features to the larger feature table before stacked ML.
[0184] For clinical metadata normalization, different methods are applied according to the data category. For numerical data (e.g., tumor size in centimeters), the available data are coerced to be numeric, and any resulting "NA" values are assigned as "0" (i.e., in logistic regression, they would 'drop out' as any weighted coefficient multiplied by 0 is 0). For categorical data, the features are factorized. For boolean data, the data are coerced to logical (i.e., binary data), except in cases where the data are incomplete, in which case a third category ("unknown") is added. Then, the normalized categorical and boolean columns are converted to dummy variable columns using the dummyVars() function of caret. For example, a single column of categorical data with three levels ("Level A", "Level B", "Level C") will be converted into a three-column data frame of binary (1 or 0) entries, where the column names are "Level A", "Level B", and "Level C". When using multiple metadata variables, such as numerical variables and categorical variables, the resulting normalized columns with dummy variables are concatenated together. Notably, whenever possible, the log odds and probability of the Mayo Clinic risk score for nodule malignancy ("pCA") were calculated using the existing clinical metadata variables in the data, as follows: {log odds = 0.0391 * age + 0.1274 * tumor_size_mm + 0.7917 * smoker_binary + 0.7838 * upper_lung_nodule_binary + 1.0407 * spiculation_binary - 6.8272}. Notably, all CODICES cohort samples had no prior history of cancer, so this was excluded from the log odds calculation; additionally, since the diagnostic model was designed to be independent but complementary to PET-CT imaging, PET-CT was optional in the pCA calculation and was also excluded from the log odds equation. Then the log odds were converted to the pCA probability, as follows: {pCA_probability = 100 * exp(log odds) / (1 + exp(log odds))}.
[0185] As an important note, the situation where missing data artificially inflates ML performance was addressed and mitigated. This issue was identified when initially incorporating'smoker status' into the low-risk clinical setting model ( Figure 7 C); although including it led to a higher AUROC between patients with lung cancer and healthy controls, an examination of the trained feature importance showed that "unknown" smokers were among the highest-ranked variables. Further examination of the metadata revealed the problem, where most patients with lung cancer knew their smoking status (394 / 401 had valid entries), but healthy subjects did not (115 / 408 had valid entries). Thus, the missing smoking status (transformed to "unknown" during metadata normalization) led to an artificial performance improvement. This was for the healthy vs. lung cancer comparison ( Figure 7C) and comparison with initial lung disease and lung cancer Figure 8 G), Part of the reason for using only metagenomic and / or proteomic data, which can be used for 100% (1030 / 1030) or 99.22% (1022 / 1030) of the total CODICES cohort respectively. In addition, only a subset of the CODICES cohort (n = 454) was used to train and evaluate subsequent clinical proteomicrobiome models using clinical metadata, where most samples contained this information (e.g., tumor size was available for 431 / 454 samples).
[0186] To calculate the ROC curve, predictions for the held-out folds (using the adjusted stacked ML model) were concatenated to produce a single set of predictions with the same number of original samples. For example, data from 808 samples were input into the stacked ML for the comparison of lung cancer with healthy controls Figure 7 C), and the prediction list had 808 rows associated with it, one for each sample. Notably, this table also included information on which samples belonged to which held-out folds, enabling the calculation of the AUROC for each CV fold, which was then aggregated to estimate the 95% performance confidence interval. Additionally, the full-length concatenated prediction list was a subset of lung cancers of individual clinical stages and histological subtypes, plus their control samples (i.e., healthy controls or lung diseases), followed by the regeneration of the ROC curve (e.g., Figure 7 C, middle and right figures). This process has been done by Mathios et al. (7) in other parts. In these cases, the CV-fold information after sample subsetting was used to estimate the confidence interval of the AUROC performance in each stage and histological subtype comparison. For clarity, this means that in Figure 7 C, for a given set of features, a single stacked model was built and evaluated as a single diagnostic test.
[0187] For the 16S rRNA cohort, which is a subset of the CODICES cohort processed independently for amplicon-based sequencing, by replacing the microbial feature table with 16S rRNA taxonomy or PICRUSt2-predicted functional pathway abundances, the same stacked ML architecture as above was used. In addition, since CEA, OPN, and clinical metadata were available for overlapping patients in the CODICES cohort, the use of amplicon-derived abundances for clinical proteomicrobiome evaluation was allowed Figure 10 D).
[0188] Another way to integrate the predictions of multiple models is'multimodal' learning, where multiple classifiers are developed independently on different features of the same samples, and a meta-learner is used to integrate the predictions from these classifiers. This is different from the stacked ensemble learning used, where a single feature table is evaluated by multiple classifiers and then by a meta-learner. Although such'multimodal' learning methods have been evaluated (data not shown), they have not been found to perform better than stacked ensemble models, and they have also been found to be more complex to optimize.
[0189] CODICES group: Shotgun metagenomic alignment, decontamination, and normalization
[0190] Fastp was used to perform quality filtering on the reads to remove adapter sequences and low-quality reads. Reads aligned against the human genome were separated from non-human reads by aligning them against the human reference genome GRCh38 using Bowtie2 with a sensitive parameter set, followed by alignment using SNAP against GRCh38. Unaligned reads against the human genome were removed using vsearch. Non-human reads were aligned against the RefSeq database version 206 ("rep206") using Bowtie2 with a sensitive parameter set. The abundances of each microbe were tallied from the alignments and used for downstream analysis.
[0191] For in silico decontamination, decontam was applied in prevalence mode to all 98 extraction blanks, which were processed in parallel with the CODICES group plasma samples, using default parameters. After viewing the histogram of prevalence p-statistics calculated by decontam ( Figure 11 A), the default cut-off value of P* = 0.1 was retained. Note that in prevalence mode, the p-statistic is set equal to the p-value of a chi-square test or Fisher's exact test, although it is not directly considered the p-value for a specific taxon. It was also noted that prevalence-based decontamination remains effective in environments with extremely low biomass. Overall, 5,630 taxa in the original rep206 data in the CODICES group were flagged as putative contaminants, and 6,455 taxa were retained (53.41% of the total); the 5,630 removed taxa accounted for 0.88% of the total read count.
[0192] For the intersection of the feature sets of Tsay et al. ("Zebra: Static and Dynamic Genome Cover Thresholds with Overlapping References." mSystems 2022; e0075822; hereinafter referred to as "Tsay"), the 16S rRNA results were shared by the corresponding first author and used to identify overlapping microorganisms in the rep206 feature set at the genus level. Specifically, since the shotgun metagenomics data of the CODICES group utilized other parts of the species (e.g., the "decontamination" and "WIS" feature sets), all species with a corresponding genus-level overlap with the 16S rRNA data collected by Tsay from bronchoalveolar lavage fluid samples were retained.
[0193] For the intersection of the feature sets with the catalog of cancer-related bacteria and fungi of the Weizmann Institute of Science (WIS), the first author of Narunsky-Haziza et al. ("Pan-cancer analyses reveal cancer type-specific fungal ecologies and bacteriome interactions. Cell", hereinafter referred to as "Narunsky") shared the list of bacterial and fungal "hits" for all tissue samples (tumors, NATs, or true normal humans [breast only]) in the two studies. Using the multi-region amplicon sequencing method for bacteria or ITS2 sequencing for fungi, this "hit" list was filtered to microorganisms with species-level results, which were then intersected with the rep206 data of the CODICES group, leaving 300 overlapping features.
[0194] Due to their low-biomass nature, studies of cancer-associated bacteria have previously shown batch effects that often need to be considered prior to downstream analysis. Past analyses in TCGA have shown that this is particularly necessary when comparing data across sequencing centers, sequencing platforms, and experimental strategies (WGS versus RNA-Seq). In the CODICES cohort, a sequential sequencing effect in the raw data was also observed that needed to be corrected prior to downstream analysis to ensure that results were not explained by sequential sequencing variation (although random grouping of additional sample types in the sequential sequencing mitigated this effect). Additionally, although batch effect correction using a combination of Voom and SNM was previously performed on TCGA cancer microbiome data and useful results were obtained, the Voom-SNM strategy is inherently a supervised strategy and difficult to use in a diagnostic approach. More specifically, if the goal of the test is to obtain biological information, such information cannot be used to supervise batch correction unless the supervisory information can be easily obtained in some other way (e.g., using age and / or sex as biological variables). After evaluating various methods (data not shown), it was found that ComBat-Seq, designed to be used with discrete counts based on a negative binomial model, completely removed the effect of the sequencing run on the raw rep206 data of the CODICES cohort when applied in an unsupervised manner during the sequencing run. Specifically, ComBat-Seq (in the sva package v.3.35.2 of R) was run using default parameters where a single vector indicated which sequencing run each sample belonged to (the "batch" variable) without any additional information (i.e., unsupervised correction). An additional advantage of the ComBat-Seq method is that both the input and output are discrete counts, which is different from Voom-SNM, which involves a log transformation and accompanying pseudo-counts to enable these log transformations. This means that the ComBat-Seq-corrected counts are in the same units as the original counts and can be directly used for downstream analysis that requires discrete count data input. For the CODICES cohort data subset ( Figure 7 C), prior to applying unsupervised ComBat-Seq normalization to the entire 1030-sample dataset of the feature set, taxonomic features were first trimmed (e.g., down to 300 WIS overlapping features, or 6530 Tsay overlapping features). Metagenomic bin abundances were treated the same, with unsupervised ComBat-Seq performed on the entire CODICES cohort, also using default parameters and a single "batch" vector indicating the sequencing run for each sample.
[0195] Then, all downstream analyses use a single, unsupervised, batch-corrected, normalized dataset, one for each feature set. For analyses subsequently using subsets of the CODICES cohort (e.g., lung cancer samples versus healthy controls), the unsupervised batch-corrected normalized dataset is a subset of the samples of interest prior to additional analysis. This means that a single version of the data, once normalized, is used for all downstream analyses for each feature set; in other words, batch-corrected subsetting avoids creating a large number of similar but slightly different datasets for each subset.
[0196] CODICES Cohort: Plasma Proteome Data Generation and Normalization
[0197] The Bio-Plex 200 platform (Bio-Rad Laboratories, Hercules, CA) was used to assess the levels of target proteins in human plasma samples. A total of 99.22% (1022 / 1030) of the samples in the CODICES cohort had sufficient plasma to obtain protein concentrations. Briefly, plasma samples were centrifuged, diluted 1:2, and subjected to a Milliplex bead-based immunoassay (Millipore Sigma, Burlington, MA) according to the manufacturer's protocol. The Milliplex HCCBP1MAG-58K (Millipore Sigma, Burlington, MA) panel was used to detect carcinoembryonic antigen (CEA) and osteopontin (OPN). The concentration of each protein analyte was determined using 5-parameter logistic curve fitting available in Bio-Plex Manager 6.2 software (Bio-Rad Laboratories, Hercules, CA) and protein standards provided with the Milliplex assay.
[0198] Samples with protein concentrations outside the range values (“OOR<” or “OOR>”) were assigned concentrations that were 10% higher (for “OOR>”) or lower (for “OOR<”) than the upper or lower limits provided by the protein standards, respectively. For CEA, the upper limit of the standard was 18,556 pg / mL and the lower limit of the standard was 25.45 pg / mL; for OPN, the upper limit of the standard was 400,000 pg / mL and the lower limit of the standard was 548.7 pg / mL. A total of 32 plates were used to process the CODICES cohort samples for protein analytes. In the CODICES cohort samples, plate-specific effects were removed in an unsupervised manner and using default settings using the accompanying R package MDimNormn (v.0.8.0), as described by Hong et al. (43). After these steps, any sample with missing values was assigned a concentration of “0” to avoid sample loss in downstream machine learning analyses, which cannot utilize missing data.
[0199] Sensitivity and specificity analysis and ROC comparison
[0200] Using the method described by Chabon et al. ("Integrating genomic features for non-invasive early lung cancer detection", Nature. 2020; 580: 245-51), sensitivity at fixed specificity and specificity at fixed sensitivity were analyzed using 1000 bootstrap resamplings. When plotting, the median and interquartile range across bootstrap resamplings were plotted. It should be noted that two other diagnostic classifiers in the field have reported their performance as 'Specificity at 97% sensitivity' or 'Sensitivity at 80% specificity', which is why the model tests were also evaluated at these cut-off values. Using the pROC (v.1.18.0) R package, ROC comparisons were made using Delong's method. Since the predictions were made on the same samples using models with different feature sets, which means the predictions were correlated, McNemar's test was used to compare sensitivity and specificity on non-bootstrapped data.
[0201] TCGA analysis
[0202] Bioinformatics alignment of 15,512 strictly host, base quality, and length-filtered (>45bp) TCGA samples with metagenomic bins showed that 14,809 samples (95.47%) had an alignment ≥1 with 1545 bins. Normal adjacent tissues (NATs) that were not part of the metagenomic assembly accounted for 73.26% (515 / 703) of the dropped samples. Metadata quality control was then performed on the remaining samples; specifically, (i) 38 samples taken from Johns Hopkins University consisted only of acute myeloid leukemia, as it was not possible to distinguish between possible sequencing center batch effects and biological effects, and (ii) samples with missing sequencing center information were removed (n = 22). The remaining 14,749 samples were then a subset of those sequenced on an Illumina HiSeq machine, which accounted for 98.00% of the samples (14454 / 14479), leaving 14454 samples for downstream analysis. The metagenomic bin abundances of TCGA were then analyzed using the method described in detail by Narunsky.
[0203] Briefly, when applying batch correction, Voom is used to transform discrete counts into pseudo-normalized data, and then SNM is used to remove batch effects in a supervised manner. The only supervision information used is "sample type" (e.g., "blood-derived normal", "primary tumor", "solid tissue normal"), and when applicable, the batch correction factor includes the sequencing center ("data_submission_center_label") and the experimental strategy ("experimental_strategy").
[0204] The raw count data was separately tested after (i) subsetting to individual sequencing centers and experimental strategies or (ii) subsetting only to WGS data, as it has lower technical batch effects than disease type ( Figure 12 D).
[0205] Machine learning was performed using a gradient boosting model (GBM) with 10-fold cross-validation with ten independent stratifications of 10% retainers. ROC and PR curves and areas were calculated for each independent 10% holdout test set in order to obtain ten sets of two-class discrimination performance for each model - effectively ten sets of 90% training - 10% testing. The hyperparameters of the GBM were fixed: {n.trees = 150, interaction.depth = 3, shrinkage = 0.1, n.minobsinnode = 1}. Then, these performance estimates for the ten folds of each model were aggregated to estimate the 95% confidence interval of performance. In cases of class imbalance, upsampling of the minority class was used; however, a requirement of ≥20 samples per class was imposed for any two-class ML comparison. Two-class comparisons involved (i) one cancer type versus primary tumors and blood assays of all other types, and (ii) unpaired primary tumors versus adjacent normal tissue. When using Voom-SNM normalized data, features were centered and scaled prior to ML; when using raw count data, only zero-variance features were removed prior to ML construction. Otherwise, no explicit feature selection or data transformation was performed. For multi-class ML using batch-corrected data ( Figure 9 I and Figure 13 D) or raw WGS data ( Figure 13 E), an extreme gradient boosting (XGBoost) machine was used, 10-fold cross-validation was performed, zero-variance features were removed, and the following fixed hyperparameters were used: {nrounds = 10, max_depth = 4, η =.1, γ = 0, colsample_bytree =.7, min_child_weight = 1, subsample =.8}.
[0206] The negative control machine learning analysis was run in the same way as above, but (i) permute the metadata of the prediction labels, or (ii) dynamically shuffle the sample IDs in the count data just before ML model building. The differences between global permutation and shuffling (i.e., once before all ML model building and testing) and dynamic permutation and shuffling (i.e., just before ML model building, but after data subsetting and labeling) were previously tested, and it was found that dynamic permutation and shuffling produced more consistent results (smaller variance) and showed more consistency with known null values (i.e., positive class prevalence of 50% AUROC and AUPR). Therefore, dynamic permutation and shuffling were used as negative controls when comparing performance to actual samples. Since the same random number seed was used for ML in both the actual analysis and the negative control analysis, they were evaluated in the same cross-validation folds. Downstream statistical analysis compared the performance of the actual analysis relative to the negative control analysis and found that the former's performance was statistically better.
[0207] Briefly, for alpha and beta diversity analysis, Qiime 2 (v.2021.11) and the corresponding plugins on a subset of samples that included individual sequencing centers, WGS, and sequencing platform (Illumina HiSeq) were used. Alpha diversity was calculated using the non-phylogenetic core metrics function of Qiime 2, rarefied to 15,000 reads / sample (approximately the 1st quartile of the sample read distribution in primary tumors and blood samples). Beta diversity analysis was performed using RPCA (robust Aitchison distance) of DEICODE, which is designed not to require rarefied data, and then the adonis implementation of Qiime 2 was used to calculate the accompanying PERMANOVA statistics.
[0208] Briefly, for differential abundance testing, ANCOM-BC was iteratively applied within the WGS sequencing center subset to evaluate a one-versus-all comparison between cancer types using the metagenomic bin abundances in primary tumors ( Figure 22 ) or blood ( Figure 25 A-25E). The following parameters were used: {p_adjust_method = "BH", zero_cutoff = 0.999, lib_cutoff = 1000, tol = 1e-5, max_iter = 100, save = false, alpha = 0.05, global = false}. Each category should have at least 10 samples before calculating abundance taxa by different methods; otherwise, the comparison will be skipped. Statistical discrimination was performed for each cancer type against all other cancer types within each subset. Then the calculated beta values, p values, and BH-adjusted q-values were used as the values for plotting volcano plots.
[0209] Hopkins group analysis
[0210] Bioinformatics alignment of 537 stringent host, base quality, and length-filtered (>45bp) samples from Cristiano et al. (“Genome-wide cell-free DNA fragmentation in patients with Cancer”. Nature. 2019;570:385-9; hereinafter referred to as “Hopkins” or “Cristiano”) to metagenomic bins gave ≥1 alignment for all 537 samples to 770 bins. As previously done, only untreated samples were used for downstream analysis, and if a patient had ≥1 sample, the earliest timepoint sample was selected, leaving a total of 491 samples (91.43% of the total) for downstream analysis. Metagenomic bin abundances in the Hopkins cohort were then analyzed using the methods described above and the methods described in detail by Narunsky.
[0211] Briefly, for machine learning comparing individual cancers to healthy or one cancer type to all other types, the following fixed hyperparameter set (identical to TCGA) was used to apply GBM to the raw metagenomic bin abundances using 10-fold cross-validation: {n.trees = 150, interaction.depth = 3, shrinkage = 0.1, n.minobsinnode = 1}. Zero-variance features were removed prior to ML, and upsampling occurred in the case of class imbalance. Performance on separate independent holdout folds was used to estimate 95% confidence intervals for AUROC and AUPR.
[0212] Briefly, for machine learning comparisons between grouped cancer samples and healthy controls ( Figure 9M), the same machine learning architecture was adopted, but 10-fold cross-validation with 10 repetitions was used to exactly match the method used by Cristiano et al. In such cases, repetitions rather than individual folds were used to estimate the confidence intervals. For plotting, the scikit-learn method for R (https: / / scikit-learn.org / stable / auto_examples / model_selection / plot_roc_crossval.html) was adjusted to estimate the average AUROC and AUPR curves in 10 repeated iterations. This can be a challenging task because the specificity cut-off points of the ten model iterations are not always equal to each other and interpolation is required. Specifically, to obtain the average performance line, linear interpolation was performed using the approx() base R function for each ROC and PR curve across 1000 equally spaced points between 0 and 1, while ensuring that each average curve starts and ends at the corners of the plot. Then, 1000 interpolated y-values between x = 0 and x = 1 were used to calculate the average ROC and PR curves and their 95% confidence intervals at each point. Superimposing these average performance lines with the 95% confidence interval bands showed good agreement.
[0213] For the negative control ML analysis, the metadata of the samples was dynamically scrambled, or their count data was dynamically shuffled, as done in TCGA, using the same random number seed to ensure matching cross-validation fold comparisons with the non-scrambled / shuffled analysis. Then, a statistical comparison of the actual and control analysis performance was performed, showing significantly better performance for the actual samples ( Figure 9 K).
[0214] For beta diversity, the DEICODE plugin of Qiime 2 was used, applying the RPCA (robust Aitchison distance) of DEICODE (the same as in the TCGA and CODICES cohorts), and subsequently the adonis implementation of Qiime 2 was used to calculate the PERMANOVA statistics.
[0215] Hong Kong Hepatocellular Carcinoma (HCC) Cohort: Data Registration and Processing
[0216] Plasma sequence data from liver cancer patients and healthy individuals were downloaded from the European Genome–Phenome Archive (EGA) accession number EGAS00001001024.42. Reads were quality-filtered by fastp, depleted of the GRCh38 host, and then aligned against metagenomic bins using Salmon. Statistical differences in alpha diversity were calculated by Wilcoxon test and beta diversity metrics by PERMANOVA based on alpha diversity, beta diversity51 of metagenomic bins, Jaccard, and RPCA, and differential bin abundances by ANCOM-BC44. Nucleotide frequency (NTF) data were generated by Budhraja et al. (PMID: 36630480), and their combination with bin abundances was used for downstream machine learning analysis.
[0217] Fragment end motif nonanucleotide frequency (NTF) analysis
[0218] Nucleotide frequency information around the fragment ends was calculated using quality-filtered sequencing reads before removing human reads. Nucleotide frequencies at nine relative positions around the fragment ends of the reads were calculated, which were previously identified as being the most informative for cancer determination (PMID: 36630480). This resulted in a table with the frequency of each base at each of the nine relative positions for each sample. This information was used for downstream machine learning analysis together with microbial, proteomic, and / or clinical data.
[0219] Lung–gut cohort: Data registration and processing
[0220] 16S rRNA sequence data from the feces of patients with LUAD and healthy individuals were downloaded from the European Nucleotide Archive (ENA) accession numbers PRJEB44169 (LUAD samples) and PRJEB33905 (healthy samples, PMID: 34586729). Reads were quality-filtered by fastp, depleted of the GRCh38 host (human reads were not expected in 16S data for consistency), and then aligned against metagenomic bins using Salmon. Alpha diversity, beta diversity were calculated by Jaccard and RPCA (PMID: 30801021) based on metagenomic bin abundances, and differential bin abundances were calculated by ANCOM-BC (PMID: 32665548). Statistical differences in alpha diversity were calculated by Wilcoxon test, and beta diversity metrics were calculated by PERMANOVA.
[0221] Blinded validation cohort analysis
[0222] The blinded validation cohort consisted of 108 samples with plasma aliquots ≤ 1 mL and paired clinical metadata without labels or diagnoses. Plasma was processed using the same methods described above for the CODICES cohort to obtain shotgun metagenomic data in independent sequencing runs. Small aliquots of plasma were also processed for proteomic information (CEA, OPN). Clinical metadata normalization was the same as for the CODICES cohort (described above), including the calculation of the Mayo Clinic risk score pCA. After reviewing the metadata, two samples were discarded based on metadata quality, specifically because of missing age ("0") or unverifiable age (year ending in "00", meaning 22 or 122 years old, or an outlier youngest or oldest in the cohort), which would affect pCA calculation. This left 106 blinded samples for evaluation of the clinical proteo-metagenomic test.
[0223] Prior to prediction on the blinded cohort, possible sequencing run variation and protein plate variation were accounted for using an unsupervised normalization method. Specifically, for the metagenomic data, the following steps were taken: (i) the raw metagenomic bin abundances of 454 samples with available clinical proteo-metagenomic information in the CODICES cohort were row-combined with the raw data in the blinded validation cohort; (ii) unsupervised ComBat-Seq was run on the combined table, where the "batch" vector contained the sequencing runs of the 454 CODICES samples and the independent validation run, without using other information and with default settings; (iii) then the metagenomic bins of the 454 CODICES samples with available clinical proteo-metagenomic information that were normalized were retained for model training; (iv) the normalized metagenomic bin data of the blinded validation cohort samples were retained for model prediction; and (v) an independent centered log-ratio transformation was performed on the retained normalized metagenomic bin data of the blinded validation cohort samples.
[0224] For proteomic data, the following steps were taken: (i) Blind validation cohort samples with protein concentrations outside the range values (“OOR<” or “OOR>”) were assigned concentrations that were 10% higher (for “OOR>”) or 10% lower (for “OOR<”) than the upper or lower limits provided by the protein standards, respectively (as done for the CODICES cohort); (ii) Protein concentrations from the blind validation cohort were then combined with the raw protein concentrations row from 454 CODICES samples; (iii) Multidimensional normalization was then used to remove plate-specific effects from the combined protein data in an unsupervised manner; (iv) Any samples in the CODICES cohort with missing values were assigned a “0” concentration to avoid sample loss (Note: All blind validation cohort samples had available protein concentrations); (v) The normalized protein concentrations of the 454 CODICES samples were retained for model training; (vi) The normalized protein concentrations of the blind validation cohort samples were retained for model prediction.
[0225] For the clinical metadata of the blind validation cohort, the normalization of the data was the same as that of the CODICES cohort. For any instances of clinical variables found in the CODICES cohort but not in the blind validation cohort (e.g., tumor_consistency = “part solid and ground glass” was only an instance in the CODICES cohort), empty zero-value dummy variable columns were added to the blind validation cohort clinical metadata. This was highly necessary because (a) stacked ML methods use dummy variables for categorical and boolean features, and (b) the stacked ML tuning model requires the presence of the same features as those initially used for training for prediction.
[0226] With the normalization of the metagenomic and proteomic data of the CODICES cohort (n = 454 samples), the final model was tuned using the above 820-step hyperparameter grid search, using these data concatenated with the clinical metadata. Specifically, the following features were included: metagenomic bins, CEA, OPN, pCA probability, block probability, smoker status, emphysema status, tumor size (cm), spiculation, tumor consistency, upper lobe nodule. As previously done, the metagenomic data were immediately subjected to centered log-ratio transformation before stacked ML. Additionally, equivalent stacked ML models were built for the same samples using only PET-CT SUV, pCA probability, or block probability.
[0227] The finally trained stacked ML models of the CODICES cohorts were then applied to the normalized, blinded validation cohort's clinical protein metagenomic data to obtain probability predictions for each sample. This process was then repeated for the stacked ML models of the CODICES cohorts constructed using only PET-CT SUV, pCA probability, or Brock probability. The probability predictions were then sent to the providers of the blinded cohort, and the ROC curves and corresponding AUROCs were returned.
[0228] CODICES, 16S rRNA, TCGA, Hopkins, and validation cohorts: Statistical analysis
[0229] Downstream analysis and plotting were generated using R version 4.03 or 4.1.1. Commonly used R packages included phyloseq (v.1.38.0), vegan (v.2.5-7), doMC (1.3.7), dplyr (v.1.0.7), reshape2 (v.1.4.4), ggpubr (0.4.0), ggsci (v.2.9), rstatix (v.0.7.0), tibble (v.3.1.6), caret (v.6.0-90), caretEnsemble (v.2.0.1), gbm (v.2.1.8), xgboost (v.1.5.0.1), randomForest (v.4.6-14), glmnet (v.4.1-3), MLmetrics (v.1.1.1), PRROC (v.1.3.1), pROC (v.1.18.0), e1071 (v.1.7-9), gmodels (v.2.18.1), ANCOM-BC (v.1.4.0), decontam (v.1.14.0), limma (v.3.50.0), edgeR (v.3.36.0), snm (v.1.42.0), sva (3.35.2), biomformat (v.1.22.0), and Rhdf5lib (v.1.16.0). The rstatix package corrected for multiple hypothesis testing when applicable. No sample size was pre-estimated, and no power calculations were performed. The gbm, randomForest, and glmnet packages were used for two-class ML; the xgboost package was used for multi-class gradient boosting ML. Except for the validation analysis using pROC, all other analyses used the PRROC package to calculate AUROC and AUPR. It was noted that the R programming language has two numerical limits when calculating small numbers including p-values: (1) double eps, or the smallest positive floating-point number x such that 1 + x!= 1, i.e., 2.220446×10 -16 ; (ii) double x min, or the smallest non-zero normalized floating point number, i.e., 2.225074×10 -308 (although this limit may be lower depending on the computational environment). Some R packages, especially ggpubr, do not report p-values less than double eps and thus represent them in the data as p < 2.2×10 -16 ; in contrast, other R packages, especially rstatix (listed below), report p-values as low as double x min , and in the data, p-values less than double x min are reported as p < 2.2×10 -308 . They are not the range of p-values.
[0230] Plasma microbiome discovery cohort
[0231] To more comprehensively capture the distribution of the plasma microbiome and evaluate its diagnostic utility, an age- and sex-matched CODICES cohort (Comprehensive Oncological Group Analysis for Early Cancer Diagnosis Identification) was constructed, consisting of 1030 untreated plasma samples, covering 11 lung cancer pathological subtypes (38.93%), 11 different etiologies of lung diseases (21.55%), and healthy controls (39.51%). Among lung cancers, nearly two-thirds were clinical stage I (n = 131, 32.59%) and stage II (n = 110, 27.36%), with smaller subsets of stage III (n = 98, 24.38%) and stage IV (n = 57, 14.18%). The smoking status (i.e., current, former, never) was known for 708 (68.73%) subjects, enabling a sub-analysis among smokers carrying the most common lung cancer risk factor. Populations of different ethnicities were also captured, with 15.63% (n = 161) of the samples originating from self-defined Black or African American, American Indian, Asian, or Pacific Islander ethnicities. Additionally, the lesion sizes of lung cancers and lung diseases were available for 428 (41.55%) subjects, with a median diameter of 2.5 cm (mean ± SD = 2.96 ± 1.91 cm). To complement the CODICES discovery cohort, a blinded validation cohort consisting of 106 patients with clinical stage I cancer or lung disease was separately obtained.
[0232] Then, 400 μL of plasma was isolated from each subject in the CODICES cohort, followed by extraction of cell-free DNA and shallow shotgun metagenomic sequencing (≈20 million reads / sample). Due to the highly fragmented nature of cell-free DNA, approximately one-third of the samples (n = 335) were additionally subjected to 16S rRNA amplicon sequencing using the shorter V6 region on a single sequencing run as validation of the method. To control for contamination, each sequencing run of shotgun metagenomic and amplicon data contained positive (mock community) and negative (experimental blank) controls, consistent with other low-biomass microbial sequencing protocols. The sample types were also randomized at each sequencing run to prevent batch confounding. Additionally, since multiple host-centric diagnostic methods have described the utility of protein-based markers in lung cancer, plasma-derived concentrations of carcinoembryonic antigen (CEA) and osteopontin (OPN) were also measured when volume permitted, covering 99.22% (n = 1022) of the CODICES cohort.
[0233] The availability of matched metagenomic, proteomic, and clinical metadata facilitated the development of multi-species (host and microbial) stacked machine learning (ML) strategies ( Figure 7 A). Specifically, after filtering out low-variance features and applying centered log-ratio transformation to the microbial counts, the concatenated data were fed into an ensemble classifier that consisted of elastic net, random forest, and gradient boosting algorithms acting in parallel with matched cross-validation (CV) folds. The scores from the three algorithms were then weighted using a logistic regression ensemble model. A ten-fold CV scheme was employed to optimize the model hyperparameters in an 820-step grid search. The predictions for the independent held-out cross-validation folds were saved to estimate the modeling performance in the discovery cohort, and a final clinical proteo-metagenomic model adjustment was performed on 454 CODICES samples with all three sets of information available before applying it to a blinded validation cohort of 106 patients ( Figure 7 A).
[0234] Evaluating known metagenomic signatures for lung cancer diagnosis in a low-risk clinical setting
[0235] Once the sequencing data and ML architecture were ready, the diagnostic ability of known metagenomic signatures was evaluated in a low-risk clinical setting across all stages for 407 healthy individuals and 401 never-treated, age- and sex-matched cancer-bearing individuals ( Figure 7B). Specifically, all 9.57×10⁸ cell-free DNA reads that did not map to the human reference genome (1.82% of the total) were aligned against the RefSeq release 206 database (“rep206”) containing bacterial, fungal, and viral genomes, of which a total of 9.60×10⁷ reads were aligned against microorganisms (0.18% of the total). Then, 98 parallel processes and the total extraction blanks included in each sequencing run were used to infer putative contaminants. Popularity-based filtering using these blanks identified and removed 5,630 species (46.59% of the taxa, with the remaining n = 6,455 taxa), which were mainly low-abundance taxa, accounting for only 0.88% of the total microbial reads ( Figure 11 A). Because the rep206 genomes are well-defined, the aggregated genomic coverage was calculated using Zebra for each sample in the CODICES cohort, across each microorganism in the database Figure 11 B). Then, a new subset of microbial taxa with a total genomic coverage of ≥1% (n = 2,172, 17.97% of the taxa) was created to compare downstream ML performance. Repeating these steps using the microorganisms identified in the experimental blanks revealed that many well-covered microorganisms in plasma had low coverage in the blanks Figure 11 C).
[0236] Next, the microorganisms identified by rep206 were cross-referenced with those previously found in lung cancer-related bronchoalveolar lavage (BAL) samples (n = 6,530 taxa) and the bacterial and fungal species identified in two highly decontaminated pan-cancer tissue center studies at the Weizmann Institute of Science (WIS) (n = 300 taxa). Notably, when calculating the Fisher's exact test for enrichment in the set with a total genomic coverage of ≥1%, there was a strong overlap in the BAL sample characteristics (p = 3.37×10⁻¹⁴⁷, X² = 667.57) and the WIS characteristics (p = 2.44×10⁻⁵⁹, X² = 263.89; Figure 11 B) showed a strong overlap, although they were of BAL and tissue origin, respectively, indicating that the plasma-derived microbial signatures can compositionally reflect their intra- and even peritissue counterparts.
[0237] Then, the diagnostic performance of these four microbial signature sets using a stacked ML strategy ( Figure 7 A and 7C) was compared. Notably, although these signature sets differed 21.8-fold in size between the largest and smallest (6,530 vs. 300 taxa), they all provided strong and similar AUROCs (range: 90.1 - 92.8%) with overlapping 95% confidence intervals Figure 7C, upper left). Subset prediction for different histological subtypes and stages further showed consistent strong diagnostic performance (minimum AUROC: 87.3%, maximum: 99.0%), with the strongest performance for small cell lung cancer (SCLC, mean AUROC = 97.03%; Figure 7 C, upper right). Unexpectedly, when examining the prediction for all cancer subtypes in individual stages, the best diagnostic performance was observed when differentiating clinical stage I lung cancer and non-cancer controls (stage I mean AUROC = 92.93%; Figure 7 C, lower left), followed by a slight but progressive performance decline at higher stages (stage IV mean AUROC = 88.4%; Figure 7 C, lower left). This trend may reflect the smaller number of samples available at higher stages in the CODICES cohort ( Figure 7 B), or a biological phenomenon where advanced cancers may be sufficient to disrupt the vascular barrier, allowing non-cancer-related microbial DNA to enter the circulation. Nevertheless, even in the untreated context, these analyses validated and extended the original observation that cell-free microbial DNA (cf-mbDNA) provides a novel method for cancer diagnosis, 8.6 times more than previously examined samples.
[0238] Since the WIS-overlapping bacterial and fungal feature sets provided similar diagnostic performance ( Figure 7 C), with a significant reduction in taxa and presumptive applicability to other cancer types, 300 species were used for additional analysis. Recognizing that WIS-overlapping bacteria previously showed a strong correlation between smokers and non-smokers, this raised questions about whether differences in smoking history potentially drove the observed diagnostic performance, as 82.79% (332 / 401) of lung cancer samples were known to have smoking exposure. Therefore, repeated stacked ML was performed after subsetting healthy subjects and CODICES subjects with cancer with subjects with known smoking history. Surprisingly, the diagnostic performance of only smokers was found to be higher than before (all AUROC = 93.0%; Figure 26 A), indicating that smoking-related pathology potentially allows additional leakage of microbial DNA into the circulation, generating a stronger diagnostic signal in the cancer context.
[0239] Classical metagenomic analysis of plasma-derived microbiome
[0240] Similarly, the presence of damaged tissues and accompanying vascular barriers in patients with lung cancer led to the hypothesis that the α-diversity in plasma from cases would be higher than in controls, even within the same subject, where lung tumor tissues had lower α-diversity compared to adjacent normal tissues. In fact, after rarefying the WIS overlapping markers, significantly higher α-diversity was observed in lung cancer plasma samples compared to healthy controls ( Figure 7 D). Calculation of robust Aitchison β-diversity distances confirmed the strength of separation identified in stacked ML, with a pseudo-F PERMANOVA statistical value of 115.92 (p = 0.001; Figure 7 E). Modeling of microbial differential abundances using ANCOM-BC subsequently demonstrated more lung cancer than health-related biomarkers, with two WIS-overlapping bacteria (i.e., Pseudomonas oleovorans and Pseudomonas mendocina) particularly driving the strong separation ( Figure 7 F, red circled dots). Notably, these two bacteria had higher relative abundances in lung cancer samples compared to healthy samples or blanks ( Figure 7 G, top), and significantly greater total genomic coverage in plasma samples compared to blanks, where Pseudomonas oleovorans showed 55.24% coverage in plasma but only 0.42% coverage in blanks. Thus, classical metagenomic analysis of cancer-related microbial biomarkers in plasma confirmed the diagnostic ability of cf-mbDNA in large cohorts of untreated individuals.
[0241] Next, it was explored whether CEA and OPN proteomic markers could enhance the diagnostic performance of metagenomics. After validating their stage-wise concentration increases, and while OPN provided better performance than CEA in early comparisons ( Figure 7 H), it was found that the combined proteogenomic classifier increased cancer discrimination to that initially observed when using all putative decontaminated rep206 features (AUROC = 92.1%; Figure 7 C, upper left; Figure 7 I)—the size of the feature set increased by 21.5-fold. Additionally, although using the same stacked ML method ( Figure 26 B), individual protein features provided only moderate cancer detection, but their performance improved at higher stages, offsetting the opposite trend of the microbial classifier. This improvement was also observable when evaluating sensitivity at 99% specificity ( Figure 7 J), which itself demonstrated state-of-the-art sensitivity for detecting stage I lung cancer (≈50% at 99% specificity) compared to existing methods evaluating epigenetic or combined ctDNA and proteomic marker panels.
[0242] To further verify that this diagnostic performance was not the result of misalignment of metagenomic reads, the short V6 region of 16S rRNA was then amplified in a subset of 335 plasma samples processed in a single sequencing run that had 16 extraction blanks and four positive control mock communities, including samples from 142 lung cancer patients and 96 healthy subjects. After processing the taxonomic composition and decontamination of the sequenced amplicons with Deblur and PICRUSt2 identification, their abundances were then input into the stacked ML pipeline. Notably, the amplicon-based taxonomic composition and functional pathways strongly distinguished lung cancer from healthy controls (AUROC range: 83.4 - 86.3%; Figure 7 K) and provided diagnostic performance similar to shotgun genomics when combined with proteomic information (AUROC: 93.6%). While simultaneously confirming the validity of the shotgun metagenomic findings, this additionally demonstrated (i) the existence and utility of plasma-derived amplicon signatures for cancer diagnosis, and (ii) the applicability of functional microbial biomarkers for cancer diagnosis. The protein amplicon-based strategy would further be highly cost-effective, utilize minimal plasma volume, and provide strong diagnostic performance.
[0243] Distinguish lung cancer from risk-matched diseased controls
[0244] Nonetheless, perhaps aside from multiple cancer tests, lung cancer screening is currently rarely performed in low-risk settings. Therefore, the diagnostic performance of the same lung cancer (LC) patients was next evaluated relative to 222 adults with lung diseases of different etiologies (LD) matched for age, gender, and smoking history ( Figure 8 A), as suggested by other studies. First, it was verified that plasma-derived CEA and OPN were again significantly increased stepwise and stage-specifically in LC relative to LD ( Figure 8 B), followed by metagenomic analysis of the classical WIS-overlapping bacterial and fungal signatures ( Figure 8 C-8F). Notably, the separation of LC and LD samples was weaker, although each comparison was still statistically significant compared to the corresponding healthy samples for LC: more similar alpha diversity ( Figure 8 C), weaker pseudo-F PERMANOVA statistics with robust Aitchison distance (F = 16.80, p = 0.001; Figure 8 D) and fewer differentially abundant features ( Figure 8 E). This variation was also evident in lung cancer-related Pseudomonas oleovorans and Pseudomonas mendocina, with the relative abundances of LD increasing in the direction of LC samples compared to healthy controls ( Figure 8 E-8F).
[0245] Similarly, the stacked ML of the LC relative to the LD of WIS-overlapping microbes provided worse discrimination than its LC-matched healthy counterparts, even when CEA and OPN were added to the analysis (AUROC = 76.2%; Figure 8 G). Additionally, when the predictions were subsetted by histological subtype and stage, stage IV detection was the only comparison with similar performance of LC relative to healthy samples ( Figure 8 G, lower right; Figure 7 C, lower right). Repeating the amplicon-based 16S rRNA approach between the same 142 LC samples and the new 97 LD samples replicated the reduced performance (overall AUROC = 70.8%; Figure 26 C). These analyses demonstrate the limitations of using known cancer-associated microbial biomarkers, particularly in the context of risk-matched controls, which are rarely used in large cancer or cancer microbiome studies. They also suggest that protein amplicon-based strategies may not provide sufficient diagnostic performance unless additional features can be added to the test.
[0246] Since the WIS-overlapping bacteria and fungi were identified outside the context of risk-matched controls and further did not examine plasma samples, which are known to contain metagenomic features that perform poorly in existing microbial databases, de novo metagenomic assembly was then performed using all blood and primary tumor whole-genome sequencing (WGS) samples from The Cancer Genome Atlas (TCGA) as well as samples from the CODICES cohort ( Figure 8 H). Specifically, blood samples from all 25 available TCGA cancer types (n = 1937 samples) were matched to cancer types in the plasma-derived CODICES cohort, while LD, healthy, and blank samples were separated into different groups and subsequently combined with metagenomic co-assemblies using metaSPAdes. At the same time, metagenomic co-assembly was also performed on 24 TCGA primary tumor types (n = 2106 samples). Overall, these analyses identified 27,819 contigs with lengths > 1500 bp, which were then pooled and binned using a deep variational autoencoder. Quality control was performed on these 4605 metagenomic bins using CheckM and then merged at the predicted strain level using PhyloPhlAn 3 and a large database of human-associated species-level genomic bins (SGBs), resulting in 1562 final pan-cancer, pan-tissue, and blood metagenomic bins ( Figure 8 H).
[0247] Notably, in every comparison except for stage IV disease, the metagenomic bins outperformed the WIS-overlapping cancer-associated features in discriminating LC samples from LD samples ( Figure 8G, orange line relative to the blue line). Combining the metagenomic bins with CEA and OPN further synergistically increased the diagnostic performance ( Figure 8 G, red line). Reapplying the de novo genomes paired with CEA and OPN for low-risk analysis revealed similar performance, especially in early disease (stage I AUROC: 90.2%); Figure 27 A), although the performance gain was most evident in the high-risk setting. In addition to better AUROC values, the shape of the bin-based ROC curves in clinical stage I and II diseases ( Figure 8 G, lower left and middle lower) also indicated higher specificity at high fixed sensitivity cutoffs (e.g., 97% sensitivity), suggesting that metagenomic bins may be particularly useful for excluding malignant diseases while reducing false positives.
[0248] Evaluation of de novo assembled metagenomic libraries in TCGA
[0249] After creating a novel set of metagenomic assembled bins with superior diagnostic performance for identifying LD in the CODICES cohort, their cancer type specificity and generalizability were then validated in two independent cohorts containing a total of 16,049 samples across 35 conditions (34 cancer types and healthy), including samples independent of metagenomic assembly.
[0250] First, all 15,512 base-quality-filtered, dual-host-depleted, length-limited (>45bp), whole-genome sequencing (WGS) and RNA sequencing (RNA-Seq) samples in TCGA were realigned relative to the metagenomic library. A total of 14,809 samples (95.47%) had non-zero alignments relative to 1,545 bins, where normal adjacent tissue (NAT) was not part of the metagenomic assembly, accounting for 73.26% (515 / 703) of the dropped samples. Notably, compared to RefSeq, the median non-human read pair alignment rate across all cancers, sample types, and experimental strategies increased 891-fold relative to the bins (Wilcoxon signed rank: p ≤ 2.23×10 -308 ; Figure 32 B). To confirm that the higher alignment rate was not caused by mapping more contaminants and since TCGA did not include sequencing blanks, non-human reads from sequencing 98 reagent-only blanks - collected during sequencing runs in all CODICES and NYU cohorts - were mapped to RefSeq and the bins, and the bins were found to significantly reduce contaminant mapping (Wilcoxon signed rank: p = 8.45×10 -18 ; Figure 32 C). Crucially, the bins in TCGA samples provided a significantly higher non-human mapping rate than the blank samples (Wilcoxon: p = 6.63×10-64 ) For 99.8% of TCGA samples, the bin mapping rate exceeded the median of the blank (0.26%). These observations were consistent when TCGA samples were stratified by WGS or RNA-Seq. Thus, compared with RefSeq, although the genomic features used were reduced by 7.65-fold, the bins significantly improved the biological microbial mapping rate while reducing reagent-based contaminant mapping.
[0251] The bin alignment rate of the experimental strategy revealed a significant increase (Wilcoxon signed rank: p ≤ 2.23×10-308; Figure 32 A), although RNA-Seq samples were excluded from metagenomic assembly; in addition, the bins coordinated the WGS (95% CI: [8.25, 8.90]%) and RNA-Seq (95% CI: [8.42, 8.61]%) alignment rates, which were different in RefSeq (WGS, 95% CI: [1.66, 1.88]%; RNA-Seq, 95% CI: [0.32, 0.39]%). Calculating the average fold change for each cancer type between the bins and RefSeq revealed that the non-human mapping rate increased by an average of 1429-fold ( Figure 32 D), with lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC) increasing by 1372-fold and 2180-fold, respectively.
[0252] Quality control of the metadata (e.g., removing sequencing centers with a single cancer type), and subsetting to a single sequencing platform that accounted for 98.00% of the samples, leaving 14,454 tissue and blood samples (93.18% of the total) for downstream analysis.
[0253] Although batch effects still persisted in the TCGA metagenomic bin abundances, they were mitigated ( Figure 12 A - B). For example, although the original principal variance component analysis (PVCA) of the raw Kraken-derived genus abundances in TCGA revealed that approximately 34% and 36% of the data variance was attributed to the sequencing center and experimental strategy, respectively, in the metagenomic bin data, they were reduced to 20.3% and 3.2% of the variance, respectively. In addition, when analyzing the raw WGS and RNA-Seq samples separately, the disease type-related variance in the WGS samples (samples used for metagenomic assembly) exceeded the variance from other technical variables without any batch correction ( Figure 12 C and D), indicating that metagenomic assembly may be able to adequately mitigate technical batch effects in retrospective analysis of cancer microbiome datasets.
[0254] Nonetheless, the initial batch correction was performed to utilize the complete TCGA cohort, using gradient boosting to calculate 10-fold cross-validated ML in a cancer-versus-all-other-cancers fashion to test cancer type specificity ( Figure 13 A-C). This analysis revealed pan-cancer discrimination between 32 primary tumor types (AUROC mean = 84.31%, 95% CI: [83.28,85.34]%); moderate discrimination between tumors and NATs (AUROC mean = 73.14%, 95% CI: [70.56,75.71]%), likely due to NAT sample loss; and very strong performance in discriminating among blood samples from 20 cancer types (AUROC mean = 96.21%, 95% CI: [95.62,96.80]%). Importantly, using metagenomic bins, blood samples from lung adenocarcinoma had the highest combined AUROC and AUPR performance in TCGA (AUROC mean = 97.92; AUPR mean = 89.99%), which was not present in bacteria- or fungi-centric analyses, indicating that co-assembly with samples in the CODICES cohort improved the accompanying cancer type specificity.
[0255] As a negative control, all ML analyses were then repeated, using permuted metadata or shuffled count data on the matching CV folds, and then comparing the expected null values with the actual performance, finding cancer type discrimination that was statistically significantly better than the null model in all cases ( Figure 14 A-B and Figure 15 A). Additionally, when all TCGA blood samples from patients with early-stage (Ia-IIc) cancer were subsetted and then ML was repeated, strong pan-cancer discrimination was found (AUROC mean = 95.23%, 95% CI: [94.05,96.42]%), including specificity for lung adenocarcinoma (AUROC mean = 98.31; AUPR mean = 94.90e%; Figure 15 B). Similar performance in early- and late-stage cancers suggests that it may be difficult to discriminate between stage I and stage IV tumors using metagenomic bins, and indeed this was the case ( Figure 15 C).
[0256] Next, to completely avoid batch correction, all raw metagenomic bin abundances for each sequencing center, experimental strategy (WGS or RNA-Seq), and sequencing platform (Illumina HiSeq) were subsetted before evaluating 10-fold ML. This was done in samples from Harvard Medical School (HMS) ( Figure 9 A, AUROC mean = 98.27%, AUPR mean = 89.56%) and all other sequencing centers ( Figure 16A-16F) provided strong tumor type discrimination. Repeated permutation of the metadata and shuffled count data against the original data again revealed pan-cancer tumor type discrimination that was significantly better than null in each subset ( Figure 9 B; Figure 17 A-17F). Application of classical metagenomic analysis further revealed significant cancer type-specific alpha and beta diversity distributions in each detectable primary tumor subset, often showing that the cancer type explained more than half of the data variance in metagenomic bin abundances ( Figure 9 C; Figure 20 A-20E; and Figure 21 A-21E). To further confirm the cancer type specificity of the bins, application of differential abundance modeling using ANCOM-BC was applied to all original data subsets, again finding differentially abundant bins in each comparison ( Figure 9 D and 9H; Figure 22 representative examples in A-22G).
[0257] Repeated tumor-versus-normal ML analysis using the original metagenomic bin subsets confirmed 9 out of 12 tumor types, although the aggregated comparison was still significantly better than the null model ( Figure 18 A-18D).
[0258] Then repeated all blood-based cancer type analyses using one versus all other ML ( Figure 9 E; Figure 19 A, 19C, 19E and 19G), negative control ML analysis ( Figure 9 F; and Figure 19 B, 19D, 19F and 19H), alpha diversity ( Figure 23 A-23E), beta diversity ( Figure 24 A-24E) and differential abundance modeling ( Figure 25 A-25E). In all of these analyses, metagenomic bin-based cancer type differences in TCGA blood samples were significant, except for alpha diversity in a single sequencing center subset (Baylor College of Medicine, Figure 23 C). Additionally, two sequencing centers (HMS, Broad Institute) with blood samples from lung adenocarcinoma sources consistently showed it to have the strongest diagnostic performance among each other cancer type ( Figure 9 E; and Figure 19 G), again indicating that CODICES group co-assembly improved lung cancer detection.
[0259] In addition, since a comparison against all others increases the no-information rate (NIR), multi-class ML was performed, in which all cancer types were considered simultaneously via the gradient boosting algorithm. Applying this to all 32 primary tumor types in the batch-corrected data replicated the performance against all other types (mean pairwise AUROC = 83.19%; Figure 13 D), with a mean balanced accuracy of 65.17% (p < 2.2×10−308) compared to 10.48% NIR. In 24 cancer types, blood-based multi-class ML (with batch-corrected data) performed better (mean pairwise AUROC = 95.50%; Figure 9 I), with a mean balanced accuracy of 79.26% (p < 2.2×10−308) compared to 8.92% NIR. Since all TCGA blood samples were WGS, and since the original WGS samples had fewer sequencing center variants than the disease types using PVCA, multi-class blood ML using the original metagenomic bin abundances had nearly identical results (mean pairwise AUROC = 95.14%; mean balanced accuracy = 79.02%; Figure 13 E) was also tested. Overall, the metagenomic bins were cancer type-specific.
[0260] Evaluating de novo assembled metagenomic bins of a cancer-associated microbial signature database of microbial reads from non-blood sources: fecal microbiome data
[0261] It has been shown that when examined against TCGA tissue and blood samples, the metagenomic bins are cancer type-specific, and then the bins are analyzed to determine if they can be used as a database of cancer-associated microbial genomes against which non-blood source sequencing reads can be aligned. To this end, it was hypothesized that the bins could provide diagnostic utility for colorectal cancer (CRC). Then, geographically distinct fecal metagenomic CRC cohorts from France (FR) and China (CN) (PMID: 25432777, 26408641) were processed and cross-compared in a subsequent meta-analysis (PMID: 30936547), providing predictions of internal cross-validation and external cohort validation performance (Figure 33A).
[0262] β-diversity revealed significant presence-absence differences between CRC-bearing and healthy subjects in the FR and CN clusters (Jaccard PERMANOVA: p = 0.001; Figure 33B), but according to reports (PMID: 25432777, 26408641), the effect sizes based on Aitchison's measures were variable and Shannon α-diversity did not change significantly. Nevertheless, applying LOOCV ML in each cluster (Figure 33C) revealed higher AUROCs (FR: 88.77%, CN: 85.64%), exceeding the results published by the original authors (FR38: 84%; CN37: 83.61%) or the meta-analysis (PMID: 30936547; FR: 85%; CN: 81%). Notably, without batch correction, the cross-application of these models revealed performance superior to LOOCV (FR vs CN: 88.73%; CN vs FR: 90.75%; Figure 33C), exceeding the international meta-analysis by up to 9% in AUC. Using internal LOOCV to lock the cut-off values yielded specificities and sensitivities of up to 82.0% and 79.2% respectively (Figure 33C). Since the FR cluster provided staging information, the predictions of LOOCV and CN cross-validation were subset to early (I-II stage) and late (III-IV stage) samples, and equally strong early performance was found (92.03 - 92.62%; Figure 33D). Calculating the bootstrap sensitivity at 92% specificity (PMID: 25432777) showed an increase and consistency of ≥15 points between LOOCV and CN cross-validation (Figure 33E, left). Although the CN cluster had lower sensitivity at 92% specificity (Figure 33E, right), its LOOCV and FR cross-validation values were similar. Thus, while improving diagnosis, the cassette is applicable to different sample types (i.e., tissue, blood, and stool), independent cancer types, and geographically distinct clusters.
[0263] Microbiota changes from feces can also inform distal lung cancer diagnosis. In the absence of publicly available fecal shotgun data for lung cancer versus controls, 16S rRNA gene amplicon data from Lim et al. (PMID: 34586729) were explored to determine its compatibility with metagenomic bins (“lung - gut” cohort, Figure 33F). Despite limited 16S rRNA alignments, there were significant presence - absence (Figure 33G) and Aitchison - based beta - diversity differences, with a significantly reduced Shannon alpha - diversity, which has not been reported previously. Then LOOCV ML, matching the method of the CRC cohort, found an AUROC of 86.52% (Figure 33H), or an AUC 10% higher than that reported (PMID: 34586729). Notably, at 92% specificity, the sensitivity of 51.6% was close to the CRC result (Figure 33I) and equal to the LOOCV result of the CN cohort (Figure 33E, right). Unfortunately, healthy saliva samples from the same cohort were not publicly available for comparison (PMID: 34586729). However, metagenomic bins can be widely generalized to improve the diagnostic performance of multiple cancers, and the data here demonstrate that aligning sequencing reads from fecal samples (shotgun metagenomic reads and 16S - targeted amplicon sequencing) with the pan - cancer metagenomic bins of the present invention enables the development of colorectal and lung cancer - specific diagnostic classifiers.
[0264] Complementary nature of metagenomic bins and fragmentomics information
[0265] Human cell - free DNA (cfDNA) is usually detected together with metagenomic cell - free DNA, but their diagnostic compatibility in cancer detection remains unknown. Therefore, two cfDNA cohorts independent of metagenomic assembly (PMID: 25646427; 31142840, Figure 29 A) were explored, and on this basis, nine DNA fragment end - nucleotide frequencies (NTF: PMID: 36630480) were obtained to synergize with metagenomic bins.
[0266] The alpha - diversity and beta - diversity of patients with HCC were significantly different ( Figure 29 B). Subsequent LOOCV ML, matching the methods of the CRC and lung - gut cohorts, revealed strong bin - based discrimination (AUROC = 91.74%; Figure 29 C, ‘bin’), with a synergistic increase in NTF (AUROC = 96.77%; Figure 29 C, ‘bin + NTF’), increasing the bootstrap median sensitivity at 99% specificity by approximately 15% to 65.6% ( Figure 29 D).
[0267] We then analyzed multiple cancer plasma samples from Cristiano et al. (hereinafter referred to as "Hopkins") Figure 29 E) and found that cancer samples had significant differences in terms of beta diversity and alpha diversity Figure 29 F). Ten-fold cross-validation (CV) of each cancer using bin abundances relative to healthy controls showed that adding NTF improved the performance of 6 out of 7 cancer types and the sensitivity of 5 out of 7 cancer types Figure 29 G-J). Aggregating all cancer controls with non-cancer controls revealed a significant improvement in NTF and bin abundances for the detection of multiple cancers across all stages, with an overall AUROC of 96.7% (DeLong's: p ≤ 4.11 × 10-8; Figure 29 K-N). Overall, bins were diagnostically useful in at least 8 cancer types and had extensive synergy with human fragmentomic signatures.
[0268] Assessment of de novo assembled metagenomic bins in an independent cohort
[0269] Despite these findings, the diagnostic performance of bins in the plasma dataset was also evaluated (i) independent of metagenomic assembly and (ii) using non-cancer controls. All 491 base-quality-filtered, dual-host-depleted, length-limited (>45bp), shallow whole-genome sequencing (WGS) plasma samples from Cristiano were re-aligned. Then, the raw metagenomic bin abundances of all cancer types were compared with 260 healthy controls using ML and found that each of them had performance superior to null values Figure 9 J), including lung cancer. Permuted and shuffled negative control ML analyses showed expected null results Figure 9 K). A comparison relative to all other cancer types also confirmed superior-to-null classification in 7 out of 8 cancer types, where only cholangiocarcinoma (n = 25) failed to achieve sufficient AUPR Figure 25 F).
[0270] Calculation of robust Aitchison beta diversity distances additionally revealed a clear separation between healthy and pan-cancer samples along axis 1 Figure 9 L; PERMANOVA: F = 4.43, p = 0.001). Subsequent ML comparisons between pan-cancer or grouped stages Figure 9 M) interestingly replicated the pattern initially observed in the CODICES cohort data Figure 27A), where the metagenomic bins provided roughly similar diagnostic performance in clinical I-III stage samples (AUROC range: 86.35 - 92.86%), and then showed the lowest performance in clinical IV stage samples (AUROC 95% CI: [76.16, 79.9]). The mechanism driving this pattern is unclear; it may artificially stem from the lower count of IV stage samples in Hopkins (n = 22 samples) and the CODICES cohort, or potentially biologically represent non-cancer-related microbial translocation due to overall vascular barrier defects. In any case, all these analyses strongly support the cancer type specificity and diagnostic utility of these novel metagenomic bins in several independent cohorts.
[0271] Construction and Preliminary Evaluation of a Multi-omics, Multi-species Lung Cancer Screening Test
[0272] To integrate heterogeneous multi-omics, multi-species data into a single test, a stacked ML strategy was designed ( Figure 7 A). The aim was to design a blood test that, if positive, would trigger subsequent confirmatory LDCT or PET-CT imaging ( Figure 10 A, Figure “2” below).
[0273] Stacked ML learning was applied to all LC and healthy samples using bin abundances with or without human information (i.e., proteins and / or NTFs) (total AUROC = 94.1%; Figure 30 A). The combined feature set significantly improved performance compared to any one alone (DeLong's test: q ≤ 1.74×10-9). In stage IV, the bins and NTFs alone showed lower performance, and 100 ML iterations with stratified random sampling excluded sample number bias; however, repeating this process with the combined feature set provided the expected trend. Bootstrap sensitivity at 99% specificity demonstrated a median sensitivity approaching 60% for stage I detection ( Figure 30 B).
[0274] To ensure that the cross-validation results were not affected by overfitting, these analyses were repeated while transforming the bin abundances into three principal component vectors using robust Aitchison PCA-projection (RPCA; PMID: 30801021) and combining protein and NTF features (14 features in total), and similar performance was found (AUROC = 92.3%; Figure 30 C). Additionally, 142 LC subjects and 96 healthy subjects were processed in parallel for plasma-derived 16S rRNA amplicons (targeted microbial amplicon sequencing), independently replicating the results (AUROC = 93.6%; Figure 30 D).
[0275] Since lung cancer screening by definition requires asymptomatic subjects with a defined smoking history and age, all screened eligible LC patients (n = 95) and age-matched healthy smokers (n = 46) (cases: 67.64 ± 6.94 years; controls: 66.37 ± 6.58 years) were subset into an internal validation cohort. After retraining the stacked ML model on all other LC-bearing subjects and healthy subjects and obtaining the cut-off value associated with 99% specificity, the retrained stacked ML model was applied to the held-out internal validation cohort ( Figure 30 E). Although the AUROC of the combined model decreased, the specificity and sensitivity were similar (specificity: 96.2 - 100%, sensitivity: 54.7 - 56.8%); the model of bins only performed the best (AUROC = 91.52%). Subsetting these predictions to stage I only, screened eligible LC patients while applying the same cut-off revealed a consistent 68.2% sensitivity at the observed specificity of 96.2 - 100%( Figure 30 F). Overall, these data demonstrate that the combination of metagenomic (human, NTF, and microbial, bins) data and plasma proteins can generate a diagnostic classifier capable of differentiating subjects with lung cancer from healthy subjects.
[0276] Development and validation of a clinical proteo-metagenomic diagnostic assay
[0277] Metagenomic bins significantly improved the discrimination of LC from LD compared to WIS overlapping biomarkers, but did not provide state-of-the-art diagnostic performance alone or in combination with proteins( Figure 8 G). The proteomic classifier was then improved by adding clinically collected and radiography-related metadata. Notably, the application of this diagnosis will occur following low-dose computed tomography (LDCT) in a high-risk clinical setting while maximizing test sensitivity to potentially obviate the need for biopsies of presumed non-malignant nodules( Figure 10 A, top); in contrast, the previously described low-risk diagnosis relying only on metagenomic or proteo-metagenomic markers will be applied to a relatively healthy population while maximizing specificity( Figure 10 A, bottom), as others have done to reduce false positives. They represent two different approaches, both of which can be enhanced by metagenomic information.
[0278] Over 400 LC and LD samples in the CODICES cohort had matched clinical metadata( Figure 10B), including most lesion diameters, shapes (such as spiculation), solidity, locations (e.g., upper lung), and nodule clinical risk scores (such as Brock probability). Notably, the median lesion diameter in this cohort was only 2.5 cm (mean ± SD = 2.95 ± 1.91 cm, n = 431). Integrating these clinical variables with metagenomic bin abundances and CEA and OPN concentrations, via stacked ML on all 454 samples, significantly increased the discrimination of malignant nodules (AUROC = 90.0%), making it greater than equivalent models constructed with matched clinical risk scores (Brock, pCA) or PET-CT-derived standardized uptake values (SUV) (AUROC range: 68.9 - 75.0%; Figure 10 C, left). Optionally, adding PET-CT SUV to this clinical proteo-metagenomic diagnosis synergistically increased the diagnostic performance (AUROC = 91.8%). Notably, subsetting these predictions to lesions with known diameters ≤ 3 cm and low clinical risk scores for malignancy (pCA ≤ 50%) improved the performance of the diagnostic model, but not that of equivalent models constructed with Brock, pCA, or PET-CT SUV ( Figure 10 C, right).
[0279] The improvement in proteo-metagenomic diagnostic performance by adding clinical metadata suggests that protein amplicon-based strategies may be similarly beneficial. Accordingly, the original LC (n = 142) versus LD (n = 97) protein amplicon analysis was repeated using additional clinical information ( Figure 26 C), finding a synergistic improvement in performance for both taxonomic composition (AUROC = 91.8%) and functional pathway abundances (AUROC = 91.3%) that was superior to clinical risk scores and PET-CT SUV ( Figure 10 D). Notably, the shape of these ROC curves ( Figure 10 C - 10D) indicates that metagenomic features can contribute to highly sensitive tests (e.g., 97% sensitivity) while still maintaining adequate specificity.
[0280] Encouraged by these improvements and intrigued by the cancer type specificity of the metagenomic bins in the Hopkins and TCGA cohorts (including between lung cancer histotypes), it was hypothesized that a diagnostic model might be able to distinguish subtypes of lung cancer in a large enough cohort. If achievable, a blood-based discrimination between these two subtypes could valuably guide clinical management, as their histologies have different impacts on therapies (e.g., pemetrexed is effective for adenocarcinoma but not for its squamous counterpart). Since 293 of the 454 samples (64.5%) were from lung adenocarcinoma (LUAD) or lung squamous cell carcinoma (LUSC) ( Figure 10B), and thus a classifier for LUAD versus LUSC was constructed using the same clinical proteomic metagenomic features as the diagnostic model. Figure 10 E). Notably, despite the small nodule size (median = 2.9 cm, mean = 3.4 cm), the diagnostic model revealed strong discrimination (AUROC = 91.5%), while an equivalent model constructed using only PET-CT SUV showed significantly worse performance (AUROC = 68.6%), as well as pseudo-random performance when using only clinical risk scores or protein information (AUROC range: 47.9% - 59.7%). These findings suggest that plasma-centered metagenomic features may be able to guide the determination of minimally invasive histological subtypes. Comparison of the performance of the lung cancer versus lung disease classifier with and without NTF information revealed that adding NTF did not improve the classifier in this specific classification scenario. Figure 31 )
[0281] Applying the diagnostic model to the blinded validation cohort
[0282] After thoroughly exploring the CODICES discovery cohort, the final clinical proteomic metagenomic diagnostic classifier was then adjusted. This classifier used all 454 LC and LD samples across different tissue types and stages. In addition to metagenomic bins, CEA, and OPN, this classifier also included clinical risk scores (Block, pCA), lesion diameter, lesion spiculation, lesion solidity, upper lobe nodule status, emphysema status, and smoking history when applicable. Equivalent stacked models using only clinical risk scores or PET-CT SUV were also constructed for comparison with later samples.
[0283] Then a blinded validation cohort from collaborators was obtained, which included 106 plasma samples from patients with an unknown number of clinically stage I lung cancers and lung diseases of different etiologies. After extracting DNA from 400 μL of plasma for independent sequencing runs and sequencing containing standard positive (mock community) and negative blank controls, non-human reads were aligned against the metagenomic bins. The metagenomic and proteomic features were then normalized to account for differences between runs, followed by feature standardization to match the diagnostic model input, including clinical metadata when applicable. Notably, the average lesion diameter in this cohort was only 1.91 cm, with multiple lesions as small as 6 mm and 8 mm. Figure 10 F)
[0284] The final diagnostic model and an equivalent model using PET-CT information were then applied to blinded cohort data to generate predictions in a single-blind fashion. Notably, the diagnostic model was significantly superior to the gold-standard PET-CT information (AUROC: 79.1% vs 64.5%; DeLong's test: p = 0.019). Given the ROC curve shapes of the 454 samples in the CODICES cohort ( Figure 10 C) and the blinded validation cohort ( Figure 10 G), it was further hypothesized that the model testing should provide significantly better specificity at a fixed 97% sensitivity. Indeed, this was the case when using the comparative McNemar's test between the model predictions fixed at 97% sensitivity and SUV as well as Bloch and pCA (all tests p < 0.001), which additionally had a poorer AUROC (Bloch: AUROC = 0.724; pCA: AUROC = 0.738). Additionally, when fixed at 80% specificity, the sensitivity was significantly superior to PET-CT SUV (p = 0.00112) and appeared to be comparable or superior to the fragmentomics approach in stage I lung cancer. Thus, the model was validated in an independent, blinded validation cohort under the most challenging conditions of clinical stage I lung cancer versus lung disease, while providing state-of-the-art diagnostic performance.
[0285] As described in Example 1, the blood and tissue metagenomes in 17,520 samples from four independent cohorts across 34 cancer types were characterized, plus healthy controls and 11 different etiologies of lung disease. Starting from known cancer-related bacterial and fungal biomarkers, the study demonstrated that as few as 300 plasma-derived microbial biomarkers could provide state-of-the-art lung cancer detection (approx. 50% sensitivity at 99% specificity) in a low-risk screening setting of newly collected, untreated individuals across all stages and histological subtypes ( Figure 7 C). It was further found that this parsimonious set of microbial biomarkers could be synergistically complemented by orthogonal, host-centered plasma proteomic markers, CEA, and OPN. If performance exceeded parsimony, it was also found that BAL-related metagenomic signatures provided up to approximately 70% sensitivity at 99% specificity in clinical stage I disease ( Figure 7 J). Notably, since these diagnostic models were constructed independently of ctDNA, fragmentomics, and epigenetic markers, the performance presented here may represent the lower limit of metagenome-enhanced cancer diagnostic performance.
[0286] Notably, WIS- and BAL-related biomarkers derived from samples within or around the tissue were significantly enriched in microbes with a total genomic coverage ≥1% in the plasma-derived CODICES cohort (p < 1×10-58 for both, Fisher's exact test). These data suggest that a large portion of the plasma metagenome is of tumor tissue origin. The fact that the diagnostic model improved when restricted to subjects with a positive smoking history also supports the hypothesis that chronic tissue damage may increase the expression of tissue-derived genomes in plasma. This theory also has parallels in the gastrointestinal environment of colon cancer, where degradation of the intestinal vascular barrier enables bacterial translocation to the liver, suggesting colorectal metastasis.
[0287] To address the potential problem of mis-mapping of metagenomic reads of human DNA, the presence and utility of plasma-derived amplicon biomarkers ( Figure 7 K) were also evaluated. In fact, both the taxonomic composition and the functional pathways inferred from 16S rRNA data surprisingly showed diagnostic performance similar to shotgun sequencing, with equivalent performance when combined with CEA and OPN ( Figure 7 K). Moreover, when the diagnostic performance based on shotgun sequencing decreased in the LC versus LD comparison, the amplicon-based performance also decreased ( Figure 26 C), and both increased similarly in the context of additional clinical proteomic features ( Figure 10 D). Thus, plasma-derived amplicon-based methods can provide a cost-, time- and volume-effective alternative to shotgun-based metagenomic assessment, warranting further investigation in lung cancer and others.
[0288] The observed decrease in diagnostic performance between LC and age-, sex- and risk-matched LD controls ( Figure 8 G) further emphasizes the importance of including these sample types in cell-free DNA screening studies to characterize the true diagnostic performance in clinical real-world scenarios. This challenge also motivated the (co-)assembly of metagenomes in 5,187 cancer-related blood and tissue samples with whole-genome sequencing. By re-aligning the identified metagenomic bins against 17,185 host-failure samples followed by comprehensive statistical and ML analyses, their cancer type specificity, diagnostic ability, ability to mitigate batch effects in retrospectively collected data, and potential to non-invasively distinguish histological subtypes of lung cancer in small nodules were thoroughly demonstrated. The broad generalizability of this metagenomic approach to other retrospectively collected cancer genomic cohorts may be crucial for characterizing the functional repertoire of cancer-related microbes.
[0289] By developing a comprehensive diagnostic model for multi-omics and multi-species (host and microbe), it was found that the metagenomic enhancement method could provide higher sensitivity than existing clinical risk scores and PET-CT imaging intensity in early diseases. Figure 10 C). Specifically, in a blinded validation cohort where the abnormalities of clinical stage I cancer are challenging relative to various lung diseases, the validation and state-of-the-art performance of the diagnostic model Figure 10 F-10H) highlighted the broad utility of plasma genomics in future cancer diagnosis.
[0290] The study had several limitations. Although the sample sizes of the new and retrospective data were large, all plasma-derived microbial information was inherently of low abundance and difficult to sequence with high coverage. Attempts to calculate the total coverage could provide a proxy for the microbes that might be present in the sample set within the cohort, but they could not address the under-sampled microbes that emerged in the available data. Similarly, high total coverage could not distinguish contaminants from non-contaminants.
[0291] Despite using multiple methods to control for possible contamination, including a large number of extraction blanks and positive controls in each sequencing run, it was impossible to exclude all false positive results. Nevertheless, using (i) an amplicon-based method with separate processing, sequencing, and decontamination; (ii) an independently validated subset of biomarkers (WIS and BAL); (iii) orthogonally developed metagenomic bins; (iv) separate data sets, including those retained in the metagenomic collection; and (v) a blinded validation cohort for independent replication of the diagnostic performance, all of these comprehensively indicated that the degree or rate of contamination did not preclude generalizable, clinically useful conclusions.
[0292] This comprehensive analysis of plasma metagenomics for early detection of lung cancer provides a generalizable approach for developing or enhancing cancer diagnosis using metagenomic information, thereby benefiting patients worldwide.
Claims
1. A method for determining a disease of a subject, the method comprising: Receiving a biological sample, electronic medical record information, and one or more radiological images of the subject; Sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and Determining the disease of the subject as an output of the prediction model when providing the one or more nucleic acid molecule sequencing reads, the electronic medical record information, and data derived from one or more radiological images of the subject as inputs to the prediction model.
2. The method according to claim 1, wherein the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as an input to the prediction model.
3. The method according to claim 1, further comprising identifying one or more protein biomarkers from the biological sample of the subject.
4. The method according to claim 2, wherein the one or more protein biomarkers from the biological sample of the subject are provided as an input to the prediction model.
5. The method according to claim 4, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.
6. The method according to claim 1, wherein the disease comprises cancer or a non-cancerous disease.
7. The method according to claim 1, wherein the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof.
8. The method according to claim 1, wherein the one or more radiological images comprise an x-ray image, a computed tomography (CT) image, a low-dose computed tomography image, a magnetic resonance imaging (MRI) image, an ultrasound image, a positron emission tomography image, a fluoroscopy image, an angiography image, or any combination thereof.
9. The method according to claim 6, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.
10. The method according to claim 1, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.
11. The method according to claim 10, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.
12. The method according to claim 1, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.
13. The method according to claim 7, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
14. The method according to claim 6, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.
15. The method according to claim 6, wherein the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, stomach adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof.
16. The method according to claim 8, further comprising calculating one or more features of the one or more radiological images, wherein the one or more features of the one or more radiological images are provided as input to the prediction model.
17. The method according to claim 16, wherein the one or more features comprise a Brock cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof.
18. The method according to claim 1, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads.
19. The method according to claim 18, wherein the genomic database comprises a human genomic database.
20. The method according to claim 18, wherein the genomic database comprises a microbial genomic database.
21. The method according to claim 20, wherein the microbial genomic database comprises de novo metagenomic assembly.
22. The method according to claim 21, wherein the de novo metagenomic assembly is derived from a biological sample representing one or more healthy states.
23. The method according to claim 22, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.
24. The method according to claim 23, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
25. The method according to claim 22, wherein the one or more health states comprise cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state.
26. The method according to claim 20, wherein the microbial genome database comprises the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof.
27. The method according to claim 1, wherein the prediction model comprises a machine learning model.
28. The method according to claim 1, wherein the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.
29. The method according to claim 28, wherein the machine learning model comprises a machine learning classifier.
30. The method according to claim 27, wherein the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.
31. The method according to claim 1, wherein the prediction model is trained using leave-one-out cross-validation.
32. The method according to claim 6, wherein the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof.
33. The method according to claim 32, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.
34. The method according to claim 1, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads.
35. The method according to claim 34, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.
36. The method according to claim 1, wherein the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
37. The method according to claim 1, wherein the sequencing comprises shotgun metagenomic sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.
38. The method according to claim 1, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.
39. The method according to claim 38, wherein the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more characteristics.
40. The method according to claim 1, wherein the prediction model is configured to distinguish cancer from non-cancerous diseases of the subject.
41. The method according to claim 18, wherein the mapping or alignment is performed using Deblur, Bowtie2, Kraken, or any combination thereof.
42. A method comprising: receiving a biological sample, electronic medical record information, data derived from one or more radiological images, and a corresponding disease of one or more subjects; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and identifying one or more features corresponding to the disease of the one or more subjects in the one or more nucleic acid molecule sequencing reads, the electronic medical record information, and the data derived from the one or more radiological images.
43. The method according to claim 42, wherein the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as input to the prediction model.
44. The method according to claim 42, wherein identifying comprises aligning the one or more sequencing reads against a genomic database.
45. The method according to claim 42, further comprising training a prediction model with the one or more features and the corresponding disease of the nucleic acid molecule sequencing reads, the electronic medical record information, and the data derived from the one or more radiological images of the one or more subjects.
46. The method according to claim 42, wherein the disease comprises cancer or non-cancerous diseases.
47. The method according to claim 42, further comprising identifying one or more features of one or more protein biomarkers of the biological sample of the subject.
48. The method according to claim 47, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.
49. The method according to claim 42, wherein the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof.
50. The method according to claim 42, wherein the one or more radiological images comprise an x-ray image, a computed tomography (CT) image, a low-dose computed tomography image, a magnetic resonance imaging (MRI) image, an ultrasound image, a positron emission tomography image, a fluoroscopy image, an angiography image, or any combination thereof.
51. The method according to claim 46, wherein the cancer comprises a tumor mass with a diameter less than 3 cm.
52. The method according to claim 42, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.
53. The method according to claim 52, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.
54. The method according to claim 42, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.
55. The method according to claim 49, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
56. The method according to claim 46, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.
57. The method according to claim 46, wherein the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof.
58. The method according to claim 42, wherein the one or more radiological image features comprise a Brock cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof.
59. The method according to claim 42, further comprising mapping or aligning the one or more nucleic acid sequencing reads against a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads.
60. The method according to claim 59, wherein the genomic database comprises a microbial genomic database.
61. The method according to claim 60, wherein the microbial genomic database comprises de novo metagenomic assembly.
62. The method according to claim 61, wherein the de novo metagenomic assembly is derived from a biological sample representing one or more healthy states.
63. The method according to claim 62, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.
64. The method according to claim 63, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
65. The method according to claim 62, wherein one or more health states comprise cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state.
66. The method according to claim 60, wherein the microbial genome database comprises the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof.
67. The method according to claim 59, wherein the genome database comprises a human genome database.
68. The method according to claim 45, wherein the prediction model comprises a machine learning model.
69. The method according to claim 45, wherein the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.
70. The method according to claim 68, wherein the machine learning model comprises a machine learning classifier.
71. The method according to claim 68, wherein the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.
72. The method according to claim 45, wherein the prediction model is trained using leave-one-out cross-validation.
73. The method according to claim 45, wherein the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof.
74. The method according to claim 73, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.
75. The method according to claim 45, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads.
76. The method according to claim 75, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.
77. The method according to claim 45, wherein the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
78. The method according to claim 42, wherein the sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.
79. The method according to claim 42, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.
80. The method according to claim 79, wherein the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundance, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundance or any combination of such characteristics, and a plurality of sequencing reads associated with the one or more characteristics.
81. The method according to claim 45, wherein the prediction model is configured to distinguish cancer and non-cancerous diseases of the subject.
82. The method according to claim 59, wherein the mapping or alignment is performed using Deblur, Bowtie2, Kraken, or any combination thereof.
83. A computer system configured to determine a disease of a subject, the computer system comprising: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, wherein the software comprises executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) receive one or more sequencing reads, electronic medical record information, and one or more images of a biological sample of the subject; and (ii) determine the disease of the subject as an output of the prediction model when providing the one or more nucleic acid molecule sequencing reads, electronic medical record information, and data derived from one or more radiological images of the subject as inputs to the prediction model.
84. The computer system according to claim 83, wherein the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided as inputs to the prediction model.
85. The computer system according to claim 83, wherein the disease comprises cancer or a non-cancerous disease.
86. The computer system according to claim 83, wherein the biological sample comprises a tissue biopsy, a liquid biopsy, or a combination thereof.
87. The computer system according to claim 83, wherein the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject.
88. The computer system according to claim 87, wherein the one or more protein biomarkers from the biological sample of the subject are provided as inputs to the prediction model.
89. The computer system according to claim 87, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, or a combination thereof.
90. The computer system according to claim 87, wherein the prediction model is trained using one or more characteristics of the nucleic acid molecule sequencing reads, the electronic medical record information, and the data derived from the one or more radiological images of the one or more subjects and the corresponding diseases.
91. The computer system according to claim 88, wherein the executable instructions comprise identifying one or more characteristics of the one or more protein biomarkers of the biological sample of the subject.
92. The computer system according to claim 87, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.
93. The computer system according to claim 87, wherein the one or more radiological images comprise x-ray images, computed tomography (CT) images, low-dose computed tomography images, magnetic resonance imaging (MRI) images, ultrasound images, positron emission tomography images, fluoroscopy images, angiography images, or any combination thereof.
94. The computer system according to claim 83, wherein the cancer comprises a tumor mass with a diameter of less than 3 centimeters.
95. The computer system according to claim 83, wherein the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads.
96. The computer system according to claim 95, wherein the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of the one or more nucleic acid molecules.
97. The computer system according to claim 83, wherein the one or more nucleic acid molecule sequencing reads comprise sequencing reads of the following: mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.
98. The computer system according to claim 86, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
99. The method according to claim 84, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.
100. The computer system according to claim 84, wherein the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, low-grade glioma of the brain, invasive breast carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, renal chromophobe cell carcinoma, renal clear cell carcinoma of the kidney, renal papillary cell carcinoma of the kidney, hepatocellular carcinoma of the liver, diffuse large B-cell lymphoma of lymphoid neoplasm, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostatic adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof.
101. The computer system according to claim 83, wherein the one or more radiological image features comprise a Brock cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof.
102. The computer system according to claim 83, wherein the executable instructions further comprise mapping or aligning the one or more nucleic acid sequencing reads relative to a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads.
103. The computer system according to claim 102, wherein the genomic database comprises a microbial genomic database.
104. The computer system according to claim 103, wherein the microbial genomic database comprises de novo metagenomic assembly.
105. The computer system according to claim 104, wherein the de novo metagenomic assembly is derived from a biological sample representing one or more healthy states.
106. The computer system according to claim 105, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.
107. The computer system according to claim 106, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
108. The computer system according to claim 105, wherein the one or more healthy states comprise cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state.
109. The computer system according to claim 103, wherein the microbial genomic database comprises a RefSeq database, a LifeNet database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination of databases thereof.
110. The computer system according to claim 102, wherein the genomic database comprises a human genomic database.
111. The computer system according to claim 83, wherein the prediction model comprises a machine learning model.
112. The computer system according to claim 83, wherein the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.
113. The computer system according to claim 111, wherein the machine learning model comprises a machine learning classifier.
114. The computer system according to claim 111, wherein the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.
115. The computer system according to claim 83, wherein the prediction model is trained using leave-one-out cross-validation.
116. The computer system according to claim 84, wherein the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof.
117. The computer system according to claim 116, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.
118. The computer system according to claim 83, wherein the executable instructions further comprise decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads.
119. The computer system according to claim 118, wherein the decontamination comprises computational decontamination, experimental control decontamination, or a combination thereof.
120. The computer system according to claim 83, wherein the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
121. The computer system according to claim 83, wherein the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.
122. The computer system according to claim 83, wherein the executable instructions further comprise determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.
123. The computer system according to claim 122, wherein the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of such characteristics, and a plurality of sequencing reads associated with the one or more characteristics.
124. The computer system according to claim 83, wherein the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject.
125. The computer system according to claim 102, wherein the mapping or alignment is performed using Deblur, Bowtie2, Kraken, or any combination thereof.
126. A method for determining a disease in a subject, the method comprising: receiving a biological sample from the subject; sequencing one or more nucleic acid molecules of the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and When one or more nucleic acid molecule sequencing reads of the subject are provided to a prediction model, a disease of the subject is determined as an output of the prediction model, wherein the prediction model is trained with one or more nucleic acid molecule sequencing reads and corresponding diseases of one or more liquid biological samples and one or more tissue biological samples of one or more subjects.
127. The method according to claim 126, wherein the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided to the prediction model as an input.
128. The method according to claim 126, wherein the disease comprises cancer, a non-cancerous disease, or a combination thereof.
129. The method according to claim 126, further comprising identifying one or more protein biomarkers from the biological sample of the subject.
130. The method according to claim 129, wherein the one or more protein biomarkers from the biological sample of the subject are provided to the prediction model.
131. The method according to claim 129, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.
132. The method according to claim 127, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters millimeters.
133. The method according to claim 126, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.
134. The method according to claim 133, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.
135. The method according to claim 126, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.
136. The method according to claim 126, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
137. The method according to claim 127, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.
138. The method according to claim 127, wherein the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof.
139. The method according to claim 126, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genomic database to determine one or more human, non-human, or a combination thereof features of the one or more nucleic acid sequencing reads provided as input to the prediction model.
140. The method according to claim 139, wherein the genomic database comprises a microbial genomic database.
141. The method according to claim 140, wherein the microbial genomic database comprises a de novo metagenomic assembly.
142. The method according to claim 141, wherein the de novo metagenomic assembly is derived from a biological sample representing one or more healthy states.
143. The method according to claim 142, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.
144. The method according to claim 143, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
145. The method according to claim 142, wherein the one or more healthy states comprise cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state.
146. The method according to claim 140, wherein the microbial genomic database comprises a RefSeq database, a LifeNet database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination of databases thereof.
147. The method according to claim 139, wherein the genomic database comprises a human genomic database.
148. The method according to claim 126, wherein the prediction model comprises a machine learning model.
149. The method according to claim 126, wherein the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.
150. The method according to claim 148, wherein the machine learning model comprises a machine learning classifier.
151. The method according to claim 148, wherein the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.
152. The method according to claim 148, wherein the prediction model is trained using leave-one-out validation.
153. The method according to claim 148, wherein the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof.
154. The method according to claim 153, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.
155. The method according to claim 126, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads, wherein the one or more decontaminated nucleic acid molecules are provided as input to the prediction model.
156. The method according to claim 155, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.
157. The method according to claim 126, wherein the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
158. The method according to claim 126, wherein the sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.
159. The method according to claim 126, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.
160. The method according to claim 159, wherein the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more characteristics.
161. The method according to claim 126, wherein the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject.
162. The method according to claim 139, wherein the mapping or alignment is done using Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
163. A method for identifying one or more non-human genomic features, the method comprising: receiving one or more liquid biological samples, one or more tissue biological samples, and corresponding diseases of one or more subjects; sequencing one or more nucleic acid molecules of the one or more liquid biological samples and the one or more tissue biological samples, thereby generating one or more sequencing reads; and Identify one or more non-human genomic features corresponding to the disease of the one or more subjects from the one or more sequencing reads.
164. The method according to claim 163, wherein the one or more sequencing reads comprise one or more sequencing reads of microbial nucleic acid molecules, and wherein the one or more sequencing reads of microbial nucleic acid molecules are provided as input to the prediction model.
165. The method according to claim 163, wherein identifying comprises aligning or mapping the one or more sequencing reads against a genomic database to determine one or more human, non-human, or combined features of the one or more nucleic acid sequencing reads.
166. The method according to claim 164, wherein the genomic database comprises a microbial genomic database.
167. The method according to claim 166, wherein the microbial genomic database comprises a de novo metagenomic assembly.
168. The method according to claim 167, wherein the de novo metagenomic assembly is derived from a biological sample representing one or more healthy states.
169. The method according to claim 168, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.
170. The method according to claim 169, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
171. The method according to claim 168, wherein the one or more healthy states comprise cancer, a pre-cancerous state, a non-malignant disease state, or a disease-free healthy state.
172. The method according to claim 166, wherein the microbial genomic database comprises a RefSeq database, a LifeNet database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination database thereof.
173. The method according to claim 163, further comprising training a prediction model with the one or more non-human genomic features and the corresponding disease of the one or more subjects.
174. The method according to claim 163, wherein the disease comprises cancer or a non-cancerous disease.
175. The method according to claim 163, further comprising identifying one or more features of one or more protein biomarkers of the one or more liquid biological samples, the one or more tissue biological samples, or a combination thereof.
176. The method according to claim 175, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.
177. The method according to claim 174, wherein the cancer comprises a tumor mass with a diameter less than 3 cm.
178. The method according to claim 163, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.
179. The method according to claim 178, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.
180. The method according to claim 163, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.
181. The method according to claim 163, wherein the liquid biological sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
182. The method according to claim 174, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.
183. The method according to claim 174, wherein the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, low-grade glioma of the brain, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, renal chromophobe cell carcinoma, renal clear cell carcinoma of the kidney, renal papillary cell carcinoma of the kidney, liver hepatocellular carcinoma, diffuse large B-cell lymphoma of lymphoid neoplasm, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof.
184. The method according to claim 164, wherein the genomic database comprises a human genomic database.
185. The method according to claim 166, wherein the prediction model comprises a machine learning model.
186. The method according to claim 166, wherein the prediction model comprises a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.
187. The method according to claim 185, wherein the machine learning model comprises a machine learning classifier.
188. The method according to claim 185, wherein the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.
189. The method according to claim 166, wherein the prediction model is trained using leave-one-out cross-validation.
190. The method according to claim 166, wherein the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof.
191. The method according to claim 190, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.
192. The method according to claim 163, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads.
193. The method according to claim 192, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.
194. The method according to claim 166, wherein the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
195. The method according to claim 166, wherein the sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.
196. The method according to claim 163, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.
197. The method according to claim 196, wherein the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundances, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundances, or any combination of features thereof, and a plurality of sequencing reads associated with the one or more characteristics.
198. The method according to claim 166, wherein the prediction model is configured to distinguish cancer from non-cancerous diseases in the subject.
199. The method according to claim 164, wherein the mapping or alignment is performed using Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
200. A computer system configured to determine a disease in a subject, the computer system comprising: (a) one or more processors; and (b) A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including software, wherein the software includes executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) Receive one or more sequencing reads of a biological sample of a subject; and (ii) When providing the one or more nucleic acid molecule sequencing reads of the subject to a prediction model, determine a disease of the subject as an output of the prediction model, wherein the prediction model is trained with one or more nucleic acid molecule sequencing reads and corresponding diseases of one or more liquid biological samples and one or more tissue biological samples of one or more subjects.
201. The computer system according to claim 126, wherein the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and wherein the one or more microbial nucleic acid molecule sequencing reads are provided to the prediction model as an input.
202. The computer system according to claim 200, wherein the disease includes cancer or a non-cancerous disease.
203. The computer system according to claim 200, wherein the executable instructions include receiving one or more protein biomarkers from the biological sample of the subject.
204. The computer system according to claim 203, wherein the one or more protein biomarkers from the biological sample of the subject are provided to the prediction model.
205. The computer system according to claim 203, wherein the executable instructions include identifying one or more characteristics of the one or more protein biomarkers of the biological sample of the subject.
206. The computer system according to claim 203, wherein the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA 21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.
207. The computer system according to claim 201, wherein the cancer includes a tumor mass having a diameter of less than 3 centimeters.
208. The computer system according to claim 200, wherein the one or more nucleic acid molecule sequencing reads include one or more amplicon-based 16S rRNA sequencing reads.
209. The computer system according to claim 208, wherein the amplicon-based 16S rRNA sequencing reads include sequencing reads of the V6 region of the one or more nucleic acid molecules.
210. The computer system according to claim 200, wherein the one or more nucleic acid molecule sequencing reads comprise sequencing reads of any of the following: mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.
211. The computer system according to claim 200, wherein the liquid biological sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
212. The method according to claim 201, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.
213. The computer system according to claim 201, wherein the cancer comprises: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine corpus endometrial carcinoma, uveal melanoma, or any combination thereof.
214. The computer system according to claim 200, wherein the executable instructions further comprise mapping or aligning the one or more nucleic acid sequencing reads against a genomic database to determine one or more human, non-human, or combined characteristics of the one or more nucleic acid sequencing reads.
215. The computer system according to claim 214, wherein the genomic database comprises a microbial genomic database.
216. The computer system according to claim 215, wherein the microbial genomic database comprises a de novo metagenomic assembly.
217. The computer system according to claim 216, wherein the de novo metagenomic assembly is derived from a biological sample representing one or more health states.
218. The computer system according to claim 217, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.
219. The computer system according to claim 218, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.
220. The computer system according to claim 217, wherein the one or more health states include cancer, pre-cancerous state, non-malignant disease state, or disease-free health state.
221. The method according to claim 215, wherein the microbial genome database includes the RefSeq database, the LifeNet database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.
222. The computer system according to claim 214, wherein the genome database includes a human genome database.
223. The computer system according to claim 200, wherein the prediction model includes a machine learning model.
224. The computer system according to claim 200, wherein the prediction model includes a neural network, a convolutional neural network, logistic regression, a random forest, a support vector machine, or any combination thereof.
225. The computer system according to claim 223, wherein the machine learning model includes a machine learning classifier.
226. The computer system according to claim 223, wherein the machine learning model includes a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.
227. The computer system according to claim 200, wherein the prediction model is trained using leave-one-out cross-validation.
228. The computer system according to claim 200, wherein the prediction model is configured to determine the stage of the cancer, the anatomic origin of the cancer, or a combination thereof.
229. The computer system according to claim 228, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.
230. The computer system according to claim 200, wherein the executable instructions further include decontaminating the one or more nucleic acid molecule sequencing reads to produce one or more decontaminated nucleic acid molecule sequencing reads.
231. The computer system according to claim 230, wherein the decontamination includes in silico decontamination, experimental control decontamination, or a combination thereof.
232. The computer system according to claim 200, wherein the prediction model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
233. The computer system according to claim 200, wherein the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.
234. The computer system according to claim 200, wherein the executable instructions further include determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.
235. The computer system according to claim 234, wherein the one or more characteristics of the one or more nucleic acid molecules comprise non-microbial taxonomic abundance, mammalian genomic coordinates, annotated genomic loci, mammalian functional genes, and / or biochemical pathway abundance or any combination of characteristics thereof, and a plurality of sequencing reads associated with the one or more characteristics.
236. The computer system according to claim 200, wherein the prediction model is configured to distinguish cancer and non-cancerous diseases of the subject.
237. The computer system according to claim 214, wherein the mapping or alignment is performed using Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
Citation Information
Patent Citations
Microbiome identification and use in diagnosis of head and neck squamous cell cancer
US20180223338A1
Method to detect colon cancer by means of the microbiome
US20180258495A1
Compositions, methods and kits for diagnosis of lung cancer
WO2019079635A1
Methods and systems for analyzing microbiota
WO2019191649A1
Detection of lung cancer using cell-free DNA fragmentation
WO2022140386A1