Multimodal method and system for disease diagnosis

A multimodal method using metagenomic assembly and machine learning models for cancer diagnosis addresses sensitivity and specificity issues in early-stage cancer detection by leveraging tumor-centric microbial databases and protein biomarkers, achieving high predictive performance and accurate cancer detection.

JP2025536883APending Publication Date: 2025-11-12LIQUID BIOPSY HOLDCO LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2025518470
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Current methods for diagnosing cancer, particularly lung cancer, face challenges in sensitivity and specificity, especially in early-stage detection, due to limitations in detecting low molecular abundance in circulation and the complexity of managing indeterminate pulmonary nodules, with existing tests showing low compliance and high false-positive rates.

Method used

A multimodal approach utilizing metagenomic assembly to create a tumor-centric reference database from whole-genome-sequenced tumor tissue and blood-derived samples, combined with machine learning models, to analyze microbial nucleic acid sequencing reads and protein biomarkers, enhancing the detection of early-stage cancer by improving mapping rates and reducing contaminants.

Benefits of technology

The method achieves high accuracy in diagnosing early-stage cancer with a predictive performance of at least 85% accuracy, sensitivity, specificity, precision, and AUPR, and AUROC, while identifying cancer-associated microbial biomarkers, thereby improving early-stage cancer detection and reducing false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536883000001_ABST
    Figure 2025536883000001_ABST
Patent Text Reader

Abstract

As described elsewhere herein, provided herein are multimodal methods and / or systems for diagnosing one or more diseases.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 412,369, filed September 30, 2022, which is incorporated herein by reference. [Background technology]

[0002] Until recently, cancer was widely considered to be a sterile tissue. Therefore, methods for diagnosing cancer have relied on the detection of low- to high-plexity human biomarkers. However, tumor size limits the molecular abundance in circulation, such as circulating tumor DNA (ctDNA), making the sensitivity of these methods challenging for detecting early-stage disease. Targeting more biomarkers, deeper sequencing, and / or more frequent testing have been used to improve early detection, but they cannot circumvent biologically imposed constraints that may hinder clinical utility.

[0003] Prior art in the related field may include US2018 / 0223338, US2018 / 0258495, WO2019 / 191649, WO2019 / 079635, WO2022 / 140386, or WO2022 / 212283. Summary of the Invention

[0004] Lung cancer is a prime example of this challenge, remaining the leading cause of cancer-related deaths worldwide, with approximately 23% of patients diagnosed at limited stage in the United States. Lung cancer screening with low-dose computed tomography (LDCT) improves early diagnosis and reduces mortality in at-risk individuals, but its benefits are limited by low patient compliance. Minimally invasive liquid biopsy can improve screening compliance but is less sensitive in early-stage disease than LDCT, which reliably detects nodules as small as 4 millimeters in diameter. For example, in a validation cohort comparing stage I disease with primarily healthy controls, a commercially available cell-free DNA (cfDNA) methylation assay reported a sensitivity of approximately 25% with a specificity of 99.3% (PMID: 33506766), a ctDNA-based assay reported an average sensitivity of approximately 30% with a specificity of 98% (PMID: 32269342), a fragmentomics assay reported a sensitivity of approximately 50% with a specificity of 80% (PMID: 34417454), and an integrated multi-omics test reported a sensitivity of approximately 40% with a specificity of 96.3% (PMID: 29348365). Conversely, the sensitivity of LDCT ranged from 59% to 100%. These discrepancies highlight the need for different approaches to early lung cancer detection.

[0005] Another diagnostic challenge stems from the difficult management of indeterminate pulmonary nodules (IPNs), which are noncalcified masses measuring 6–30 mm in diameter and whose malignant status is unclear. Incidental pulmonary nodules are reported in approximately 30% of chest CT scans, with approximately >1.5 million nodules detected annually in the United States. Although patients with IPNs are at higher risk for lung cancer than the general population, most IPNs are benign, complicating their evaluation. Furthermore, 25.8% of transthoracic needle biopsies result in clinical complications. For this reason, the standard of care (SOC) for IPN management is careful monitoring of nodule size and / or PET-CT scans to rule out malignancy. However, PET-CT evaluation is complicated by diabetic hyperglycemia, inconsistent standardized uptake values ​​(SUVs), and false-positive results due to infectious etiologies. Liquid biopsy can aid in the diagnostic determination of IPNs, but improving PET-CT SOC would require high sensitivity for small nodule sizes (≤3 cm). To our knowledge, there is only one non-PET-CT test available for determining IPN malignancy, with an AUROC of 0.76 in the validation cohort (97% sensitivity, 44% specificity; PMID: 29496499), demonstrating an unmet need for innovation.

[0006] We have previously characterized intracellular, cancer-type-specific communities of tumor microbiomes, whose genomes are detectable in the circulation and constitute orthogonal biomarkers for human-derived molecules (PMIDs: 32214244; 32467386; 36179670). However, this study identified only a small fraction of the total metagenomic content by either enriching for microbe-specific amplicons or capturing trace DNA or RNA fragments with matches in publicly available reference databases. Even after repeated, highly sensitive host depletion and quality filtering, 98.8% of the approximately 4.4 billion non-human DNA or RNA reads in The Cancer Genome Atlas (TCGA) could not be mapped to any known organism in the RefSeq30 (Release 200) multidomain database (PMID: 36179670). This suggests that considerable cancer-associated microbial diversity remains unexplored.

[0007] Starting with the scientific problem of unmappable microbial reads, the methods and / or systems described elsewhere herein, in some embodiments, can generate a tumor-centric reference database via metagenomic assembly for 5187 whole-genome-sequenced tumor tissue and / or blood-derived cancer samples. In some cases, the metagenomic assembly can include de novo metagenomic assembly. In some cases, the tumor-centric reference data can include up to about 1562 metagenomic bins. In some cases, the tumor-centric reference data can include at least about 1562 metagenomic bins. In some embodiments, these bins increase the median mapping rate of non-human DNA in TCGA by at least about 891-fold, while simultaneously reducing the median mapping rate of reagent-based contaminants by at least about 7.6-fold compared to publicly available reference genomes. As reviewed elsewhere herein, in some embodiments, the methods and / or systems of the present invention utilize these cancer-derived metagenomic bins to create two distinct multi-omics, multi-species tests for lung cancer (e.g., early stage) detection that combine bin-derived information with other data features (e.g., plasma proteins and clinical risk scores described elsewhere herein) via predictive (e.g., machine learning) modeling. In some cases, a first model identifies lung cancer in a subject (e.g., a subject from an otherwise healthy population), and a second model can, for example, determine the malignant status of a detected lung cancer or nodule IPN in the subject. In some cases, each of the first and second models demonstrates strong predictive performance in a validation subset of stage I disease, with the model(s) having a predictive performance of at least about 85% accuracy, specificity, sensitivity, precision, AUPR, AUROC, or any combination thereof. Thus, by exploring unexplored microbial diversity, the present disclosure demonstrates and illustrates the utility of plasma-derived metagenomics for diagnosing early-stage cancer, and further demonstrates that metagenomic bins serve as a pan-cancer (i.e., multi-cancer type, not limited to lung cancer) database capable of identifying cancer-associated microbial biomarkers.

[0008] Aspects of the present disclosure provided herein describe a method for determining disease in a subject, the method comprising receiving a biological sample, electronic medical record information, and one or more radiological images from the subject; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and determining disease in the subject as an output of a predictive model when data derived from the one or more nucleic acid molecule sequencing reads, electronic medical record information, and one or more radiological images of the subject are provided as input to the predictive model. In some embodiments, the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input. In some embodiments, the method further comprises identifying one or more protein biomarkers from the subject's biological sample. In some embodiments, the predictive model is provided with the one or more protein biomarkers from the subject's biological sample. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof. In some embodiments, the one or more radiological images comprise images of X-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing.In some embodiments, amplicon-based 16S rRNA sequencing sequences the V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the method further comprises calculating one or more features of the one or more radiological images, wherein the one or more features of the one or more radiological images are provided as input to the predictive model. In some embodiments, the one or more features comprise a Brock cancer probability score, a diameter of the lesion, spiculation of the lesion, solidity of the lesion, or any combination thereof. In some embodiments, the method further comprises mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof features of the one or more nucleic acid sequencing reads. In some embodiments, the genome database comprises a microbial genome database.In some embodiments, the microbial genome database comprises a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a health condition. In some embodiments, the biological sample comprises a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the health condition comprises cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition. In some embodiments, the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a dendrogram vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the predictive model is trained using leave-one-out validation. In some embodiments, the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises decontaminating the one or more nucleic acid molecular sequencing reads to generate one or more decontaminated nucleic acid molecular sequencing reads. In some embodiments, the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun metagenomic sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and a non-cancerous disease in the subject. In some embodiments, the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.

[0009] Another aspect of the present disclosure provided herein describes a method including receiving a biological sample, electronic medical record information, data derived from one or more radiological images, and a corresponding disease from one or more subjects; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and identifying one or more features in the one or more nucleic acid molecule sequencing reads, electronic medical record information, and data derived from the one or more radiological images corresponding to the disease from the one or more subjects. In some embodiments, the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the one or more microbial nucleic acid molecule sequencing reads are provided as input to a predictive model. In some embodiments, the identifying includes aligning the one or more sequencing reads to a genome database. In some embodiments, the method further includes training the predictive model using the nucleic acid molecule sequencing reads, electronic medical record information, and data derived from the one or more radiological images and one or more features of the corresponding disease from the one or more subjects. In some embodiments, the disease includes cancer or a non-cancerous disease. In some embodiments, the method further includes identifying one or more features of one or more protein biomarkers in the subject's biological sample. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof. In some embodiments, the one or more radiological images comprise images of x-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof.In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing comprises sequencing the V6 region of the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the one or more radiological features comprise Brock's cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof. In some embodiments, the method further comprises mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genome database comprises a microbial genome database.In some embodiments, the microbial genome database comprises a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a health condition. In some embodiments, the biological sample comprises a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the health condition comprises cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition. In some embodiments, the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a dendrogram vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the predictive model is trained using leave-one-out validation. In some embodiments, the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises decontaminating one or more nucleic acid molecular sequencing reads to generate one or more decontaminated nucleic acid molecular sequencing reads. In some embodiments, the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.In some embodiments, the predictive model determines disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and a non-cancerous disease in the subject. In some embodiments, the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.

[0010] Another aspect of the present disclosure provided herein describes a computer system configured to determine a disease in a subject, the computer system including: (a) one or more processors; and (b) a non-transitory computer-readable storage medium containing software, the software including executable instructions that, when executed, cause the one or more processors of the computer system to: (i) receive one or more sequencing reads of a biological sample, electronic medical record information, and one or more images of the subject; and (ii) determine the disease in the subject as an output of a predictive model when data derived from the one or more nucleic acid molecule sequencing reads, electronic medical record information, and one or more radiological images of the subject are provided as input to the predictive model. In some embodiments, the one or more nucleic acid sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the biological sample comprises a tissue biopsy, a liquid biopsy, or a combination thereof. In some embodiments, the executable instructions include receiving one or more protein biomarkers from the subject's biological sample. In some embodiments, the predictive model is provided with one or more protein biomarkers from a subject's biological sample. In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, or a combination thereof. In some embodiments, the predictive model is trained using data derived from nucleic acid molecular sequencing reads, electronic medical record information, and one or more radiological images of one or more subjects and one or more features of the corresponding disease. In some embodiments, the executable instructions include identifying one or more features of the one or more protein biomarkers from the subject's biological sample.In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the one or more radiological images comprise images of x-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads. In some embodiments, the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise sequencing reads of mammalian RNA, mammalian DNA, DNA from mammalian cells, RNA from mammalian cells, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, DNA from non-human cells, RNA from non-human cells, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), squamous cell carcinoma of the lung (LUSC), small cell lung cancer (SCLC), or any combination thereof.In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the one or more radiological features comprise Brock's cancer probability score, lesion diameter, lesion spiculation, lesion solidity, or any combination thereof. In some embodiments, the executable instructions further comprise mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genome database comprises a microbial genome database. In some embodiments, the microbial genome database comprises a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a health condition. In some embodiments, the biological sample comprises a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the health condition comprises cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition. In some embodiments, the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model.In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a variance vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the predictive model is trained using leave-one-out validation. In some embodiments, the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the executable instructions further comprise decontaminating one or more nucleic acid molecular sequencing reads to generate one or more decontaminated nucleic acid molecular sequencing reads. In some embodiments, the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the predictive model determines disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the executable instructions further comprise determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and a non-cancerous disease in the subject.In some embodiments, the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.

[0011] Another aspect of the present disclosure provided herein describes a method for determining a disease in a subject, comprising receiving a biological sample from the subject, sequencing one or more nucleic acid molecules in the biological sample, thereby generating one or more nucleic acid molecule sequencing reads, and determining the disease in the subject as an output of a predictive model when the one or more nucleic acid molecule sequencing reads of the subject are provided to the predictive model, wherein the predictive model is trained using one or more nucleic acid molecule sequencing reads of one or more liquid biological samples and one or more tissue biological samples from one or more subjects and the corresponding disease. In some embodiments, the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the one or more microbial nucleic acid molecule sequencing reads are provided as input to the predictive model. In some embodiments, the disease comprises cancer, a non-cancerous disease, or a combination thereof. In some embodiments, the method further comprises identifying one or more protein biomarkers from the subject's biological sample. In some embodiments, the predictive model is provided with one or more protein biomarkers from the subject's biological sample. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters or millimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of one or more nucleic acid molecules.In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the method further includes mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads provided as input to the predictive model. In some embodiments, the genome database includes a microbial genome database. In some embodiments, the microbial genome database includes a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representative of a health condition. In some embodiments, the biological sample includes a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample includes plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.In some embodiments, the health condition comprises cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition. In some embodiments, the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a stochastic vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the predictive model is trained using leave-one-out validation. In some embodiments, the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further includes decontaminating one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads, and the one or more decontaminated nucleic acid molecules are provided as input to the predictive model. In some embodiments, the decontamination includes in silico decontamination, experimental control decontamination, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing includes shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.In some embodiments, the method further includes determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject. In some embodiments, the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.

[0012] Another aspect of the present disclosure provided herein describes a method for identifying one or more non-human genomic features, the method including receiving one or more liquid biological samples, one or more tissue biological samples, and a corresponding disease from one or more subjects; sequencing one or more nucleic acid molecules in the one or more liquid biological samples and the one or more tissue biological samples, thereby generating one or more sequencing reads; and identifying one or more non-human genomic features corresponding to the disease from the one or more subjects. In some embodiments, the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the one or more microbial nucleic acid molecule sequencing reads are provided as input to a predictive model. In some embodiments, the identifying includes aligning or mapping the one or more sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genome database includes a microbial genome database. In some embodiments, the microbial genome database includes a de novo metagenomic assembly. In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a health condition. In some embodiments, the biological sample comprises a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the health condition comprises cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition. In some embodiments, the microbial genome database comprises the RefSeq database, the Web of Life database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof. In some embodiments, the method further comprises training a predictive model using one or more non-human genomic features and corresponding diseases of the one or more subjects. In some embodiments, the disease comprises cancer or a non-cancerous disease.In some embodiments, the method further comprises identifying one or more characteristics of one or more protein biomarkers in one or more liquid biological samples, one or more tissue biological samples, or a combination thereof. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, amplicon-based 16S rRNA sequencing sequences the V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biological sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a stochastic vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the predictive model is trained using leave-one-out validation. In some embodiments, the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises decontaminating one or more nucleic acid molecular sequencing reads to generate one or more decontaminated nucleic acid molecular sequencing reads. In some embodiments, the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.In some embodiments, the predictive model determines disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and a non-cancerous disease in the subject. In some embodiments, the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.

[0013] Another aspect of the present disclosure provided herein describes a computer system configured to determine a disease in a subject, the computer system including: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, wherein the software includes executable instructions that, when executed, cause the one or more processors of the computer system to: (i) receive one or more sequencing reads from a biological sample of the subject; and (ii) determine a disease in the subject as an output of a predictive model when the one or more nucleic acid molecule sequencing reads from the subject are provided to the predictive model; the predictive model is trained using the one or more nucleic acid molecule sequencing reads from one or more liquid biological samples and one or more tissue biological samples of one or more subjects and the corresponding disease. In some embodiments, the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the one or more microbial nucleic acid molecule sequencing reads are provided as input to the predictive model. In some embodiments, the disease includes cancer or a non-cancerous disease. In some embodiments, the executable instructions include receiving one or more protein biomarkers from the biological sample of the subject. In some embodiments, the predictive model is provided with one or more protein biomarkers from a subject's biological sample. In some embodiments, the executable instructions include identifying one or more features of the one or more protein biomarkers from the subject's biological sample. In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter.In some embodiments, the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads. In some embodiments, the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise sequencing reads of mammalian RNA, mammalian DNA, DNA from mammalian cells, RNA from mammalian cells, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, DNA from non-human cells, RNA from non-human cells, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biological sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), squamous cell carcinoma of the lung (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the executable instructions further include mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genome database comprises a microbial genome database. In some embodiments, the microbial genome database comprises a de novo metagenomic assembly.In some embodiments, the de novo metagenomic assembly is derived from a biological sample representing a health condition. In some embodiments, the biological sample comprises a tissue sample, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the health condition comprises cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition. In some embodiments, the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a stochastic vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof. In some embodiments, the predictive model is trained using leave-one-out validation. In some embodiments, the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the executable instructions further comprise decontaminating one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads. In some embodiments, the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.In some embodiments, the predictive model determines disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof. In some embodiments, the executable instructions further comprise determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and a non-cancerous disease in the subject. In some embodiments, the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.

[0014] Aspects of the present disclosure described herein, in some embodiments, describe a method for determining disease in a subject, the method comprising: (a) receiving a biological sample, electronic medical record information, and radiology data from the subject; (b) sequencing a plurality of non-human nucleic acid molecules in the biological sample, thereby generating a plurality of microbial sequencing reads; and (c) processing the plurality of microbial sequencing reads, electronic medical record information, and radiology data with a trained predictive model, thereby determining disease in the subject with at least about 80% accuracy. The trained predictive model is trained using a plurality of microbial abundances and corresponding cancer types, and the trained predictive models include a first predictive model and a second predictive model, the first predictive model processes the plurality of microbial sequencing reads, electronic medical record information, and radiology data, and the second trained predictive model processes the output of the first predictive model. In some embodiments, the biological sample comprises a liquid biopsy. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, urine, feces, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the sequencing comprises shotgun sequencing. In some embodiments, the method further comprises receiving concentrations of one or more plasma proteins. In some embodiments, the trained predictive model comprises one or more machine learning models. In some embodiments, the method further comprises aligning sequencing reads of the biological sample to a human reference genome to identify a plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, the method further comprises aligning the plurality of non-human nucleic acid molecule sequencing reads to a database of microbial genomes to identify a plurality of microbial sequencing reads. In some embodiments, the database comprises a de novo metagenomic assembly comprising a genome contig. In some embodiments, the genome contig comprises one or more metagenomic bins. In some embodiments, aligning the plurality of non-human nucleic acid molecule sequencing reads to the de novo metagenomic assembly generates aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads.In some embodiments, the trained predictive model is configured to process aligned bin abundances of a plurality of non-human nucleic acid molecule sequencing reads of the subject. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing of a plurality of nucleic acid molecules of the biological sample. In some embodiments, the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of the V6 region of the plurality of nucleic acid molecules of the biological sample. In some embodiments, the method further comprises determining one or more features of the radiology data, wherein the one or more features of the radiology data are processed by the trained predictive model. In some embodiments, the one or more features of the radiology data comprise a Brock cancer probability score, a diameter of a cancerous lesion, spiculation of a cancerous lesion, solidity of a cancerous lesion, or any combination thereof. In some embodiments, the disease comprises cancer. In some embodiments, the cancer comprises a tumor mass having a diameter of less than about 3 centimeters or less than about 8 millimeters. In some embodiments, the trained predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, determining the disease of the subject comprises distinguishing between cancer and a non-cancerous disease in the subject.

[0015] Aspects of the present disclosure described herein, in some embodiments, describe a system configured to determine a disease in a subject, the system comprising: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, the software including executable instructions that, when executed, cause the one or more processors of the computer system to: (i) receive a plurality of non-human nucleic acid molecule sequencing reads of a biological sample of the subject, electronic medical record information, and radiology data; and (ii) process a plurality of microbial nucleic acid molecule sequencing reads of the plurality of non-human nucleic acid molecule sequencing reads of the subject, the electronic medical record information, and the radiology data with a trained predictive model, thereby determining the disease in the subject with at least about 80% accuracy, the trained predictive model being trained using a plurality of microbial abundances and corresponding cancer types, the trained predictive models including a first predictive model and a second predictive model, the first predictive model processing the plurality of microbial nucleic acid molecule sequencing reads, the electronic medical record information, and the radiology data, and the second predictive model processing the output of the first predictive model. In some embodiments, the biological sample comprises a liquid biopsy. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, urine, feces, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the plurality of microbial nucleic acid molecule sequencing reads are generated by shotgun sequencing. In some embodiments, the executable instructions comprise receiving concentrations of one or more plasma proteins. In some embodiments, the trained predictive model comprises one or more machine learning models. In some embodiments, the executable instructions cause one or more processors to align the plurality of nucleic acid molecule sequencing reads of the biological sample to a human reference genome library to identify a plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, the executable instructions cause one or more processors to align the plurality of non-human nucleic acid molecule sequencing reads to a database of microbial genomes to identify a plurality of microbial sequencing reads. In some embodiments, the database comprises a de novo metagenomic assembly comprising genome contigs.In some embodiments, the genome contig comprises one or more metagenomic bins. In some embodiments, the executable instructions cause one or more processors to align a plurality of non-human nucleic acid molecule sequencing reads to the de novo metagenomic assembly to generate aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads. In some embodiments, a trained predictive model is configured to process the aligned bin abundances of the plurality of non-human nucleic acid molecule sequencing reads of the subject. In some embodiments, the executable instructions cause one or more processors to receive a plurality of amplicon-based 16S rRNA sequencing reads of a plurality of nucleic acid molecules of the biological sample. In some embodiments, the plurality of amplicon-based 16S rRNA sequencing reads include sequencing reads of the V6 region of the plurality of nucleic acid molecules of the biological sample. In some embodiments, the executable instructions cause one or more processors to determine one or more features of the radiation data, wherein the one or more features of the radiation data are processed by the trained predictive model. In some embodiments, the one or more features of the radiological data include a Brock cancer probability score, a diameter of a cancerous lesion, spiculation of a cancerous lesion, solidity of a cancerous lesion, or any combination thereof. In some embodiments, the disease includes cancer. In some embodiments, the cancer includes a tumor mass having a diameter of up to about 3 centimeters or up to about 8 millimeters. In some embodiments, the trained predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, determining the disease in the subject includes distinguishing between cancer and a non-cancerous disease in the subject.

[0016] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0017] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 illustrates a flow diagram of data types and data stream structures for stacked machine learning training, as described in some embodiments herein. [Figure 2] FIG. 1 shows a flow diagram of metagenomic data isolated and / or identified from a human subject, as described in some embodiments herein. [Figure 3] FIG. 1 shows a flow diagram of proteomic data derived from human subjects, as described in some embodiments herein. [Figure 4] FIG. 1 shows a flow diagram of data types derived from human subject samples and human subject medical records used to train and / or develop diagnostic classifiers, as described in some embodiments herein. [Figure 5] FIG. 1 shows a flow diagram of the clinical proteometagenomic features used to generate a lung cancer classifier, as described in some embodiments herein. [Figure 6] 1 shows a diagram of a computer system configured to implement the methods of the present disclosure, as described in some embodiments herein. [Figure 7] Figures 7A-K show graphs of experimental data and metagenomic features broadly discriminating between treatment-naive lung cancer and healthy samples across diverse lung cancer histology types and stages using stacked machine learning, as described in some embodiments herein. [Figure 8] Figures 8A-H show experimental data and graphs using metagenomic plasma biomarkers to distinguish between lung cancer and lung disease and develop metagenomic bins to improve discrimination performance, as described in some embodiments herein. [Figure 9]9A-M show experimental data and graphs of the performance of a metagenomic bin-based pan-cancer classifier, as described in some embodiments herein, in diagnosing the presence and type of cancer in blood and plasma of two independent cohorts. [Figure 10] Figures 10A-I show experimental data, graphs, and / or workflow diagrams for the development and validation of a clinical proteometagenomic classifier for lung nodule malignancy determination, as described in some embodiments herein. [Figure 11] Figures 11A-C show experimental data and graphs of decontamination and aggregate genome coverage of plasma-derived microbiomes in a sample cohort, as described in some embodiments herein. [Figure 12] Figures 12A-D show graphs of experimental data and batch-corrected analysis of the microbiome of The Cancer Genome Atlas (TCGA) biological samples using metagenomic bin features, as described in some embodiments herein. [Figure 13] 13A-E show experimental data and graphs of the performance of the TCGA classifier for cancer discrimination using metagenomic bins across cancer types, as described in some embodiments herein. [Figure 14] 14A-B show experimental data and graphs from a control analysis validating the performance of the TCGA tissue-based classifier by comparing machine learning models built with scrambled metadata or shuffled samples, as described in some embodiments herein. [Figure 15] Figures 15A-C show graphs of experimental data and control analyses validating the performance of the TCGA blood-based classifier and further machine learning using blood samples from low-stage cancers or comparing primary tumors from low and high clinical stages, as described in some embodiments herein. [Figure 16]Figures 16A-F show experimental data and graphs of the performance of a TCGA classifier using subset raw data with metagenomic bin abundances for the discrimination of cancer types of primary tumors, as described in some embodiments herein. [Figure 17] 17A-F show experimental data and graphs of a raw data control analysis to validate the performance of the TCGA primary tumor classifier, as described in some embodiments herein. [Figure 18] Figures 18A-D show experimental data and graphs of the performance of the TCGA classifier using subset raw data with metagenomic bins and control analysis for discrimination of primary tumor versus adjacent normal tissue, as described in some embodiments herein. [Figure 19] Figures 19A-H show experimental data and graphs of the performance of the TCGA hematology classifier using subset, raw, metagenomic bin abundances for cancer discrimination and control analysis, as described in some embodiments herein. [Figure 20] 20A-E show experimental data and graphs of metagenomic bin alpha diversity across TCGA primary tumors, as described in some embodiments herein. [Figure 21] Figures 21A-E show experimental data and graphs of classical metagenomic analysis of cancer type specificity in TCGA primary tumor samples when using raw metagenomic bin abundances, as described in some embodiments herein. [Figure 22] 22A-G show experimental data and graphs of differential abundance of metagenomic bins between primary tumor cancer types in TCGA, as described in some embodiments herein. [Figure 23] 23A-E show experimental data and graphs of metagenomic bin alpha diversity across TCGA blood samples, as described in some embodiments herein. [Figure 24]Figures 24A-E show experimental data and graphs of classical metagenomic analysis showing cancer type specificity in TCGA blood samples when using raw metagenomic bin abundances, as described in some embodiments herein. [Figure 25] 25A-F show experimental data and graphs of differential abundance of metagenomic bins between blood samples and their associated cancer types in TCGA, as described in some embodiments herein. [Figure 26] Figures 26A-C show experimental data and graphs of the diagnostic performance of metagenomic, protein, and amplicon analyses in the Comprehensive Oncobiome Analysis for the Diagnostic Identification of Early Stage Cancer in Subjects (CODICES) cohort, as described in some embodiments herein. [Figure 27] 1 shows experimental data and graphs of the diagnostic performance of metagenomic bins in the CODICES cohort, as described in some embodiments herein. [Figure 28] FIG. 1 shows a flow diagram of data types derived from human subject samples and human subject medical records used to train and / or develop diagnostic classifiers, as described in some embodiments herein. [Figure 29] Figures 29A-N show experimental data and graphs of the diagnostic performance of metagenomic and DNA fragmentomics (nucleotide frequency) analyses derived from publicly available cell-free DNA datasets, as described in some embodiments herein. [Figure 30] Figures 30A-F show experimental data and graphs of the lung cancer vs. healthy diagnostic performance of metagenomic, plasma protein, and DNA fragmentomics (nucleotide frequency) analyses derived from the CODICES cohort, as described in some embodiments herein. [Figure 31] 1 shows experimental data and graphs of lung cancer vs. lung disease performance of metagenomic, plasma protein, and DNA fragmentomics (nucleotide frequency) analyses derived from the CODICES cohort, as described in some embodiments herein. [Figure 32]Figures 32A-D show improved whole genome and RNA sequencing mapping rates to metagenomic assembly bins compared to the publicly available RefSeq206 genome database. [Figure 33] Figures 33A-I show the performance of colon cancer and lung cancer classifiers for machine learning models generated from fecal microbiome data, where microbial abundances are derived from alignment of sequencing reads to metagenomic assemblies ("bins") of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] In some embodiments, the present disclosure describes methods for determining, identifying, and / or distinguishing malignant diseases (e.g., malignant lung tumors, alternatively referred to as tumors) from benign diseases by analyzing and / or assaying a combination of two or more data types from a patient sample. In some cases, the disclosure describes methods and / or systems for determining and / or diagnosing a disease. In some cases, the disease may include cancer. In some cases, the disease may include a non-cancerous disease. In some cases, the methods and / or systems described elsewhere herein may distinguish and / or differentiate between cancerous and non-cancerous diseases in a subject. In some cases, the cancer may include a tumor mass having a diameter of less than about 3 centimeters or less than about 8 millimeters. In some embodiments, the data types from the patient sample may include one or more analyte types from the patient sample. The one or more analyte types may include the presence and / or abundance of one or more nucleic acid molecules of non-human origin (e.g., microorganisms, bacteria, viruses, and / or fungi) in a liquid-based (e.g., blood-derived) patient sample, a non-microbial human-derived analyte (e.g., protein, human genomic nucleic acid molecule), or a combination thereof. In some cases, the liquid biopsy sample may include plasma, serum, whole blood, urine, feces, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some cases, the data type may include a patient medical history and / or diagnostic data type. In some embodiments, the present specification provides methods and / or systems configured and / or capable of detecting and / or determining the presence of early clinical stage cancer (e.g., stage I and II lung cancer) in asymptomatic subject(s) at high risk for lung cancer due to the subject's smoking history. For example, such subjects may include those who meet the U.S. Preventive Services Task Force recommendations for annual low-dose computed tomography screening (adults aged 50-80 years with a 20-pack-year smoking history who currently smoke or have quit smoking within the past 15 years).The method(s) described elsewhere herein may include training one or more predictive models (e.g., machine learning algorithms) on a dataset (e.g., feature table), where the dataset includes the aforementioned data types, thereby identifying disease (e.g., cancer) correlation patterns between these data types. Specifically, the methods and / or systems described elsewhere herein may train one or more predictive models on analyte types of microbial abundance (relative or absolute) obtained from sequencing one or more microbial nucleic acid molecules (e.g., determining microbial read counts by next-generation sequencing), alone or in combination with measured plasma protein concentrations. In some cases, the sequencing may include shotgun sequencing. In some embodiments, the training dataset may include microbial abundance, measured plasma proteins, human genomic information (e.g., DNA fragmentation patterns or chromosomal copy number changes), a cancer risk score calculated based on information obtained from a patient's medical record, or any combination thereof. In some embodiments, the methods and / or systems described elsewhere herein may train a predictive model (e.g., a multimodal model) using one or more data types to generate a trained diagnostic model (e.g., a trained multimodal diagnostic model).

[0020] In some cases, the methods and / or systems described elsewhere herein achieve a performance metric (e.g., accuracy, sensitivity, specificity, NPV, PPV, AUROC, and / or AUPR) of at least about 70% or at least about 0.7 of the predictive model when assessing, diagnosing, and / or detecting cancer in a subject from one or more data types. In some cases, the use of nucleic acid analyte types of non-human origin when training one or more predictive models described elsewhere herein may improve the performance of the trained model in distinguishing between a first lung disease state and a second lung disease state (e.g., the presence of lung cancer nodules and non-cancerous lung nodules). In some cases, non-cancerous lung nodules may result from etiologies including sarcoidosis, interstitial pulmonary fibrosis, bronchiectasis, pneumonia, chronic obstructive pulmonary disease, and / or hamartoma. Traditionally, such nodule-bearing conditions can be distinguished from true lung malignancies by invasive intrathoracic biopsy followed by histopathological examination. The methods and / or systems described elsewhere herein offer advantages over typical pathology reports in tests (e.g., intrathoracic tissue biopsies) because the disclosed methods and / or systems do not rely on observed subjective measures of tissue structure, cellular atypia, or any other subjective measures traditionally used to diagnose cancer. In some embodiments, the disclosed methods and / or systems may instead rely on measured, quantified presence and / or abundance of one or more data types. In some cases, the disclosed methods and / or systems may demonstrate increased performance metrics by narrowing down features only in microbial sources, as described elsewhere herein, rather than modified human (i.e., cancerous) sources, which are often modified at very low frequencies in a background of "normal" human sources. Furthermore, the disclosed methods and / or systems may be performed with and / or on blood-derived samples, which, compared to, for example, invasive biopsies, are minimally invasive samples and therefore pose little risk to patients and can be repeated over time at low cost.

[0021] In some embodiments, the methods and / or systems described elsewhere herein utilize, analyze, and / or train predictive models on metagenomic assemblies generated from non-human sequencing reads obtained from tumor whole-genome sequencing data. In some cases, the metagenomic assembly may include de novo metagenomic assembly. In some cases, the method may determine the microbial components of a tumor type by using the non-human components of the whole-genome sequencing data and use these components as a reference database for subsequent NGS sequencing read alignment. The metagenomic assemblies utilized by the methods and / or systems described elsewhere herein are derived, in part, from treatment-naive biopsy tumor samples in The Cancer Genome Atlas (TCGA), thereby representing diverse metagenomes from over 30 different cancer types and / or (sub)types. In some embodiments, cell-free DNA sequencing reads from a subject (e.g., patient) sample are computationally aligned to one or more tumor-derived metagenomes, as described elsewhere herein, to identify the presence and / or abundance of tumor-associated microbial features. These sample-associated microbial abundances can then be combined with other data types (e.g., plasma protein concentrations and / or clinical data features) to form a multimodal feature set, which is input into a trained diagnostic model to generate a probable cancer score. This score, as described elsewhere herein, can be used by a medical professional and / or clinician to determine whether a subsequent more invasive biopsy, e.g., a breast biopsy, may be required, or whether continuous non-invasive monitoring of the lung nodule via radiological imaging, e.g., low-dose computed tomography and / or further liquid biopsy-based analysis and / or assay methods may be required.

[0022] In some embodiments, metagenomic data may be combined with other data types to form a multimodal input dataset for training a predictive model (e.g., one or more machine learning algorithms). In some embodiments, a trained diagnostic model is generated using multimodal data derived from one or more patients and / or subjects with known health histories. For example, in the case of lung cancer, metagenomic data from a known healthy subject 101, a subject with lung cancer 102, and a subject with a non-cancerous lung disease (e.g., sarcoidosis, interstitial pulmonary fibrosis, bronchiectasis, pneumonia, chronic obstructive pulmonary disease, etc.) 103 may be combined with proteomic data (104-106) and clinical data (107-109) from the same subjects to be used as a training dataset for an ensemble (“stacked”) predictive model (e.g., a stacked machine learning algorithm) 110 that can learn from data containing both numerical and categorical data types. In some embodiments, stacked predictive model predictions from different model types (e.g., logistic regression, random forest, and support vector machine) may be combined to generate a single final predictive model 111, as shown in Figure 1. The results of training this stacked predictive model may include a diagnostic classifier, which may be used to distinguish one or more subjects with cancer from healthy subjects and / or one or more subjects with cancer from non-cancer (i.e., lung disease) subjects, based on analysis of one or more data types of subject samples not previously used in training the predictive model.

[0023] In some embodiments, metagenomic data (101-103, FIG. 1; and / or 112, FIG. 2) may be derived from one or more training and / or test subject biological samples. In some embodiments, the one or more training and / or test subject biological samples may include human 113 and non-human 115 nucleic acid molecules. In some cases, one or more sequences of the human 113 and / or non-human 115 nucleic acid molecules may be determined by sequencing. In some cases, sequencing may include next-generation sequencing. In some cases, sequencing may include shotgun sequencing. In some cases, sequencing one or more sequences of the human 113 and / or non-human 115 nucleic acid molecules may generate and / or produce one or more sequencing reads of the human 113 and / or non-human 115 nucleic acid molecules. In some cases, one or more genomic analyses and / or assays, as described elsewhere herein, may be performed and / or applied to the human 113 sequencing reads to generate a human genomic feature set for subsequent learning, analyzing, and / or training 117 predictive models. For example, human genome 114 analysis may include detection of cancer-associated DNA mutations, copy number alterations, DNA fragmentation profiles, DNA end analysis (e.g., analysis of nucleotide frequencies at the ends of DNA fragments), chromosomal instability profiles, and epigenetic profiles (e.g., DNA methylation patterns), or any combination thereof. In some embodiments, DNA fragmentation analysis (e.g., DNA end analysis) may provide a nucleotide frequency of one or more nucleotide bases. In some cases, a nucleic acid molecule sequence read length of 100 base pairs (bp) may include about 1 to about 30 bases analyzed for nucleotide frequency. In some cases, a nucleic acid molecule sequence read length of 100 bp may include about 1 to about 25 bases analyzed for nucleotide frequency. In some cases, the nucleotide frequency includes the occurrence of a given nucleotide in the analyzed base segment. In some cases, at least three nucleotide bases are analyzed for the nucleic acid molecule sequence. In some embodiments, one or more analyses and / or assays may be performed on the non-human components 115 of the metagenomic dataset to generate a feature table for learning, training, and / or analyzing a predictive model.In some cases, the non-human (e.g., microbial) genome analysis and / or assay 116 may include measuring the taxonomic abundance of microorganisms in a sample, inferring the abundance of biochemical pathways represented by identified microorganisms, aligning sequencing reads to a metagenomic assembly (“bins”) to obtain bin abundances, determining amplicon sequence variants represented in a targeted amplicon sequencing dataset, or any combination thereof. In some cases, a plurality of nucleic acid molecules of a biological sample (e.g., a liquid biopsy) may be aligned to a human reference genome to identify a plurality of non-human nucleic acid molecule sequencing reads. In some cases, the non-human nucleic acid molecule sequencing reads may be aligned to a database of microbial genomes to identify a plurality of microbial sequencing reads. In some cases, the database of microbial genomes may include a de novo metagenomic assembly including genome contigs. In some cases, the genome contigs may include one or more metagenomic bins. In some embodiments, a plurality of non-human nucleic acid molecule sequencing reads may be aligned to a de novo metagenomic assembly to generate aligned bin abundances for the plurality of non-human nucleic acid molecule sequencing reads. In some cases, as described elsewhere herein, one or more predictive models can process the aligned bin abundances of one or more subjects to determine the disease of the subject, as described elsewhere herein.In some cases, the targeted amplification sequencing dataset can include the amplification-based 16S rRNA sequencing read dataset of multiple nucleic acid molecules of biological samples.In some cases, the amplicon-based 16S rRNA sequencing read can include the sequencing read of the V6 region of multiple nucleic acid molecules of the subject's biological samples, as described elsewhere herein.

[0024] In some embodiments, proteomic data (104-106, FIG. 1; 118, FIG. 3) derived from one or more training and / or test subject biological samples may include blood-derived serum or plasma proteins 119. In some embodiments, plasma proteins 120 may include carcinoembryonic antigen (CEA), osteopontin (OPN), cancer antigen 125 (CA125), cancer antigen 19-9 (CA19-9), cancer antigen 15-3 (CA15-3), interleukin-8 (IL-8), prolactin (PRL), and cytokeratin 19 fragment (CYRA21-1). The concentrations of these plasma proteins (typically in ng / mL) can be determined via immunoassays, as detailed herein, and used as inputs 117 (along with other data types (e.g., FIG. 1)) to train a predictive model (e.g., a machine learning algorithm) and / or to analyze a previously trained predictive model.

[0025] In some embodiments, data from a human subject 121 may include proteomic data 122 in combination with non-genomic data 131, and may include various data features obtained from the subject's medical history, for example, as shown in Figure 4. In some embodiments, the data features may include the subject's smoking history 123, clinical scores used to assess cancer risk 124, clinical findings and family health history 125, and data derived from medical imaging 126. These data types may be analyzed alone (in the absence of metagenomics data (Figure 4)) to train a predictive model, or may be combined with metagenomics data features (132, 133) as shown in Figure 28, and / or may be analyzed and / or input 117 into a previously trained predictive model to determine a diagnostic output, as described elsewhere herein.

[0026] In some embodiments, a trained predictive model (e.g., a trained machine learning classifier) ​​(130, FIG. 5 ) capable of distinguishing between malignant cancer (e.g., lung cancer) nodules and benign cancer nodules (e.g., benign lung cancer nodules) can be determined and / or obtained as an output of the stacked predictive model 110 using a predictive model input feature table 117 consisting of metagenomic bin abundances 127 derived from plasma cell-free DNA sequencing data, protein concentrations of plasma proteins CEA and OPN 128, and / or clinical data obtained from a subject's medical record 129. In some cases, the clinical data can include radiology data. In some cases, the radiology data can include radiological images. In some cases, the radiological images can be generated by X-rays, computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), or any combination thereof. In some cases, one or more features of the radiological data can be determined and utilized to diagnose and / or determine disease in a subject and / or train one or more predictive models, alone or in combination with other data types, as described elsewhere herein. In some cases, one or more features of the radiology data may be analyzed, determined, and / or identified. In some cases, the one or more features of the radiology data may include a Brock cancer probability score, a diameter of a cancerous lesion, spiculation of a cancerous lesion, solidity of a cancerous lesion, or any combination thereof. In some embodiments, the clinical data features may include a Mayo lung cancer probability score, a Brock lung cancer probability score, a subject's smoking status, or any combination thereof. In some cases, the clinical data features may include features determined from medical imaging, such as lung tumor / nodule size, tumor solidity, evidence of spiculation, location of a lung nodule, presence of emphysema, or any combination thereof.

[0027] Predictive Model The disclosed methods and / or systems may utilize and / or access external capabilities of artificial intelligence, predictive models, and / or machine learning techniques to identify one or more microbial signatures of enriched (e.g., hybridization-enriched) biological samples of one or more individual subjects. In some cases, the microbial signatures determined from the subject hybridization-enriched biological samples may predict cancer and / or non-cancerous disease in one or more subjects. In some cases, the signatures may be used to train one or more predictive models (e.g., one or more machine learning algorithms) described elsewhere herein. These signatures may be used to predict disease (e.g., cancer, non-cancerous disease, disorder, or any combination thereof). Using such trained predictive models, healthcare providers (e.g., physicians) may make informed and accurate risk-based decisions, thereby improving the quality of care and monitoring provided to patients with cancer, non-cancerous disease, disorder, or any combination thereof. In some cases, the one or more predictive models may include a first predictive model configured to process and / or receive one or more data types as input, as described elsewhere herein, and a second predictive model configured to receive and / or process the output of the first predictive model, where the output of the second predictive model determines and / or diagnoses the disease in the subject. In some cases, the at least three predictive models may then provide output to a single predictive model (e.g., a logistic regression model) that may output the determination, detection, and / or diagnosis of the disease in the subject. In some examples, the at least three predictive models may be trained using one or more data types, as described elsewhere herein. In some cases, the single predictive model that receives input from the at least three predictive models may weight the outputs of the at least three predictive models with different weights for each of the models.

[0028] The methods and / or systems described elsewhere herein may analyze the presence and / or abundance of microorganisms (e.g., the abundance of microorganisms of a particular genus and / or taxon) in a biological sample enriched by hybridization probes, as described elsewhere, which may non-specifically bind to microbial nucleic acids. The presence and / or abundance of microorganisms may be used to determine one or more microbial and / or non-microbial characteristics, which may predict cancer and / or non-cancerous disease in one or more subjects. In some cases, the methods and / or systems described elsewhere herein may train a predictive model using one or more microbial and / or non-microbial characteristics indicative of cancer and / or non-cancerous disease in a subject. In some cases, the trained predictive model may then be used to generate (e.g., predict) the likelihood of cancer and / or non-cancerous disease in one or more subjects different from the one or more subjects utilized to train the predictive model. The trained predictive model may include an artificial intelligence-based model, such as a machine learning-based classifier, configured to process one or more microbial nucleic acid molecule sequence reads obtained from a hybridization-enriched biological sample to generate a likelihood that the subject has a disease or disorder. The model can be trained using the presence or abundance of hybridization-enriched biological samples from one or more cohorts of patients (e.g., cancer patients, patients with a non-cancerous disease, disease- and cancer-free patients, cancer patients undergoing treatment for cancer, patients undergoing treatment for a non-cancerous disease, or a combination thereof). In some cases, the predictive model can be trained to provide treatment predictions for treating cancer in one or more patients that are not part of the training dataset of the predictive model. Such a predictive model can output a treatment recommendation for one or more patients that are not part of the training dataset when provided with inputs of the patient presence and abundance of one or more microorganisms in the hybridization-enriched biological samples.

[0029] The predictive model may include one or more predictive models. The model may include one or more machine learning algorithms. The machine learning algorithms may include model(s) of support vector machines (SVMs), naive Bayes classification, random forests, neural networks (e.g., deep neural networks (DNNs), recurrent neural networks (RNNs), deep RNNs, long short-term memory (LSTM) recurrent neural networks (RNNs), gated recurrent units (GRUs), gradient boosting machines, random forests, or other supervised or unsupervised learning algorithms, statistics, linear regression, k-nearest neighbors, k-means, decision trees, logistic regression, or any combination thereof. The model(s) may be used for classification or regression. The model may provide an estimated output of one or more ensembles of models composed of multiple predictive models and may utilize techniques such as gradient boosting in constructing decision trees. The model may be trained using one or more training datasets including one or more microbial features, patient data such as the patient's medical history, the patient's family medical history, the patient's vitals (e.g., blood pressure, pulse, temperature, oxygen saturation), or any combination thereof.

[0030] The predictive model may include any number of machine learning algorithms. In some embodiments, the random forest machine learning algorithm may be an ensemble of bagged decision trees. The ensemble may be at least about 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 500, 1000, or more bagged decision trees. The ensemble may be up to about 1000, 500, 250, 200, 180, 160, 140, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, or fewer bagged decision trees. The ensemble can be about 1-1000, 1-500, 1-200, 1-100, or 1-10 bagged decision trees.

[0031] In some embodiments, the machine learning algorithm may have various parameters, such as a learning rate, a mini-batch size, a number of epochs to train, momentum, learning weight decay, or neural network layers.

[0032] In some embodiments, the learning rate may be between about 0.00001 and 0.1.

[0033] In some embodiments, the mini-batch size may be between about 16 and 128.

[0034] In some embodiments, the neural network can include neural network layers. The neural network can have at least about 2 to 1000, or more, neural network layers.

[0035] In some embodiments, the number of epochs for training may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 500, 1000, 10000, or more.

[0036] In some embodiments, the momentum can be at least about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or more, hi some embodiments, the momentum can be up to about 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, or less.

[0037] In some embodiments, the learning weight decay may be at least about 0.00001, 0.0001, 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, or more. In some embodiments, the learning weight decay may be up to about 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, 0.0001, 0.00001, or less.

[0038] In some embodiments, the machine learning algorithm may use a loss function, which may be, for example, a regression loss, a mean absolute error, a mean bias error, a hinge loss, an Adam optimizer, and / or a cross entropy.

[0039] In some embodiments, the parameters of the machine learning algorithm may be tuned using humans and / or computer systems.

[0040] In some embodiments, the machine learning algorithm may prioritize certain features. The machine learning algorithm may prioritize features that may be more relevant to detecting cancer, non-cancerous diseases, disorders, or any combination thereof. If a feature is classified more frequently than another feature in determining cancer, non-cancerous diseases, and / or disorders, the feature may be more relevant to detecting cancer, non-cancerous diseases, and / or disorders. In some cases, features may be prioritized using a weighting system. In some cases, features may be prioritized based on probability statistics, such as the frequency and / or quantity of the feature's occurrence. The machine learning algorithm may prioritize features with the assistance of a human and / or a computer system.

[0041] In some cases, the predictive model may prioritize certain features to reduce computational costs, save processing power, save processing time, increase reliability, and / or reduce random access memory usage.

[0042] The training dataset may be generated, for example, from one or more cohorts of patients with a diagnosis of a common cancer, non-cancerous disease, or disorder. The training dataset may include one or more microbial features in the form of microbial presence and / or abundance in enriched biological samples (e.g., hybridization-enriched biological samples) of one or more subjects. The features may include the corresponding cancer diagnosis of one or more subjects for the microbial features. In some cases, the features may include patient information such as the patient's age, patient medical history, other medical conditions, current or past medications, clinical risk score, and time since last observation. For example, a set of features collected from a given patient at a given time point may collectively act as a signature that can indicate the patient's health state or condition at the given time point.

[0043] The label can include, for example, a clinical outcome such as the presence, absence, diagnosis, and / or prognosis of cancer, a non-cancer disease, a disorder, or a combination thereof in a subject (e.g., a patient). The clinical outcome can include treatment efficacy (e.g., whether the subject is a positive or negative responder to a cancer and / or disease-based treatment).

[0044] The input features may be structured by aggregating the data into bins, or alternatively by using one-hot encoding. The input may also include feature values ​​or vectors derived from the aforementioned inputs, such as cross-correlations.

[0045] The training dataset can be constructed from one or more microbial presence and / or abundance features in a hybridization-enriched biological sample, or a combination of one or more microbial presence and / or abundance features with one or more somatic nucleic acid molecules of the enriched biological sample that are indicative of cancer, a non-cancerous disease, a disorder, or any combination thereof.

[0046] The model may process the input features to generate output values ​​including one or more classifications, one or more predictions, or a combination thereof. For example, such classifications or predictions may include a binary classification of the subject as either cancer is present or absent, a non-cancerous disease is present or absent, a disorder is present or absent, or any combination thereof. In some cases, one or more predictive models (e.g., machine learning algorithms) may classify a subject among a group of categorical labels (e.g., "no cancer, non-cancerous disease, and / or disorder," "obvious cancer, non-cancerous disease, and / or disorder," and "probability of cancer, non-cancerous disease, and / or disorder"); the likelihood (e.g., relative likelihood or probability) of developing a particular cancer, non-cancerous disease, and / or disorder; a score indicating the presence of cancer, non-cancerous disease, and / or disorder, a "risk factor" for the patient's likelihood of death, a confidence interval for any numerical prediction; or any combination thereof. Various machine learning techniques may be cascaded such that the output of a machine learning technique can also be used as input features to subsequent layers or subsections of the model.

[0047] To train a model (e.g., by determining model weights and correlations) to generate real-time classifications or predictions, the model may be trained using a training dataset and / or one or more training features described elsewhere herein. Such datasets and / or features may be large enough to generate statistically significant classifications or predictions. For example, a dataset may include a database of data including the presence and / or abundance of fungi, viruses, archaea, microorganisms, bacteria, or any combination thereof in biological samples from one or more subjects.

[0048] A dataset may be divided into subsets (e.g., separate or overlapping), such as a training dataset, a development dataset, and a test dataset. For example, a dataset may be divided into a training dataset comprising 80% of the dataset, a development dataset comprising 10% of the dataset, and a test dataset comprising 10% of the dataset. The training dataset may comprise about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The development dataset may comprise about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The test dataset may comprise about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, leave-one-out cross-validation may be used. The training set (e.g., training data set) can be selected by random sampling of a set of data corresponding to one or more patient cohorts to ensure sampling independence. Alternatively, the training set (e.g., training data set) can be selected by proportional sampling of a set of data corresponding to one or more patient cohorts to ensure sampling independence.

[0049] To improve the accuracy of model predictions and reduce overfitting of the model, the dataset can be augmented to increase the number of samples in the training set. For example, data augmentation can include rearranging the order of observations in the training records. To accommodate datasets with missing observations, methods for imputing missing data, such as forward fitting, back fitting, linear interpolation, and multitask Gaussian processes, can be used. The dataset can be filtered and / or batch corrected to remove or reduce confounding factors. For example, a subset of patients can be excluded within the database.

[0050] The model may include one or more neural networks, such as a neural network, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), and / or a deep RNN. The recurrent neural network may include units, which may be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, the model may include an algorithmic architecture including a neural network with a set of input features (e.g., microbial features, patient and / or subject vital measurements, patient and / or subject medical history, patient and / or subject demographics, or any combination thereof), as described elsewhere herein. Neural network techniques, such as dropout or regularization, may be used during model training to prevent overfitting. The neural network may include multiple sub-networks, each configured to generate classifications or predictions of different types of output information, which may be combined to form the overall output of the neural network. The machine learning model may alternatively utilize statistical or related algorithms including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, any combination thereof, and / or ensembles and gradient boosted variations thereof.

[0051] If the model generates a classification or prediction of cancer, a non-cancerous disease, a disorder, or a combination thereof, a notification (e.g., an alert or alarm) may be generated and sent to a healthcare provider, such as a doctor, nurse, or other member of the patient's treatment team in a hospital. The notification may be sent via automated phone call, short message service (SMS), multimedia message service (MMS) message, email, and / or an alert in a dashboard. The notification may include output information such as a prediction of cancer, a non-cancerous disease, and / or a disorder; a predicted probability of cancer, a non-cancerous disease, and / or a disorder; an expected time to onset of the cancer, a non-cancerous disease, and / or a disorder; a confidence interval of the likelihood or time; a recommended course of treatment for the cancer, a non-cancerous disease, and / or a disorder; or any combination of such information. In some cases, the expected time to onset of cancer may include a time from detection of the precancerous lesion of at least one year, at least two years, or at least three years.

[0052] To verify the performance of the model, different performance metrics can be generated. For example, the area under the receiver operating characteristic curve (AUROC) can be used to determine the diagnostic, prognostic, screening, or any combination thereof ability of the model. For example, the model can use an adjustable classification threshold, such that the specificity and sensitivity are adjustable, and the receiver operating characteristic curve (ROC) can be used to identify different operating points corresponding to different values ​​of specificity and sensitivity.

[0053] In some cases, such as when the dataset is not large enough, cross-validation may be performed to assess the robustness of the model across different training and testing datasets.

[0054] The following definitions may be used to calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), area under the precision / recall curve (AUPR), AUROC, or any combination thereof: A "false positive" may refer to an outcome in which a positive outcome or result is generated erroneously or prematurely (e.g., before or without the actual onset of cancer, non-cancerous disease, and / or disorder). A "true positive" may refer to an outcome in which a positive outcome or result is correctly generated when a patient has cancer, a non-cancerous disease, and / or disorder (e.g., the patient exhibits symptoms of, or the patient's records indicate, cancer, a non-cancerous disease, and / or disorder). A "false negative" may refer to an outcome in which a negative outcome and / or result is produced but the patient has a cancer, non-cancerous disease, and / or disorder (e.g., the patient exhibits symptoms of, or the patient's records indicate, a cancer, non-cancerous disease, and / or disorder). A "true negative" may refer to an outcome in which a negative outcome or result is produced (e.g., before or without the actual onset of, a cancer, non-cancerous disease, and / or disorder).

[0055] The model may be trained until certain predetermined conditions for accuracy or performance are met, such as having a minimum desired value corresponding to a diagnostic accuracy measure. For example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of occurrence of cancer, a non-cancerous disease, and / or a disorder in a subject. As another example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of worsening and / or recurrence of a cancer, a non-cancerous disease, and / or a disorder for which the subject has previously been treated. Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, AUPR, and AUROC, which correspond to the diagnostic accuracy of detecting or predicting cancer, a non-cancerous disease, and / or a disorder.

[0056] For example, such a predetermined condition may be that the sensitivity of predicting cancer, non-cancerous disease, and / or disorder includes, for example, values ​​of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0057] As another example, such a predetermined condition may be that the specificity of predicting cancer, non-cancerous disease, and / or disorder includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0058] As another example, such a predetermined condition may be that the positive predictive value (PPV) for predicting cancer, non-cancerous diseases, and / or disorders includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0059] As another example, such a predetermined condition may be that the negative predictive value (NPV) of the prediction of cancer, non-cancerous disease, and / or disorder includes, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0060] As another example, such a predetermined condition may be that the area under the curve (AUC) of a receiver operating characteristic (ROC) curve (AUROC) for the prediction of cancer, non-cancerous disease, and / or disorder comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0061] As another example, such a predetermined condition may be that the area under the precision / recall curve (AUPR) predicting cancer, non-cancerous diseases, and / or disorders comprises a value of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0062] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases, and / or disorders with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0063] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases, and / or disorders with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0064] In some embodiments, a model can be trained or configured to predict cancer, non-cancer diseases, and / or disorders with a positive predictive value (PPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0065] In some embodiments, a model can be trained or configured to predict cancer, non-cancer diseases, and / or disorders with a negative predictive value (NPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0066] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases, and / or disorders with an area under the receiver operating characteristic (ROC) curve (AUC) (AUROC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0067] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases, and / or disorders with an area under the precision-recall curve (AUPR) of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0068] The training datasets may be collected from training subjects (e.g., humans), and each training dataset of the training datasets may include a corresponding diagnostic status of a given training dataset of the subject, indicating whether the subject has been diagnosed with a disease (e.g., cancer, a non-cancerous disease, and / or a disorder) or has not been diagnosed with a disease.

[0069] In some embodiments, the model is a neural network or a convolutional neural network. See Vincent et al., 2010, "Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion," J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, "Exploring strategies for training deep neural networks," J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.

[0070] In some embodiments, the data are de-dimensionalized using independent component analysis (ICA), such as that described in Lee, T.-W. (1998): Independent component analysis: Theory and applications, Boston, Mass: Kluwer Academic Publishers, ISBN 0-7923-8261-7, and Hyvaerinen, A.; Karhunen, J; Oja, E. (2001): Independent Component Analysis, New York: Wiley, ISBN 978-0-471-40540-5, which are incorporated herein by reference in their entireties.

[0071] In some embodiments, the data is dedimensionalized using principal component analysis (PCA), such as that described in Jolliffe, IT (2002). Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag. doi:10.1007 / b98835. ISBN 978-0-387-95442-4, which is incorporated herein by reference in its entirety.

[0072] SVM is an SVM, Cristianini and Shawe-Taylor. Theory, Wiley, New York, Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, and Furey et al. al., 2000, Bioinformatics 16, 906-914 (each of which is incorporated herein by reference in its entirety). When used for classification, SVMs separate a given set of binary labeled data using a hyperplane that is maximally far from the labeled data. When linear separation is not possible, SVMs can work in conjunction with the technique of "kernels" that automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.

[0073] Decision tree classifiers are reviewed by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Decision tree-based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One specific algorithm that may be used is a classification and regression tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests—Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.

[0074] Clustering (e.g., unsupervised and supervised clustering model algorithms) is described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings within a data set. To identify natural groupings, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than to samples in other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure is determined. Similarity measures are discussed in Section 6.7 of Duda 1973, and one way to begin a clustering investigation is to define a distance function and calculate a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster will be significantly shorter than the distance between reference entities in different clusters. However, as described in Duda 1973, p. 215, clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x,x') can be used to compare two vectors x and x'. Traditionally, s(x,x') is a symmetric function whose value is large when x and x' are "similar" in some way. An example of a non-metric similarity function s(x,x') is provided in Duda 1973, p. 218. Once a method for measuring "similarity" or "dissimilarity" between points in a dataset is selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. The partition of the dataset that maximizes the criterion function is used to cluster the data. See Duda 1973, p. 217.Criterion functions are discussed in Section 6.8 of Duda 1973. More recently, Duda et al., Pattern Classification, 2nd edition, John Wiley & Sons, Inc., New York, has been published. Clustering is described in detail on pages 537-563. For details on clustering techniques, see Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3rd ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is incorporated herein by reference. Specific exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum-of-squares algorithms), k-means clustering, fuzzy k-means clustering algorithms, and Jarvis-Patrick clustering. In some embodiments, the clustering comprises unsupervised clustering, where no preconceptions about which clusters should be formed are imposed when the training set is clustered.

[0075] Regression models such as multicategory logit models are described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, Inc., New York, Chapter 8, which is incorporated herein by reference in its entirety. In some embodiments, the model utilizes the regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, which is incorporated herein by reference in its entirety. In some embodiments, gradient boosting models are used, for example, for the classification algorithms described herein; these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019). "Gradient Boosting." Hands-On Machine Learning with R. Chapman & Hall. pp. 221-245. ISBN 978-1-138-49568-5, which is incorporated herein by reference in its entirety. In some embodiments, ensemble modeling techniques are used, and these ensemble modeling techniques are described in the classification model implementations herein and in Zhou Zhihua (2012). Ensemble Methods: Foundations and Algorithms. Chapman and Hall / CRC. ISBN 978-1-439-83003-1, which is incorporated herein by reference in its entirety.

[0076] In some embodiments, the machine learning analysis is performed by a device executing one or more programs (e.g., one or more programs stored in non-persistent or persistent memory) that include instructions for performing the data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., a processing core) and memory (e.g., one or more programs stored in non-persistent or persistent memory) that include instructions for performing the data analysis. In some embodiments, the predictive model may be stored in computer memory and executed by one or more processors of the system, as described elsewhere herein.

[0077] system The present disclosure, in some embodiments, describes a computer system that can be programmed to implement one or more methods of the present disclosure, as described elsewhere herein. Figure 6 shows a computer system 600 that can be programmed or otherwise configured to predict cancer, a non-cancerous disease, or a combination thereof, train a predictive model, generate a recommended treatment, or perform any combination of the methods described elsewhere herein. The computer system 600 can be a user's electronic device or a computer system located remotely relative to the electronic device. The electronic device can be a mobile electronic device.

[0078] In some embodiments, computer system 600 may include one or more central processing units 606 (CPUs, also referred to herein as “processors” and “computer processors”), which may be single-core or multi-core processors, or multiple processors for parallel processing. In some embodiments, computer system 600 may also include memory and / or memory locations 604 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 602 (e.g., a hard disk), a communication interface 608 (e.g., a network adapter) for communicating with one or more other systems, and / or peripheral devices 610, such as cache, other memory, data storage devices, and / or electronic display adapters. Memory 604, storage unit 602, interface 608, and peripheral devices 610 may communicate with CPU 606 via a communication bus (solid lines), such as a motherboard. Storage unit 602 may be a data storage unit (or data repository) for storing data. Computer system 600 may be operably coupled to a computer network (“network”) 612 using communication interface 608. Network 612 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. In some cases, network 612 may include a telecommunications and / or data network. Network 612 may include one or more computer servers that may enable distributed computing, such as cloud computing. Network 612, in some cases with the assistance of computer system 600, may implement a peer-to-peer network, which may enable devices coupled to computer system 600 to act as clients or servers.

[0079] CPU 606 may perform one or more steps of the methods described elsewhere herein, which may be embodied in a series of machine-readable instructions, e.g., a program or software. The instructions may be stored in a memory location, such as memory 604. The instructions may be directed to CPU 606, which may then program or configure CPU 606 to implement the methods of the present disclosure described elsewhere herein. Examples of operations performed by CPU 606 may include fetch, decode, execute, and write back.

[0080] The CPU 606 may be part of a circuit, such as an integrated circuit. One or more other components of the system 600 may also be included in the circuit. In some cases, the circuit may include an application specific integrated circuit (ASIC).

[0081] The storage unit 602 may store files such as drivers, libraries, and / or saved programs (e.g., including one or more steps of methods described elsewhere herein). The storage unit 602 may store user data, such as classifications and / or predictions provided by one or more models, training user data, test user data, user preferences, and user programs, or any combination thereof. In some cases, the computer system 600 may include one or more additional data storage units that are external to the computer system 600, such as located on a remote server in communication with the computer system 600 via an intranet or the Internet.

[0082] Computer system 600 can communicate with one or more remote computer systems via network 612. For example, computer system 600 may communicate with a user's remote computer system. Examples of remote computer systems may include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user may access computer system 600 via network 612.

[0083] The methods described herein may be implemented by machine (e.g., one or more processors) executable code stored on electronic storage locations of computer system 600, such as, for example, on memory 604 or electronic storage unit 602. Machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by one or more processors 606. In some cases, the code may be retrieved from storage unit 602 and stored on memory 604 for immediate access by processor(s) 606. In some situations, electronic storage unit 602 may be excluded, and machine-executable instructions are stored in memory 604.

[0084] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or compiled manner.

[0085] In some embodiments, a system described elsewhere herein includes one or more processors and a non-transitory computer-readable storage medium containing software, the software including executable instructions that, when executed, cause the one or more processors of the computer system to receive one or more nucleic acid molecule sequencing reads of a subject's biological sample, wherein the subject has a disease, and the one or more nucleic acid molecule sequencing reads are obtained from one or more nucleic acid molecules enriched by one or more probes exposed to the subject's biological sample; map the one or more nucleic acid molecule sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more nucleic acid molecule sequencing reads; and identify one or more microbial features of the one or more non-human sequencing reads to classify the disease in the subject.

[0086] Aspects of the systems and methods described elsewhere herein, such as computer system 600, may be embodied in programming. Various aspects of the technology may be considered "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in a type of machine-readable medium. The machine-executable code may be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer, processor, etc., or associated modules of tangible memory, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage for software programming at any time. All or portions of the software may be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another computer, e.g., from a management server or host computer to the computer platform of an application server. Thus, another type of medium that may carry software elements includes optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical terrestrial communications networks, and via various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0087] Thus, a machine-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer(s), which may be used to implement the databases, etc., shown in the figures. Volatile storage media may include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media may include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, or DVD-ROMs, any other optical media, punched card paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROMs, and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0088] The computer system 600 may include or be in communication with an electronic display 614, which may include a user interface (UI) 616, for example, to provide a display for visualization of prediction results, and may include one or more interfaces and / or panels for training predictive models, managing and / or manipulating subject and / or patient data, or any combination thereof. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces. ***

[0089] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

[0090] While the method steps described and claimed herein illustrate and / or describe each method or set of operations according to an embodiment, those skilled in the art will recognize many variations based on the teachings described herein. These steps may be completed in different orders. Steps may be added or omitted. Some steps may include sub-steps. Many steps may be repeated as many times as beneficial. One or more steps may be performed simultaneously.

[0091] One or more steps of each method or set of operations may be performed using one or more of the circuitry described herein, e.g., logic circuitry such as a processor or programmable array logic for a field programmable gate array. For example, a circuit may be programmed to provide one or more of the steps of each of the methods or sets of operations, and the program may include program instructions stored on a computer-readable memory or programmed steps of a logic circuit such as, e.g., programmable array logic or a field programmable gate array.

[0092] definition Unless otherwise defined, all technical terms, notations, and other technical and scientific or terminology used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some instances, terms having a commonly understood meaning are defined herein for clarity and / or ease of reference, and the inclusion of such definitions herein should not necessarily be construed as representing a substantial difference from what is commonly understood in the art.

[0093] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the present disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual values ​​within that range. For example, the description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, e.g., 1, 2, 3, 4, 5, and 6. This applies regardless of the size of the range.

[0094] As used in this specification and claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a sample" includes multiple samples, including mixtures thereof.

[0095] The terms "determining," "measuring," "evaluating," "assessing," "assaying," and "analyzing" are often used interchangeably herein to refer to forms of measurement. These terms include determining whether an element is present or not (e.g., detecting). These terms can include quantitative, qualitative, or quantitative and qualitative determinations. Evaluating can be relative or absolute. "Detecting the presence of" can include measuring the amount of something present in addition to determining whether it is present or absent, depending on the context.

[0096] The terms "subject," "individual," or "patient" are used interchangeably herein. A "subject" can be a biological entity containing expressed genetic material. The biological entity can be, for example, a plant, animal, or microorganism, including bacteria, viruses, fungi, and protozoa. A subject can be tissues, cells, and their progeny of a biological entity obtained in vivo or cultured in vitro. A subject can be a mammal. A mammal can be a human. A subject can be diagnosed or suspected of being at high risk for a disease. In some cases, a subject is not necessarily diagnosed or suspected of being at high risk for a disease.

[0097] The term "in vivo," as used herein, describes events that take place inside a subject's body.

[0098] The term "ex vivo," as used herein, describes an event that occurs outside of a subject's body. An ex vivo assay is not performed on a subject. Rather, it is performed on a sample that is separate from the subject. An example of an ex vivo assay performed on a sample is an "in vitro" assay.

[0099] The term "in vitro," as used herein, is used to describe events that occur within a container for holding laboratory reagents such that the material is separated from the biological source from which it is obtained. In vitro assays can include cell-based assays in which live or dead cells are used. In vitro assays can also include cell-free assays in which no intact cells are used.

[0100] The terms "metagenome," "metagenomes," and "metagenomic," as used herein, are used to refer to the sum total of all genomic information represented in a sample, regardless of species of origin, and thus include both human and non-human genomic information.

[0101] The term "metagenomic assembly," as used herein, is used to refer to the process of reconstructing microbial genomes from metagenomic sequencing data, and "metagenomic assemblies" is used to refer to the products of the metagenomic assembly process.

[0102] The term "contig," as used herein, refers to a non-redundant nucleic acid genomic sequence formed by joining one or more smaller sequences based on sequence overlap. The length of a contig can, in some cases, be from about several kilobases (kb) to about several hundred kb.

[0103] The term "bin," as used herein, is used to refer to a collection of one or more contigs that are grouped together because they may originate from the same parent genome.

[0104] The term "bin abundance" or "bin abundance," as used herein, is used to refer to the number of nucleic acid sequencing reads aligned to one or more bins.

[0105] The terms "nodule," "neoplasm," and "tumor," as used herein, are often used interchangeably and are intended to refer to an abnormal growth of tissue, regardless of its malignant status.

[0106] As used herein, the term "about" a number refers to a numerical value plus or minus 10% of that number. The term "about" a range refers to a range that is 10% less than the minimum value of the range and 10% more than the maximum value of the range.

[0107] The use of absolute or sequential terms, such as "will," "will not," "shall," "shall not," "must," "must not," "first," "initially," "next," "subsequently," "before," "after," "lastly," and "finally," when used herein, are not intended to limit the scope of the embodiments disclosed herein but are exemplary.

[0108] Any systems, methods, software, compositions, and platforms described herein are modular and not limited to sequential steps, and therefore terms such as "first" and "second" do not necessarily imply a priority, order of importance, or sequence of actions.

[0109] As used herein, the terms "treatment" or "treating" refer to a pharmaceutical or other intervention regimen to obtain a beneficial or desired result in a recipient. Beneficial or desired results include, but are not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit may refer to the eradication or amelioration of the condition or underlying disorder being treated. Alternatively, therapeutic benefit may be achieved by the eradication or amelioration of one or more physiological symptoms associated with the underlying disorder, such that an improvement is observed in the subject, although the subject may still be afflicted with the underlying disorder. A prophylactic effect includes delaying, preventing, or eliminating the onset of a disease or condition, delaying or eliminating the onset of symptoms of a disease or condition, slowing, halting, or reversing the progression of a disease or condition, or any combination thereof. For prophylactic benefit, subjects at risk of developing a particular disease or who report one or more physiological symptoms of a disease may receive treatment, even though they may not have been diagnosed with the disease.

[0110] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. [Example]

[0111] Example 1: Early detection of lung cancer using blood and tissue metagenomes Most patients with stage I lung cancer have a cell-free tumor DNA fraction of less than 0.1%, precluding the sensitivity of host-based diagnostics that rely on separating normal and tumor-derived molecules. As described in Example 1, we investigated an alternative and orthogonal approach centered on cell-free metagenomics that circumvents the limitations of host-centric methods, resulting in state-of-the-art detection of lung cancers as small as 8 mm in diameter. We performed de novo metagenomic assembly in 5,187 whole-genome-sequenced blood and tissue samples, identified 1,562 pan-cancer bins, and demonstrated their diagnostic capability and cancer type specificity in 17,079 samples across four independent cohorts. Furthermore, amplicon-based (16S rRNA) plasma sequencing was found to approach the diagnostic performance of shotgun-sequenced metagenomics, suggesting cost-effective opportunities while validating the approach. Furthermore, because this strategy is independent of and complementary to host-centric methods, a clinical proteometagenomic diagnostic test was developed and validated that surpassed the diagnostic performance of PET-CT and clinical risk scores in a blinded validation cohort of 106 patients with stage I cancer and lung disease of diverse etiologies (mean ± SD = 1.91 ± 0.84 cm). Collectively, the findings described herein firmly establish the utility of cell-free metagenomics as a generalizable and sensitive strategy for early cancer detection, warranting its application in additional cancer types.

[0112] Existing liquid biopsy methods for lung cancer have size-dependent sensitivity that prevents optimal detection of early-stage disease. Orthogonal, cancer-associated, microbial biomarkers can overcome some of these limitations. The study described in Example 1 (i) describes a metagenomics-centric lung cancer diagnostic test, (ii) performs metagenomic assembly in cancer blood and tissue, (iii) evaluates amplicon-based plasma sequencing for cancer detection, and (iv) demonstrates the superiority of this approach over PET-CT in the setting of stage I disease.

[0113] Until recently, cancer was primarily considered a sterile disease of the human genome, and methods for diagnosing its presence and / or type therefore relied on the detection of human-centric biomarkers. While various diagnostic strategies have been developed, ranging from low-plex PCR-based mutation panels to high-plex assays of millions of DNA fragments, the size of cancer lesions, along with the associated sensitivity of tests, inherently limits the availability and abundance of cancer-derived molecules in circulation. While these challenges can be mitigated, in part, by measuring more biomarkers, sequencing more deeply, and / or testing more frequently (e.g., monthly), ultimately, there are strict limitations imposed by biology on how much cancer-derived material can be present in a vial of blood drawn during early disease, precluding widespread diagnosis of these tumors.

[0114] Among difficult-to-detect cancers, lung cancer is the leading cause of cancer-related deaths worldwide and is frequently detected in later stages of the disease. Despite recent advances in liquid biopsy, blood tests capable of detecting stage I lung cancer lesions, especially those ≤3 cm in diameter (i.e., T1 clinical stage), have proven challenging. For example, in this context, the validation cohort of Lung-CLiP, a recent state-of-the-art circulating tumor DNA (ctDNA)-based test, achieved an area under the receiver operating characteristic curve (AUROC) of only 0.69 among stage I cancers compared with risk-matched controls. Another recent fragmentomics diagnostic demonstrated an AUROC of 0.76 among stage I cancers in its validation cohort, but its performance in the validation cohort was not reported beyond approximately 50% sensitivity at 80% specificity. With the exception of PET-CT scans, the only clinically available blood test for the detection of early-stage malignancies in small nodules (≤3 cm) is the Integrated Clinical Proteomics Panel, which had an AUROC of 0.76 in its validation cohort. These results highlight the need to approach the problem of early lung cancer detection differently.

[0115] Based on intracellular cancer-type-specific communities of intratumoral fungi, bacteria, and viruses whose genomic content is detectable in plasma-derived cell-free DNA (15, 16), we describe a method to address the problem of early-stage lung cancer detection via microbial biomarkers that is independent of and complementary to host-centric approaches, as described in Example 1 and elsewhere herein. First, we evaluated the ability of a simplified signature containing 300 plasma-derived cancer-associated, fungal, and bacterial biomarkers to diagnose early-stage disease in a low-risk clinical setting in 808 treatment-naive patients, demonstrating state-of-the-art performance and validating our approach with amplicon-based plasma sequencing. In the high-risk patient setting, we developed metagenomic assembly bins of cancer-associated microbial DNA utilizing 5,187 whole-genome-sequenced blood and tissue samples. These metagenomic bins provided superior diagnostic performance over 300 known cancer-associated microbial biomarkers for 222 treatment-naive lung disease samples of diverse etiologies. Furthermore, we demonstrated superior diagnostic performance by analyzing their abundance in 17,079 samples from four independent cohorts. Metagenomic bins were found to be broadly cancer type-specific and useful for identifying early-stage lung cancer. Finally, a subset of 454 patients from the cohort with matched clinical and imaging metadata was used to construct a clinical proteometagenomics diagnostic and validated its state-of-the-art diagnostic performance in a blinded validation cohort of 106 patients with nodules as small as 6 mm. As described in Example 1, this study demonstrated the utility and generalizability of plasma-derived metagenomes for early-stage cancer detection.

[0116] subject All subject samples were obtained from biobanked retrospective cohorts at New York University, University of California, San Francisco, Miami Cancer Center, University of North Carolina at Chapel Hill, and two commercial sources. Sample collection was conducted in accordance with protocols approved by each institution's institutional review board. All subjects signed written informed consent to provide samples for research. Subjects with lung cancer or lung nodules included those with suspicious nodules on chest imaging (low-dose computed tomography or chest x-ray) who underwent diagnostic biopsy or surgical resection. Histopathological diagnosis was performed on all biopsied / resected lung tumors, and malignant tumors were assigned to lung cancer subtypes and stages based on histopathology. Disease severity in subjects with sarcoidosis or COPD was scored according to the Scadding staging system and the GOLD criteria, respectively. Patients with a history of cancer or recent (within the past 2 months) antibiotic use were excluded from the study.

[0117] Four hundred fifty-four (44%; 342 lung cancer, 112 non-cancer) samples were originally collected at the NYU Lung Cancer Biomarker Center from subjects with lung nodules who underwent diagnostic biopsy or tissue resection between 2006 and 2021. Eighty-nine subject plasma samples were obtained from Subpopulations and Intermediate Outcome Measures. In the COPD study, SPIROMICS; all 89 subjects were current smokers, 47 were diagnosed with GOLD stage II (mild) COPD, and the remaining 42 were diagnosed with GOLD stage III (severe) COPD. Twenty-five subject samples with pulmonary sarcoidosis were obtained from the UCSF Sarcoidosis Research Program; 12 of the 25 had Scadding stage 2 sarcoidosis, and 13 of the 25 had Scadding stage 4 sarcoidosis. Of the 402 cancer samples, the two major histological subtypes of lung cancer, lung adenocarcinoma (LUAD) and squamous cell carcinoma (LUSC), comprised 57.2% (230) and 29.8% of the cancer samples, respectively. The number (percentage) of cancers represented by pathological stage was as follows: stage I, 131 (32.59%); stage II, 110 (27.36%); stage III, 98 (24.38%); stage IV, 57 (14.18%); and unknown, 6 (1.49%). Smoking status (current / former / never) was known for 708 (68.73%) of the cohort subjects. All CODICES cohort samples (1030) underwent shotgun metagenomic sequencing, and a subset of 335 samples (142 lung cancer, 97 lung disease, and 96 healthy) underwent targeted 16S rRNA gene sequencing. All cancer plasma samples were obtained from treatment-naive patients, and samples from all disease states were age (±5 years) and sex matched with healthy donor plasma samples from two commercial sources.

[0118] Sample processing and sequencing library preparation Total circulating DNA was extracted from 400 μL of plasma from each sample using the QIAamp Circulating Nucleic Acid Kit (QIAGEN 55114) according to the manufacturer's instructions and purified with Agencourt AMPure XP beads (Beckman Coulter). Sequencing libraries were prepared from cfDNA using the KAPA HyperPlus Kit (Roche Diagnostics) and proprietary dual-index primers. The final libraries were analyzed using an Agilent 4200 TapeStation system (high-sensitivity DNA kit) and quantified by qPCR using the NEBNext Library Quantification Kit for Illumina (New England Biolabs). Paired-end 2 × 150 bp sequencing was performed on a NovaSeq 6000 instrument with an S4 flow cell (Illumina).

[0119] 16S sequencing The circular V6 region of 16S ribosomal RNA was first amplified by PCR using primers 967F (0.3 μM, 5′-CNACGCGAAGAACCTTANC-3′) and 1064R (0.3 μM, 5′-CGACRRCCATGCANCACCT-3′) and 10 ng of input DNA. The cycling program was set as follows: 95°C for 5 min, 20 cycles of 98°C for 20 s, 52°C for 15 s, and 72°C for 15 s, with a final extension of 5 min at 72°C. These PCR products were subsequently subjected to two rounds of 10-cycle PCR to incorporate universal adapters suitable for Illumina sequencing platforms and unique sample-specific dual indexing, following the approach described by Glenn et al. (39). PCR conditions were the same as above, except that the annealing temperature was increased to 60°C and the primer concentration was increased to 0.5 μM. All PCR products were purified using AMPure XP beads (Beckman Coulter #A63881) according to the manufacturer's instructions. The final libraries were quantified via qPCR using the Qubit DNA Assay (Thermo Fisher Scientific) and the NEBNext Library Quantification Kit for Illumina (New England Biolabs), and their quality was checked on an Agilent 4200 TapeStation system. Sequencing analysis was performed on an Illumina MiSeq instrument using the Reagent Kit v2 (Illumina, #MS-102-2002).

[0120] Shotgun metagenomic sequencing analysis Reads were quality filtered using fastp to remove adapter sequences and low-quality reads. Reads that aligned to the human genome were separated from non-human reads by aligning them to the complete human reference genome T2T-CHM13 (v2.0) using Bowtie2 with sensitive parameters. Reads that did not align to the human genome were dereplicated using vsearch. Non-human reads were aligned to the RefSeq database (release 206) using Bowtie2 with sensitive parameters. The abundance of each microorganism was summed from the alignment and used for downstream machine learning analysis.

[0121] Plasma proteome analysis Target protein levels in human plasma samples were assessed using the Bio-Plex 200 platform (Bio-Rad, Hercules, CA). Briefly, plasma samples were centrifuged, diluted 1:2, and subjected to Milliplex bead-based immunoassays (Millipore Sigma, Burlington, MA) according to the manufacturer's protocol. The Milliplex HCCBP1MAG-58K (Millipore Sigma, Burlington, MA) panel was used to detect the following protein analytes: cancer antigen 15-3 (CA15-3), cancer antigen 19-9 (CA19-9), carcinoembryonic antigen (CEA), cancer antigen 125 (CA125), interleukin-8 (IL-8), prolactin (PRL), cytokeratin 19 fragment (CYFRA21-1), and osteopontin (OPN). The concentration of each protein analyte was determined using a five-parameter logistic curve fit available with Bio-Plex Manager 6.2 software (Bio-Rad, Hercules, CA) and the protein standards provided with the Milliplex assay.

[0122] TCGA and plasma cell-free microbial DNA metagenome assembly, binning, and analysis First, preprocessed (i.e., trimmed, quality-controlled, and filtered for human reads) whole-genome sequencing (WGS) samples were co-assembled by cancer and sample type. Co-assembly by sample and cancer type was motivated by previous findings that microbes are differentially distributed across cancer types and that most individual samples did not have sufficient read depth to assemble microbial genomes. Co-assembly was performed via metaSPAdes (v. 3.13.1) with 1TB RAM limit allocated across 10 threads and k-mer sizes of 21, 33, 55, 77, 99, and 127. The resulting contigs from each co-assembly were filtered for lengths greater than 1500 and separated as either prokaryotic or eukaryotic using EukRep (v. 0.6.6). Each set of prokaryotic or eukaryotic contigs was binned using Vamb (v.3.0.3) with default parameters on contig abundance profiles estimated independently for each sample via the MetaBAT2 (v.2.12.1) jgi_summarize_bam_contig_depths function. Abundance profiles for each sample were estimated by mapping reads to binned contigs using the Salmon (v.0.13.1) quantification function with the -meta flag enabled. Quality metrics for the resulting prokaryotic refined bin sets were calculated using CheckM (v.1.0.13). All prokaryotic bins were filtered based on CheckM statistics of >10% completeness and <5% contamination.

[0123] Microbial diversity analysis Alpha and beta diversity were calculated using Qiime2 on the microbial abundance table of Rep206 species filtered for overlap with the WIS dataset. For alpha diversity, samples were first purified to the minimum number that would retain the top quartile of samples. Statistical differences between groups were calculated using a two-tailed Mann-Whitney-Wilcoxon test. Beta diversity distances and coordinates were calculated using the Robust Aitchison PCA metric in DEICODE. Statistical differences between groups based on beta diversity were measured using a PERMANOVA test.

[0124] 16S sequencing analysis Raw 16S reads were quality filtered using fastp to remove adapter sequences and low-quality reads. Unique 16S sequences were identified and distinguished from sequencing noise using Deblur in Qiime2. As a secondary method, sequences were also clustered at 90% OTUs using vsearch. Deblur sub-operational taxonomic units (sOTUs) and 90% OTUs were taxonomically identified using a trained Greengenes 13_8 99% OTU classifier and the classify-sklearn method in Qiime2. The resulting feature table was used for downstream machine learning analyses.

[0125] PICRUSt2 was also run on the 90% OTU sequences. First, the 90% OTU table was filtered to a minimum frequency of 10 sequences and a minimum prevalence of 10 samples. Next, PICRUSt2 was run on tables filtered by KEGG ortholog, enzyme, and pathway output data types. These data tables were used for downstream machine learning analyses.

[0126] Genome coverage analysis The genome coverage of each genome in Rep206 was calculated. Briefly, for each microbial genome, the total number of uniquely covered bases in the genome was calculated by aggregating alignments from all plasma samples in the dataset. The same process was applied to all sequencing blank samples. Then, for each Rep206 genome in the plasma and blank samples, the percentage of covered bases across the total bases in the genome was calculated.

[0127] Stack Machine Learning Approach Different ML architectures can provide different performance on the same feature set. This observation was found to be relevant when attempting to link multi-omics multi-species (host and microbial) data together, particularly when linking categorical data (e.g., smoker status) and quantitative data (e.g., metagenomic abundance) on the same sample. In general, logistic regression-like classifiers performed better on categorical data and protein abundance, but worse on metagenomic data. This suggested the need to significantly optimize the hyperparameters of a given single ML model or identify different mechanisms by using multiple models to reach a compromise in model performance on multiple types of data.

[0128] Inspired by recent multi-omics models (PMID: 34875674) and modern convolutional neural networks (PMID: 31631918) for predicting breast cancer therapy response, stacked ensemble modeling was implemented to integrate predictions from multiple models, each of which would ideally "pick up" different aspects or distributions in the data. Combinations using two or more of the following models were explored: random forest, gradient boosting machine, logistic regression, elastic net, and support vector machine (linear, radial, and radialSigma). Internal testing (data not shown) revealed that optimal performance was often achieved using three models in parallel: random forest, elastic net, and gradient boosting machine. Therefore, these three base algorithms were used for ensemble modeling, followed by a meta-learner including logistic regression to weight the probability outputs of the base algorithms, while calculating the final prediction based on the weighting. To ensure that the base algorithms were aligned and provided predictions on the same holdout folds at the same time, a standardized 10-fold cross-validation index was predefined and fed into the ensemble modeling. The caretEnsemble (https: / / github.com / zachmayer / caretEnsemble) R package was used to perform model tuning, CV-fold synchronization, hyperparameter tuning of the base algorithms, and weight adjustment of the meta-learner.

[0129] For the base algorithm, the number of hyperparameter combinations was fixed using the "tuneLength" argument of caretEnsemble in the caretList() function and set equal to 20. This value represents the maximum number of iterations to try for each hyperparameter in the model, which may have several hyperparameters, so the total number of hyperparameters evaluated is the combination and sum of the values. For example, the Random Forest base algorithm has a single adjustable hyperparameter "mtry", which means that up to 20 mtry values ​​are tried, as follows: {2, 3, 4, 7, 10, 16, 25, 39, 60, 91, 140, 215, 328, 503, 770, 1178, 1802, 2757, 4219, 6454}. Obviously, the Gradient Boosting Machine has two adjustable hyperparameters ("n.trees" and "interaction.depth"), which means that for each, up to 20*20 values, including combinations, are tried, or up to 400 values ​​in total. These hyperparameters ranged as follows: n.trees, 50 to 1000 × 50; interaction.depth, 1 to 20 × 1. Similarly, the elastic net model has two adjustable hyperparameters ("alpha", "lambda"), resulting in 20 × 20 combinations.Possible automatically generated alpha parameters include: {0.1, 0.147368421052632, 0.194736842105263, 0.242105263157895, 0.289473684210526, 0.336842105263158, 0.384210526315789, 0.431578947368421, 0.478947368421053, 0.526315789473684, 0.573684210526316, 0.621052631578947, 0.668421052631579, 0.71578947368421, 0.763157894736842, 0.810526315789474, 0.857894736842105, 0.905263157894737, 0.952631578947368, 1}. Possible automatically generated lambda parameters include: {0.00662229466399123, 0.00824606200985824, 0.0102679723752198, 0.0127856492677635, 0.0159206531946638, 0.0198243509450728, 0.0246852240035688, 0.0307379689652752, 0.0382748293462366, 0.0476597059206647, 0.0593457268717406, 0.0738971260921729, 0.0920164859802709, 0.114578660090195, 0.142673013517157, 0.177656020501926, 0.221216758814626, 0.275458463170505, 0.343000075305506, 0.427102693834312}. If the number of features was less than the allowed hyperparameters, those respective hyperparameters were excluded from the grid search. Unless the feature set was subset (e.g., PET-CT SUVs only), the hyperparameter search for the base algorithm was performed stepwise (within each model) across up to 820 different combinations.

[0130] For model comparisons with minimal feature sets, such as only CEA or OPN or PET-CT SUV, a stacked ML architecture was used, with CV-fold matching to ensure comparable performance values. The only change was that for single features, elastic nets could not be included in stacked ML due to the inability to regularize single features. As an alternative in these situations, elastic nets were swapped for simple logistic regression and used with random forest and gradient boosting machine-based algorithms and meta-learners.

[0131] To (i) ensure approximately equal representation of samples across sequencing runs and diagnoses within individual CV-folds, (ii) ensure that CV-folds were consistent during base algorithm training, and (iii) provide reproducibility and comparability across feature sets, a CV-fold index was calculated and stored in the metadata for each comparison type after subsetting and before stackML. For example, in the lung cancer vs. healthy comparison using metagenomic data (Figure 7C), the following was done: (a) Starting with the metadata for the entire CODICES cohort, samples were subset into 808 healthy subjects or subjects with cancer; (b) using the data subsets, predefined 10-fold CV indexes were calculated based on stratified random sampling across sequencing runs and diagnoses in the metadata to ensure equal representation, and these indexes were saved as new metadata columns; (c) during stackML, predefined CV-folds were used for the base algorithm and meta-learner. Furthermore, these CV-folds were determined in the sample metadata database, meaning that all feature sets were trained and evaluated using the same CV-folds. Therefore, the observed performance of different feature sets (Figure 7C) is directly comparable to each other, as are the AUROC confidence intervals calculated by using the performance on the individual CV-folds.

[0132] For metagenomic data normalization, unsupervised batch-corrected metagenomic count data (see above) were transformed using a centered log-ratio transformation prior to stackML. This was done using each subset immediately prior to stackML, meaning that it occurred on the 808 healthy and lung cancer samples in each stackML (Figure 7C) rather than the full CODICES cohort (n = 1030).

[0133] Because protein concentrations were already corrected for each plate using multidimensional normalization (see above) (PMID: 27570895), they were not further transformed. Normalized CEA and OPN values, if included, were simply concatenated as features into the larger feature table prior to stack ML.

[0134] For clinical metadata normalization, different approaches were applied depending on the class of data. For numerical data (e.g., tumor size in centimeters), available data were coerced to numeric values, and any resulting "NA" values ​​were assigned a "0" (i.e., in logistic regression, any weighting factor multiplied by 0 is 0, so they "drop out"). For categorical data, features were considered. For Boolean data, data were coerced to logical values ​​(i.e., binary data) unless the data were incomplete, and a third category ("unknown") was added. The normalized categorical and Boolean columns were then converted into dummy variable columns using the caret dummyVars() function. For example, a single column containing categorical data with three levels ("Level A," "Level B," and "Level C") was converted into a three-column data frame of binary (1 or 0) entries with column names of "Level A," "Level B," and "Level C." When multiple metadata variables were used, such as a numerical variable with a categorical variable, the resulting normalized columns with dummy variables were concatenated together. Of note, the log odds and probability of Mayo clinical risk score for nodal malignancy ("pCA") were calculated using existing clinical metadata variables in our data, where possible, as follows: {logOdds=0.0391*age+0.1274*tumor_size_mm+0.7917*smoker_binary+0.7838*upper_lung_nodule_binary+1.0407*spiculation_binary-6.8272}. Of note, all CODICES cohort samples had no history of cancer, so this was excluded from the logOdds calculation. Additionally, because the diagnostic model is designed to be independent of, but complementary to, PET-CT imaging, PET-CT, which is optional in the pCA calculation, was also excluded from the logOdds formula. The logOdds were then converted to pCA probability as follows: {pCA_probability=100*exp(logOdds) / (1+exp(logOdds))}.

[0135] Importantly, we addressed and mitigated situations where missing data artificially inflated ML performance. This issue was originally identified when we included "smoker status" in the low-risk clinical environment model (Figure 7C). While it resulted in a much higher AUROC between subjects with lung cancer and healthy controls, inspection of the importance of the trained features indicated that "unknown" smokers were among the highest-ranking variables. Further inspection of the metadata then revealed a problem: most patients with lung cancer had known smoking status (394 / 401 valid entries), but healthy subjects did not (115 / 408 valid entries). Thus, missing smoking status, converted to "unknown" during metadata normalization, contributed to the artificially boosted performance. This is, in part, why only metagenomic and / or proteomic data were used for the healthy vs. lung cancer comparison (Figure 7C) and early lung disease vs. lung cancer comparison (Figure 8G), which were available for 100% (1030 / 1030) or 99.22% (1022 / 1030) of the entire CODICES cohort, respectively. Furthermore, subsequent clinical proteometagenomic models utilizing clinical metadata were trained and evaluated using only a subset (n = 454) of the CODICES cohort in which most samples contained this information (e.g., tumor size was available for 431 / 454 samples).

[0136] To calculate ROC curves, predictions for the holdout folds (using the adjusted stackML model) were concatenated to generate a single prediction set with the same number of original samples. For example, data from 808 samples were input into stackML for the lung cancer vs. healthy comparison (Figure 7C). The prediction list had 808 rows associated with it, one row per sample. Notably, this table also contained information about which samples belonged to which holdout fold, allowing for the calculation of the AUROC for each CV-fold, which was then aggregated to estimate 95% confidence intervals for performance. Additionally, the full-length concatenated prediction list was subsetted for individual clinical stages and histological subtypes of lung cancer, as well as their control samples (i.e., healthy controls or lung disease), followed by regeneration of the ROC curves (e.g., Figure 7C, center and right panels). This process has been performed elsewhere by Mathios et al. (7). In these cases, the CV-fold information after sample subsetting was used to estimate confidence intervals for AUROC performance for each stage and histological subtype comparison. For clarity, this means that in Figure 7C, for a given feature set, a single stack model was built and evaluated as a single diagnostic test.

[0137] For the 16S rRNA cohort, which was a subset of the CODICES cohort processed independently for amplicon-based sequencing, we also used the stacked ML architecture described above by replacing the microbial feature table with 16S rRNA taxonomy or PICRUSt2-predicted functional pathway abundances. Furthermore, CEA, OPN, and clinical metadata were available for overlapping patients in the CODICES cohort, allowing clinical proteometagenomic assessment with amplicon-derived abundances (Figure S10D).

[0138] Another method for integrating predictions from multiple models is "multimodal" learning, in which multiple classifiers are independently developed for different features of the same sample, and the predictions from those classifiers are integrated using a meta-learner. This differs from the stacked ensemble learning used, in which a single feature table is evaluated by multiple classifiers and subsequently by a meta-learner. Such "multimodal" learning approaches (data not shown) have been evaluated but have not been found to perform better than stacked ensemble models and are also found to be more complex to optimize.

[0139] CODICES cohort: shotgun metagenomic alignment, decontamination, and normalization Reads were quality filtered using fastp to remove adapter sequences and low-quality reads. Reads that aligned to the human genome were separated from non-human reads by aligning them to the human reference genome GRCh38 using Bowtie2 with a sensitive parameter set, followed by alignment to GRCh38 using SNAP. Reads that did not align to the human genome were dereplicated using vsearch. Non-human reads were aligned to the RefSeq database (release 206) ("rep206") using Bowtie2 with a sensitive parameter set. The abundance of each microorganism was summed from the alignment and used for downstream analysis.

[0140] For in silico decontamination, decontamination was applied in prevalence mode using all 98 extraction blanks. This was processed in parallel with the CODICES cohort plasma samples using default parameters. After examining the histogram of calculated prevalence p-statistics for decontamination (Figure 11A), we maintained the default cutoff of P* = 0.1. In prevalence mode, the p-statistic is set equal to the p-value from a chi-squared or Fisher's exact test, but we noted that it is not directly treated as a p-value for making decisions about specific taxa. We also noted that prevalence-based decontamination is effective even in very low biomass environments. In total, 5,630 taxa in the raw rep206 data in the CODICES cohort were flagged as putative contaminants, and 6,455 taxa were retained (53.41% of the total). The 5,630 removed taxa accounted for 0.88% of the total reads.

[0141] For feature set intersection with Tsay et al. (“Zebra: Static and Dynamic Genome Cover Thresholds with Overlapping References,” mSystems 2022; e0075822, hereafter “Tsay”), 16S rRNA results were shared by each first author and used to identify overlapping microorganisms in the rep206 feature set at the genus level. Specifically, because shotgun metagenomics data from the CODICES cohort utilized species elsewhere (e.g., in the “Decontamination” and “WIS” feature sets), all species with corresponding genus-level overlap with 16S rRNA data collected by Tsay from bronchoalveolar lavage samples were retained.

[0142] For feature set intersection with the Weizmann Institute of Science (WIS) catalog of cancer-associated bacteria and fungi, the first author of Narunsky-Haziza et al. ("Pan-cancer analyses reveal cancer type-specific fungal ecologies and bacteriome interactions. Cell"; hereafter "Narunsky") shared a bacterial and fungal "hit" list across all tissue samples (tumor, NAT, or true normal [breast only]) from the two studies. This "hit" list was filtered to microorganisms with species-level results using a multiregion amplicon sequencing approach for bacteria or ITS2 sequencing for fungi. This was then intersected with the rep206 data from the CODICES cohort, leaving 300 overlapping features.

[0143] Due to their low biomass nature, investigations of cancer-associated bacteria have previously demonstrated batch effects that often need to be considered before downstream analysis. Previous analyses in TCGA demonstrated this is particularly necessary when comparing data across sequencing centers, sequencing platforms, and experimental strategies (WGS vs. RNA-Seq). The CODICES cohort also confirmed that inter-sequencing run effects in raw data must be corrected before downstream analysis to ensure results are not explained by inter-sequencing run variability (although additional randomization of sample types between sequencing runs mitigated this effect). Furthermore, while batch effect correction using a combination of Voom and SNM in TCGA cancer microbiome data has previously been addressed and yielded useful results, the Voom-SNM strategy is inherently supervised, making it difficult to use in diagnostic approaches. More specifically, if the purpose of the test is to obtain such information, batch correction cannot be supervised using biological information unless the supervision information is readily available in some other way (e.g., using age and / or sex as biological variables). After evaluating various methods (data not shown), we found that ComBat-Seq, designed to operate on discrete counts based on a negative binomial model, completely eliminated the sequencing run effect on the raw rep206 data of the CODICES cohort when applied in an unsupervised manner across sequencing runs. Specifically, ComBat-Seq (in the sva package v.3.35.2 in R) was run using default parameters, including a single vector (the "batch" variable) indicating which sequencing run each sample belonged to, without any additional information (i.e., unsupervised corrections). One additional advantage of the ComBat-Seq approach is that both input and output are discrete counts, unlike Voom-SNM, which requires log transformations and accompanying pseudocounts to allow for those log transformations. This means that ComBat-Seq-corrected counts are in the same units as the original counts and can be used directly in downstream analyses that require separate count data inputs.For a subset of the CODICES cohort data (Figure 7C), we first reduced the taxonomic features (e.g., to 300 WIS overlapping features or 6530 Tsay overlapping features) before applying unsupervised ComBat-Seq normalization to the entire 1030-sample dataset for that feature set. Treating metagenomic bin abundances similarly, we performed unsupervised ComBat-Seq on the entire CODICES cohort, again using default parameters and a single "batch" vector representing the sequencing run for each sample.

[0144] All downstream analyses then used a single unsupervised batch-corrected normalized dataset (one for each feature set). For analyses using subsets of the CODICES cohort (e.g., lung cancer samples vs. healthy controls), the unsupervised batch-corrected normalized dataset was then subset to those samples of interest before further analysis. This meant that, once normalized, a single version of the data was used for all downstream analyses for each feature set; in other words, the batch-corrected subsetting avoided creating duplicates of similar but slightly different datasets for each subset.

[0145] The CODICES cohort: Generation and normalization of plasma proteomic data Target protein levels in human plasma samples were assessed using the Bio-Plex 200 platform (Bio-Rad, Hercules, CA). A total of 99.22% (1022 / 1030) of the samples in the CODICES cohort had sufficient plasma to obtain protein concentrations. Briefly, plasma samples were centrifuged, diluted 1:2, and subjected to Milliplex bead-based immunoassays (Millipore Sigma, Burlington, MA) according to the manufacturer's protocol. Carcinoembryonic antigen (CEA) and osteopontin (OPN) were detected using the Milliplex HCCBP1MAG-58K (Millipore Sigma, Burlington, MA) panel. Concentrations of each protein analyte were determined using a five-parameter logistic curve fit available in Bio-Plex Manager 6.2 software (Bio-Rad, Hercules, CA) and the protein standards provided with the Milliplex assay.

[0146] Samples with protein concentrations outside the range ("OOR<" or "OOR>") were assigned a concentration 10% higher (for "OOR>") or 10% lower (for "OOR<") than the upper or lower limit provided by the protein standard, respectively. For CEA, the upper standard limit was 18,556 pg / mL, and the lower standard limit was 25.45 pg / mL. For OPN, the upper standard limit was 400,000 pg / mL, and the lower standard limit was 548.7 pg / mL. A total of 32 plates were used to process the CODICES cohort samples for protein analytes. Plate-specific effects were removed in an unsupervised manner among the CODICES cohort samples using multidimensional normalization using their accompanying R package MDimNormn (v. 0.8.0) with default settings, as described by Hong et al. (43). After these steps, any samples with missing values ​​were assigned a "0" concentration to avoid sample dropout in downstream machine learning analyses that could not utilize the missing data.

[0147] Sensitivity and specificity analysis and ROC comparison Sensitivity at a fixed specificity and specificity at a fixed sensitivity were analyzed using 1000 bootstrap resamplings, using the approach described by Chabon et al. ("Integrating genomic features for non-invasive early lung cancer detection," Nature. 2020;580:245-51). When plotting, median and interquartile range across bootstrap resamplings were plotted. We noted that two other diagnostic classifiers in this field report their performance as either "specificity at 97% sensitivity" or "sensitivity at 80% specificity." This is why model testing is also evaluated at these cutoffs. ROC comparisons were performed using Delong's method with the pROC (v.1.18.0) R package. Because predictions were made on the same samples using models with different feature sets (meaning the predictions were correlated), comparisons of sensitivity and specificity were calculated on non-bootstrapped data using McNemar's test.

[0148] TCGA analysis Bioinformatic alignment of 15,512 strictly host-, base-quality-, and length-filtered (>45 bp) TCGA samples to metagenomic bins showed that 14,809 samples (95.47%) had ≥1 alignment to 1,545 bins. Normal adjacent tissue (NAT) tissues that were not part of the metagenomic assembly constituted 73.26% (515 / 703) of the removed samples. The remaining samples were then quality-controlled for metadata. Specifically, (i) 38 samples consisting solely of acute myeloid leukemia were removed from Johns Hopkins due to the inability to distinguish between possible sequencing center batch effects and biological effects, and (ii) samples with missing sequencing center information were removed (n=22). The remaining 14,749 samples were then subset to those sequenced on an Illumina HiSeq machine. This accounted for 98.00% (14,454 / 14,479) of the samples, leaving 14,454 samples for downstream analysis. Metagenomic bin abundances in TCGA were then analyzed using the method described in detail by Narunsky.

[0149] Briefly, batch correction, if applied, was used to convert discrete counts to pseudo-normally distributed data using Voom, followed by SNM to remove batch effect(s) in a supervised manner. The only supervised information used was "sample type" (e.g., "blood-derived normal," "primary tumor," "solid tissue normal"), and batch correction factor(s) included sequencing center ("data_submitting_center_label") and experimental strategy ("experimental_strategy"), if applicable.

[0150] The raw count data were examined separately after (i) subsetting to individual sequencing centers and experimental strategies, or (ii) subsetting to WGS data only, as this was less subject to technical batch effects than disease type (Figure 12D).

[0151] Machine learning was performed using gradient boosting models (GBM) with 10-fold cross-validation with 10 independent stratified 10% holdouts. ROC and PR curves and areas were calculated for each independent 10% holdout test set, yielding 10 sets of two-class discrimination performance for each model (effectively 10 sets of 90% training and 10% testing). The GBM hyperparameters were fixed: {n.trees=150, interaction.depth=3, shrinkage=0.1, n.minobsinnode=1}. These performance estimates for the 10 folds were then aggregated for each model to estimate 95% confidence intervals of performance. In cases of class imbalance, minority class upsampling was used, but ≥20 samples per class were required to perform ML comparisons of any two classes. Two-class comparisons included (i) one cancer type versus all other cancer types for primary tumor and blood analysis, and (ii) unpaired primary tumor versus adjacent normal. When using Voom-SNM normalized data, features were centered and scaled before ML. When using raw count data, only zero-variance features were removed before ML construction. Otherwise, no explicit feature selection or data transformation was performed. For multi-class ML using either batch-corrected data (Figures 9I and 13D) or raw WGS data (Figure 13E), an extreme gradient boosting (XGBoost) machine was used with 10-fold cross-validation, zero-variance feature removal, and the following fixed hyperparameters: {nrounds=10, max_depth=4, eta=0.1, gamma=0, colsample_bytree=0.7, min_child_weight=1, subsample=0.8}.

[0152] Negative control machine learning analyses were performed as described above, except that either (i) the predicted label metadata was scrambled or (ii) sample IDs in the count data were dynamically shuffled immediately before ML model building. The differences between global scrambling and shuffling (i.e., once before all ML models were built and tested) and dynamic scrambling and shuffling (i.e., immediately before ML model building but after data subsetting and labeling) were previously examined, and dynamic scrambling and shuffling yielded more consistent results (less variance) and demonstrated greater agreement with known null values ​​(i.e., 50% AUROC and AUPR positive class prevalence). Therefore, dynamic scrambling and shuffling was used as a negative control when comparing performance with actual samples. Because the same random number seeds were used for both the actual and negative control analyses, they were evaluated on the same cross-validation split. Downstream statistical analyses compared the performance of the actual analysis with the negative control analysis and found statistically superior performance for the former.

[0153] Briefly, alpha and beta diversity analyses were performed on a subset of samples, including individual sequencing centers, WGS, and sequencing platforms (Illumina HiSeq), using Qiime2 (v. 2021.11) and their respective plugins. Alpha diversity was calculated using the non-phylogenetic core metric function in Qiime2 after purifying to 15,000 reads / sample (the first quartile of sample read distribution between primary tumor and blood samples). Beta diversity analysis was performed using RPCA (Robust Aitchison Distance) in DEICODE. DEICODE does not require purifying data by design, and the ADONIS implementation in Qiime2 subsequently calculated the PERMANOVA statistic.

[0154] Briefly, for differential abundance testing, ANCOM-BC was applied iteratively within WGS sequencing center subsets to assess one-versus-all-others comparisons between cancer types using metagenomic bin abundances in primary tumors (Figure 22) or blood (Figures 25A-25E). The following parameters were used: {p_adj_method="BH", zero_cut=0.999, lib_cut=1000, tol=1e-5, max_iter=100, conserve=FALSE, alpha=0.05, global=FALSE}. A minimum of 10 samples was enforced in each class before calculating differentially abundant taxa; otherwise, the comparison was skipped. Statistical discrimination was performed for each cancer type within each subset, and then volcano plots were created using the calculated beta, p, and BH-adjusted q values.

[0155] Hopkins cohort analysis Bioinformatics alignment of 537 strictly host-, base-quality-, and length-filtered (>45 bp) samples from Cristiano et al. (“Genome-wide cell-free DNA fragmentation in patients with cancer.” Nature. 2019;570:385-9; hereafter referred to as “Hopkins” or “Cristiano”) to metagenomic bins yielded ≥1 alignments for all 537 samples to 770 bins. As before, only treatment-naive samples were used for downstream analysis. If a patient had ≥1 sample, the earliest timepoint sample was selected, leaving 491 samples (91.43% of the total) for downstream analysis. Metagenomic bin abundance in the Hopkins cohort was then analyzed using the methods described above and detailed by Narunsky et al.

[0156] Briefly, for machine learning for individual cancers or for comparisons of one cancer type versus all other types, GBM on raw metagenomic bin abundances was applied using 10-fold cross-validation with the following fixed hyperparameter set (same as TCGA): {n.trees=150, interaction.depth=3, shrinkage=0.1, n.minobsinnode=1}. Zero-variance features were removed prior to ML, and upsampling occurred in case of class imbalance. Performance on each independent holdout fold was used to estimate 95% confidence intervals for AUROC and AUPR.

[0157] Briefly, for the machine learning comparison between grouped cancer samples and healthy controls (Figure 9M), we used the same machine learning architecture, but with 10 replicates using 10-fold cross-validation, consistent with the approach used by Cristiano et al. In this case, confidence intervals were estimated using replicates rather than individual folds. For plotting, we adapted the R scikit-learn approach (https: / / scikit-learn.org / stable / auto_examples / model_selection / plot_roc_crossval.html) to estimate the average AUROC and AUPR curves across the 10 replicates. This can be a challenging task because the specificity breaks for the 10 model replicates are not always equal to each other, requiring interpolation. Specifically, to obtain the average performance lines, we performed linear interpolation using the approx() base R function for each ROC and AUPR curve across 1000 equally spaced points between 0 and 1, ensuring that each average curve began and ended at a corner of the plot. The mean ROC and PR curves and associated 95% confidence intervals were then calculated at each point using the 1000 interpolated y-values ​​between x = 0 and x = 1. These mean performance lines were overlaid with ribbons of the 95% confidence intervals, which showed good agreement.

[0158] For the ML analysis of the negative control, samples were dynamically scrambled for their metadata or dynamically shuffled for their count data, as is done in TCGA, and the same random number seed was used to ensure that cross-validation partition comparisons and non-scrambled / shuffled analyses were consistent. Statistical comparisons between the actual and control analysis performances were then performed, showing significantly better performance for the actual samples (Figure 9K).

[0159] For beta diversity, DEICODE's RPCA (robust distance) (same as for TCGA and CODICES cohorts) was applied using the DEICODE plugin in Qiime2, followed by calculation of PERMANOVA statistics using the Qiime2 add-in implementation.

[0160] Hong Kong Liver Cancer (HCC) Cohort: Data Access and Processing Plasma sequence data from liver cancer patients and healthy individuals were downloaded from the European Genome-Phenome Archive (EGA) accession number EGAS00001001024. 42 reads were quality filtered using fastp, host-depleted against GRCh38, and then aligned to metagenomic bins using Salmon. From metagenomic bins, alpha diversity, beta diversity using Jaccard and RPCA,51 and differential bin abundance using ANCOM-BC.44, statistical differences in alpha diversity were calculated using Wilcoxon tests, and beta diversity metrics were calculated using PERMANOVA. Nucleotide frequency (NTF) data were generated by Budhraja et al. (PMID: 36630480), and their combination with bin abundance was used for downstream machine learning analysis.

[0161] Fragment terminal motif 9 nucleotide frequency (NTF) analysis Nucleotide frequency information around the fragment ends was calculated using quality-filtered sequencing reads before removing human reads. Nucleotide frequencies were calculated for nine relative positions around the read fragment ends. These positions were previously identified as being most informative for cancer classification (PMID: 36630480). This resulted in a table with the frequency of each base at each of the nine relative positions for each sample. This information was used in conjunction with microbial, proteomic, and / or clinical data for downstream machine learning analysis.

[0162] Lung-gut cohort: data access and processing Stool 16S rRNA sequence data from patients with LUAD and healthy individuals were downloaded from the European Nucleotide Archive (ENA) accessions PRJEB44169 (LUAD sample) and PRJEB33905 (healthy sample, PMID: 34586729). Reads were quality filtered by fastp, host-depleted against GRCh38 (for consistency, human reads were not expected to contribute to the 16S data), and then aligned to metagenomic bins using Salmon. From metagenomic bin abundances, alpha and beta diversity were calculated using Jaccard and RPCA (PMID: 30801021), and differential bin abundances were calculated using ANCOM-BC (PMID: 32665548). Statistical differences in alpha diversity were calculated using Wilcoxon tests and beta diversity metrics by PERMANOVA.

[0163] Blinded validation cohort analysis The blinded validation cohort included 108 samples, each sent in pairs with a ≤1 mL plasma aliquot and clinical metadata (no label or diagnosis). Plasma was processed using the same methods as described above for the CODICES cohort to obtain shotgun metagenomic data in an independent sequencing run. A small amount of plasma was also processed for proteomic information (CEA, OPN). Clinical metadata was normalized to the same metadata as in the CODICES cohort (described above), including calculation of the Mayo Clinical Risk Score (pCA). After reviewing the metadata, two samples were discarded based on the quality of their metadata, particularly due to missing ages ("0") or unverifiable ages (years ending in "00" were either 22 or 122 years, representing either the youngest or oldest outlier in the cohort), which would have affected the calculation of pCA. This left 106 blinded samples for evaluation of the clinical proteometagenomic study.

[0164] Before making predictions for the blinded cohort, we took into account possible sequencing run and protein plate variation using an unsupervised normalization approach. Specifically, for the metagenomics data, the following steps were taken: (i) raw metagenomic bin abundances from 454 samples in the CODICES cohort with available clinical proteometagenomic information were row-joined with raw data from the blinded validation cohort; (ii) unsupervised ComBat-Seq was run on the overlaid table using a "batch" vector containing the 454 CODICES samples and the sequencing runs for the independent validation run, without any other information or default settings; (iii) the 454 CODICES samples normalized with the metagenomic bins with available clinical proteometagenomic information were reserved for model training; (iv) the normalized metagenomic bin data of the blinded validation cohort samples was reserved for model prediction; and (v) the normalized metagenomic bin data reserved for the blinded validation cohort samples was independently centered log-ratio transformed.

[0165] For the proteomic data, the following steps were taken: (i) blinded validation cohort samples with protein concentrations that were out of range ("OOR<" or "OOR>") were assigned a concentration that was 10% higher (for "OOR>") or 10% lower (for "OOR<") than the upper or lower limit provided by the protein standard (as in the CODICES cohort); (ii) protein concentrations from the blinded validation cohort were row-combined with the raw protein concentrations from the 454 CODICES samples; (iii) plate-specific Effects were removed from the combined protein data in an unsupervised manner using multidimensional normalization; (iv) any samples in the CODICES cohort with missing values ​​were assigned a "0" concentration to avoid sample dropout (Note: all blinded validation cohort samples had protein concentrations available); (v) normalized protein concentrations for 454 CODICES samples were reserved for model training; (vi) normalized protein concentrations of blinded validation cohort samples were reserved for model prediction.

[0166] For the clinical metadata of the blinded validation cohort, the data were normalized in the same way as for the CODICES cohort. For any clinical variable instance found in the CODICES cohort that was not found in the blinded validation cohort (e.g., tumor_solidity="partially solid and shattered glass" was an instance only in the CODICES cohort), an empty, zero-value dummy variable column was then added to the clinical metadata of the blinded validation cohort. This was particularly necessary because (a) the stacked ML approach uses dummy variables for categorical and Boolean features, and (b) the stacked ML-tuned model requires the same features to be present to make predictions that were originally used for training.

[0167] The metagenomic and proteomic data for the CODICES cohort were normalized (n = 454 samples), and final models were trained using these data in parallel with clinical metadata using the 820-step hyperparameter grid search described above. Specifically, the following features were included: metagenomic bin, CEA, OPN, pCA probability, Brock probability, smoker status, emphysema status, tumor size (cm), spiculation, tumor solidity, and superior pulmonary nodules. As before, the metagenomic data were centered log-ratio transformed immediately prior to stacked ML. Additionally, equivalent stacked ML models were built for the same samples using only PET-CTSUV, pCA probability, or Brock probability.

[0168] The final trained CODICES cohort stacked ML model was then applied to the normalized blinded validation cohort clinical proteometagenomic data to obtain probability predictions for each sample. This process was then repeated for CODICES cohort stacked ML models built using only PET-CT SUVs, pCA probabilities, or Brock probabilities. The probability predictions were then sent to the blinded cohort provider, who returned the ROC curves and their respective AUROCs.

[0169] CODICES, 16S rRNA, TCGA, Hopkins, and validation cohorts: statistical analysis Downstream analyses and plots were generated in R (either version 4.03 or 4.1.1). Common R packages used included phyloseq (v. 1.38.0), vegan (v. 2.5-7), doMC (1.3.7), dplyr (v. 1.0.7), reshape2 (v. 1.4.4), ggpubr (0.4.0), ggsci (v. 2.9), rstatix ​​(v. 0.7.0), tibble (v. 3.1.6), caret (v. 6.0-90), caretEnsemble (v. 2.0.1), gbm (v. 2.1.8), xgboost (v. 1.5.0.1), and randomForest (v. 4.0.1). The following packages were used: glmnet (v. 4.1-3), MLmetrics (v. 1.1.1), PRROC (v. 1.3.1), pROC (v. 1.18.0), e1071 (v. 1.7-9), gmodels (v. 2.18.1), ANCOM-BC (v. 1.4.0), decontam (v. 1.14.0), limma (v. 3.50.0), edgeR (v. 3.36.0), snm (v. 1.42.0), sva (v. 3.35.2), biomformat (v. 1.22.0), and Rhdf5lib (v. 1.16.0). The Rstatix ​​package was corrected for multiple hypothesis testing when applicable. Sample size was not estimated a priori, and no power calculations were performed. The gbm, randomForest, and glmnet packages were used for two-class ML. The xgboost package was used for multi-class gradient boosting ML. AUROC and AUPR were calculated using the pRROC package for all analyses except for the validation analysis, which used pROC. We note that the R programming language has two numerical limitations regarding calculating small numbers, including p-values: (1) double eps such that 1 + x != 1, or the smallest positive floating-point number x (which is 2.220446 × 10 -16 (ii) double x min , or the smallest non-zero normalized floating-point number (which is 2.225074 × 10 -308(However, this limit may be lower depending on the computing environment.) Some R packages, notably ggpubr, do not report p-values ​​below double eps, so for our data, p<2.2 × 10 -16 Conversely, other R packages, notably rstatix ​​(listed below), use double x min reported a low p-value of double x min p-values ​​less than p<2.2×10 -308 These are not ranges of p-values.

[0170] Plasma microbiome discovery cohort To more comprehensively capture the distribution of the plasma microbiome and evaluate its diagnostic utility, we constructed the age- and sex-matched CODICES cohort (Comprehensive Oncobiome Analysis for Early Cancer Diagnostic Identification) including 1,030 treatment-naive plasma samples across 11 pathological subtypes of lung cancer (38.93%), 11 diverse etiologies of lung disease (21.55%), and healthy controls (39.51%). Among lung cancers, nearly two-thirds were clinical stage I (n=131, 32.59%) and stage II (n=110, 27.36%), with stage III (n=98, 24.38%) and stage IV (n=57, 14.18%) comprising a smaller subset. Smoking status (i.e., current, former, never) was known for 708 (68.73%) subjects, allowing for subanalysis within smokers with the most common risk factors for lung cancer. Samples from ethnically diverse populations, including subjects who self-identified as Black or African American, American Indian, Asian, or Pacific Islander, were also captured, accounting for 15.63% (n=161). Additionally, 428 subjects (41.55%) had radiologically based lesion sizes for both lung cancer and lung disease, with a median diameter of 2.5 centimeters (mean ± SD = 2.96 ± 1.91 cm). To complement the CODICES discovery cohort, a separate blinded validation cohort was obtained, including 106 patients with either clinical stage I cancer or lung disease.

[0171] Next, 400 μL of plasma was isolated from each subject in the CODICES cohort, followed by cell-free DNA extraction and shallow shotgun metagenomic sequencing (approximately 20 million reads per sample). Approximately one-third of the samples (n = 335) were further processed for 16S rRNA amplicon sequencing using the shorter V6 region due to the highly fragmented nature of cell-free DNA in a single sequencing run as an additional validation of the approach. To control for contamination, all sequencing runs for both shotgun metagenomic and amplicon data included positive (mock community) and negative (experimental blank) controls, in accordance with other low-biomass microbial sequencing protocols. Sample type was also randomized in all sequencing runs to prevent batch confounding. Furthermore, because multiple host-centric diagnostic approaches have described the utility of protein-based markers in lung cancer, we also measured plasma-derived concentrations of carcinoembryonic antigen (CEA) and osteopontin (OPN), covering 99.22% (n = 1022) of the CODICES cohort, where quantities permitted.

[0172] The availability of matched metagenomic, proteomic, and clinical metadata prompted the development of a multispecies (host and microbial) stack machine learning (ML) strategy (Figure 7A). Specifically, after filtering for low-variability features and centered log-ratio-transformed microbial counts, the concatenated data was fed into an ensemble classifier comprising elastic net, random forest, and gradient boosting algorithms operating in parallel using matched cross-validation (CV) folds. Scores from the three algorithms were then weighted using a logistic regression ensemble model. Model hyperparameters were optimized within an 820-step grid search using a 10-fold CV scheme. Predictions from an independent holdout cross-validation fold were saved to estimate model performance in the discovery cohort, and final clinical proteometagenomic model tuning was performed on 454 CODICES samples with all three sets of available information before application to a blinded validation cohort of 106 patients (Figure 7A).

[0173] Evaluation of known metagenomic signatures for lung cancer diagnosis in clinically low-risk settings We deployed sequencing data and ML architectures to evaluate the diagnostic performance of known metagenomic features in 407 healthy individuals versus 401 treatment-naive, age- and sex-matched, low-risk individuals with cancer across all stages of clinical trials (Figure 7B). Specifically, all 9.57 × 10 8 cell-free DNA reads (1.82% of the total) that did not map to the human reference genome were aligned against the RefSeq Release 206 database ("rep206"), which contains bacterial, fungal, and viral genomes. A total of 9.60 × 10 7 reads aligned to microbial genomes (0.18% of the total). A total of 98 extraction blanks, processed in parallel and included in each sequencing run, were then used to estimate putative contaminants. Prevalence-based filtering using these blanks identified and removed 5,630 species (46.59% of the taxa, remaining n = 6,455 taxa). These were primarily low-abundance taxa, accounting for only 0.88% of total microbial reads (Figure 11A). Because the rep206 genome is well defined, aggregate genome coverage was calculated across all microorganisms in the database using Zebra, using all samples from the CODICES cohort (Figure 11B). A new subset of microbial taxa with an aggregate genome coverage of ≥1% (n = 2172, 17.97% of taxa) was then created to compare downstream ML performance. Repeating these steps using microorganisms identified in the experimental blank revealed that many well-covered microorganisms in plasma had low coverage in the blank (Figure 11C).

[0174] We next crossed the microorganisms identified in rep206 with those previously found in lung cancer-associated bronchoalveolar lavage (BAL) samples (n = 6530 taxa) and bacterial and fungal species (n = 300 taxa) identified in two highly decontaminated pan-cancer tissue-centric studies from the Weizmann Institute of Science (WIS). Notably, when calculating Fisher's exact test for enrichment in aggregate genome coverage sets of ≥ 1%, both BAL sample features (p = 3.37 × 10-147, X2 = 667.57) and WIS features (p = 2.44 × 10-59, X2 = 263.89; Figure 11B) showed strong overlap, despite their BAL and tissue origins, respectively. This suggests that plasma-derived microbial features may compositionally reflect their tissue and even peri-tissue counterparts.

[0175] We then compared the diagnostic performance of these four microbial feature sets using our stacked ML strategy (Figures 7A and 7C). Notably, although these feature sets varied 21.8-fold in size between the largest and smallest (6,530 vs. 300 taxa), all of them provided strong and similar AUROCs (range: 90.1–92.8%) with overlapping 95% confidence intervals (Figure 7C, upper left). Subset predictions for various histological subtypes and stages further demonstrated consistently strong diagnostic performance (minimum AUROC: 87.3%, maximum: 99.0%), with the strongest performance for small cell lung cancer (SCLC, AUROCavg = 97.03%; Figure 7C, upper right). Unexpectedly, when examining predictions across all cancer subtypes at individual stages, we observed the best diagnostic performance when discriminating clinical stage I lung cancer versus non-cancer controls (stage I AUROCavg = 92.93%; Figure 7C, bottom left), followed by a slight but gradual decline in performance at higher stages (stage IV AUROCavg = 88.4%; Figure 7C, bottom left). This trend may reflect either the scarcity of samples available at higher stages in the CODICES cohort (Figure 7B) or a biological phenomenon in which late-stage cancers may sufficiently disrupt the vascular barrier to allow circulating non-cancer-associated microbial DNA. Nevertheless, even in the context of treatment-naive cancers, these analyses validate and extend our original observation that cell-free microbial DNA (cf-mbDNA) offers a novel class of cancer diagnostics, 8.6-fold greater than previously tested samples.

[0176] Because the WIS overlapping bacterial and fungal feature set had substantially fewer taxa and putative applicability to other cancer types, providing similar diagnostic performance (Figure 7C), we used 300 species for further analysis. WIS overlapping bacteria have previously been shown to be strongly associated with smokers and non-smokers, and since 82.79% (332 / 401) of our lung cancer samples had known smoking exposure, we suspected that differences in smoking history might have influenced the observed diagnostic performance. Therefore, we repeated stacked ML after subsetting healthy and CODICES subjects with cancer to those with known smoking history. Surprisingly, diagnostic performance was found to be higher than previously among smokers alone (AUROC = 93.0%; Figure 26A). This suggests that smoking-related pathology potentially allows further leakage of microbial DNA into the circulation, generating a stronger signal for diagnosis in the setting of cancer.

[0177] Classical metagenomic analysis of plasma-derived microbiomes Similarly, although lung tumor tissue has lower alpha diversity compared to adjacent normal tissue within the same subject, the presence of damaged tissue and concomitant vascular barriers between patients with lung cancer led to the hypothesis that alpha diversity in plasma from cases would be higher than in controls. Indeed, after filtering for WIS overlapping markers, significantly higher alpha diversity was observed between lung cancer plasma samples compared to healthy controls (Figure 7D). The strength of separation was confirmed by calculating robust Aitchison beta diversity distances. This was identified with stacked ML, with a pseudo-F PERMANOVA statistic of 115.92 (p = 0.001; Figure 7E). Microbial differential abundance modeling with ANCOM-BC then indicated that lung cancer overwhelmed health-related biomarkers, particularly the two WIS-overlapping bacteria Pseudomonas oleovorans and Pseudomonas mendocina, provided strong separation (Figure 7F, circled red dots). Notably, both of these bacteria had relatively higher abundance in lung cancer samples than in either healthy samples or blanks (Figure 7G, top), and had substantially higher aggregate genome coverage in plasma samples compared to blanks, with P. oleovorans showing 55.24% coverage in plasma but only 0.42% coverage in blanks. Thus, classical metagenomic analysis of cancer-associated microbial biomarkers in plasma confirms the diagnostic ability of cf-mbDNA in large cohorts of treatment-naive individuals.

[0178] Next, we explored whether the CEA and OPN proteomic markers could enhance the diagnostic performance of metagenomics. After examining their increasing concentrations at each stage, OPN provided better performance than CEA in the initial comparison (Figure 7H), and we found that the combined proteogenomic classifier increased cancer discrimination compared to that initially observed when using a feature set 21.5 times larger than all putatively decontaminated rep206 features (AUROC = 92.1%; Figure 7C, top left, Figure 7I). Furthermore, while protein features individually provided only moderate cancer detection with the same stacked ML approach (Figure 26B), their performance improved at higher stages, counteracting the opposite trend of our microbial classifier. This improvement was also observable when assessing sensitivity at 99% specificity (Figure 7J). This demonstrated state-of-the-art sensitivity (approximately 50% at 99% specificity) for detecting stage I lung cancer compared with existing methods that evaluated epigenetic panels or the combination of ctDNA and proteomic marker panels.

[0179] To further verify that this diagnostic performance was not the result of metagenomic read misalignment, we then amplified the short V6 region of 16S rRNA in a subset of 335 plasma samples processed in a single sequencing run using 16 extraction blanks and four positive control mock assemblages, including samples from 142 lung cancer patients and 96 healthy subjects. After processing and decontaminating the sequenced amplicons with Deblur for taxonomic composition and identifying them with PICRUSt2, their abundances were input into the stackML pipeline. Notably, both amplicon-based taxonomic composition and functional pathways strongly discriminated between lung cancer and healthy controls (AUROC range: 83.4–86.3%; Figure 7K) and, when combined with proteomic information, provided diagnostic performance similar to shotgun metagenomics (AUROC: 93.6%). This simultaneously validates the shotgun metagenomic findings and further demonstrates (i) the existence and utility of plasma-derived amplicon signatures for cancer diagnosis, and (ii) the applicability of functional microbial biomarkers for cancer diagnosis. A proteoamplicon-based strategy would also be highly cost-effective, utilize minimal amounts of plasma, and provide robust diagnostic performance.

[0180] Discrimination of lung cancer against risk-matched disease controls Nevertheless, lung cancer screening, except perhaps for multiple cancer screening, is currently rare in low-risk settings. Therefore, we next evaluated the diagnostic performance of lung cancer screening in 222 age-, sex-, and smoking-history-matched adults (n = 161 smokers, or 72.5%) with diverse etiologies of lung disease (LD) compared with the same lung cancer (LC) patients, as recommended by other studies (Figure 8A). First, we again verified that plasma-derived CEA and OPN were significantly increased in a stage-specific manner in LC compared with LD (Figure 8B). This was followed by classical metagenomic analysis of WIS-overlapping bacterial and fungal signatures (Figure 8C-8F). Notably, the separation of LC and LD samples was statistically significant, but weaker, compared with LC vs. healthy samples for all comparisons: alpha diversity was more similar (Figure 8C), the robust Aitchison distance showed a weaker pseudo-F PERMANOVA statistic (F = 16.80, p = 0.001; Figure 8D), and fewer differential abundance features were observed (Figure 8E). This change was also evident in lung cancer-associated P. oleovorans and P. mendocina, where LD relative abundance increased in the direction of LC samples compared with healthy controls (Figures 8E-8F).

[0181] Similarly, stacked ML of WIS overlapping microbial LC vs. LD provided worse discrimination than its LC vs. healthy counterparts, even when CEA and OPN were added to the analysis (AUROC = 76.2%; Figure 8G). Furthermore, when predictions were subset across histological subtypes and stages, stage IV detection was the only comparison that showed similarity to LC vs. healthy samples (Figure 8G, bottom right; Figure 7C, bottom right). Repeating the amplicon-based 16S rRNA approach between the same 142 LC and 97 new LD samples reproduced the poor performance (AUROC All = 70.8%; Figure 26C). These analyses suggested the limitations of using known cancer-associated microbial biomarkers, especially in the context of risk-matched controls, which are rarely used in large-scale cancer or cancer microbiome studies. They also suggest that proteoamplicon-based strategies may not provide sufficient diagnostic performance unless additional features can be added to the test.

[0182] Because WIS overlapping bacteria and fungi were identified outside the context of risk-matched controls and plasma samples known to contain metagenomic features underrepresented in existing microbial databases were not further examined, we next performed de novo metagenomic assembly using all blood and primary tumor whole-genome sequencing (WGS) samples from The Cancer Genome Atlas (TCGA) along with samples from the CODICES cohort (Figure 8H). Specifically, we separated LD, healthy, and blank samples into distinct groups, and then performed metagenomic co-assembly using metaSPAdes while overlaying blood samples from all 25 available TCGA cancer types (n = 1937 samples) with matching cancer types in the plasma-derived CODICES cohort. In parallel, we also performed metagenomic co-assembly for 24 TCGA primary tumor types (n = 2106 samples). Collectively, these analyses identified 27,819 contigs >1500 bp in length, which were subsequently pooled and binned using a deep variational autoencoder. These 4605 metagenomic bins were quality-controlled with CheckM and then overlaid at the predicted strain level using PhyloPhlAn3 and a large database of human-related species-level genome bins (SGBs), resulting in 1562 final pan-cancer, pan-tissue, and blood metagenomic bins (Figure 8H).

[0183] Notably, metagenomic bins outperformed WIS overlapping cancer-associated features in discriminating between LC and LD samples in all comparisons except for stage IV disease (Figure 8G, orange vs. blue lines). Combining metagenomic bins with CEA and OPN further synergistically increased diagnostic performance (Figure 8G, red lines). Reapplication of de novo metagenomics paired with CEA and OPN to low-risk analysis revealed similar performance, particularly in early-stage disease (stage I AUROC: 90.2%; Figure 27A), but the improved performance was most pronounced in the high-risk setting. In addition to better AUROC values, the shape of the bin-based ROC curves in clinical stage I and II disease (Figure 8G, bottom left and bottom center) suggested higher specificity at high fixed sensitivity cutoffs (e.g., 97% sensitivity), suggesting that metagenomic bins may be particularly useful for ruling out malignant disease while mitigating false positives.

[0184] Evaluation of de novo assembled metagenomic bins in TCGA After creating a novel set of metagenomic assembly bins with good diagnostic performance for identifying LD in the CODICES cohort, we validated their cancer type specificity and generalizability in two independent cohorts containing 16,049 samples across 35 conditions (34 cancer types and healthy) with samples independent of the metagenomic assembly.

[0185] First, we realigned all 15,512 base-quality filtered, double-host depleted, length-limited (>45 bp), whole genome sequencing (WGS) and RNA sequencing (RNA-Seq) samples in TCGA to metagenomic bins. A total of 14,809 samples (95.47%) had non-zero alignments to 1545 of the bins, and normal adjacent tissue (NAT) samples were not part of the metagenomic assembly, accounting for 73.26% (515 / 703) of the dropout samples. Notably, the pairwise alignment rate of non-human reads across all cancers, sample types, and experimental strategies increased by a median of 891-fold to bins compared to RefSeq (Wilcoxon signed-rank test: p ≤ 2.23 × 10 ). -308 (Figure 32B). To confirm that the higher alignment rate was not caused by mapping more contaminants, and because TCGA did not include sequencing blanks, we mapped non-human reads derived from sequencing 98 reagent-only blanks collected across all CODICES and NYU cohort sequencing runs to RefSeq and bins, and found that bins significantly reduced contaminant mapping (Wilcoxon signed-rank test: p=8.45×10 -18 ; Figure 32C). Importantly, bins provided significantly higher non-human mapping rates in TCGA samples than in blanks (Wilcoxon: p = 6.63 x 10 -64 ), 99.8% of TCGA samples exceeded the median mapping rate to blank bins (0.26%). These observations were consistent when TCGA samples were stratified by WGS or RNA-Seq. Thus, compared to RefSeq, bins dramatically increased biological microbial mapping rates while mitigating reagent-based contaminant mapping, despite using 7.65-fold fewer genomic features.

[0186] The bin alignment rate for our experimental strategy revealed a significant increase, even though RNA-Seq samples were excluded from the metagenomic assembly (Wilcoxon signed-rank test: p ≤ 2.23 × 10-308; Figure 32A). Furthermore, bins were consistent with the alignment rates for WGS (95% CI: [8.25, 8.90]%) and RNA-Seq (95% CI: [8.42, 8.61]%), which differed for RefSeq (WGS, 95% CI: [1.66, 1.88]%; RNA-Seq, 95% CI: [0.32, 0.39]%). Calculation of the average fold change between bins and RefSeq by cancer type revealed an average 1429-fold increase in non-human mapping rates (Figure 32D), with lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC) increasing 1372-fold and 2180-fold, respectively.

[0187] Quality control of metadata (e.g., removing sequencing centers with a single cancer type) and subsetting to a single sequencing platform, which accounted for 98.00% of the samples, left 14,454 tissue and blood samples (93.18% of the total) for downstream analysis.

[0188] Although batch effects still persisted in TCGA metagenomic bin abundances, they were mitigated (Figure 12A-B). For example, our original principal variance component analysis (PVCA) of raw Kraken-derived genera abundances in TCGA revealed that approximately 34% and 36% of the data variance was attributable to sequencing center and experimental strategy, respectively, which were reduced to 20.3% and 3.2% of the variance in the metagenomic bin data, respectively. Furthermore, when raw WGS and RNA-Seq samples were analyzed separately, disease-type-related variance among WGS samples (those used for metagenomic assembly) exceeded the variance due to other technical variables without batch correction (Figure 12C and D). This suggests that metagenomic assembly may be able to sufficiently mitigate technical batch effects in retrospectively analyzed cancer microbiome datasets.

[0189] Nevertheless, batch calibration was first performed to utilize the full TCGA cohort, and 10-fold cross-validated ML analyses were performed using gradient boosting in one cancer versus all other modes to test cancer type specificity (Figure 13A-C). This analysis revealed pan-cancer discrimination among 32 primary tumor types (AUROCavg = 84.31%, 95% CI: [83.28, 85.34]%), moderate discrimination between tumors and NAT (AUROCavg = 73.14%, 95% CI: [70.56, 75.71]%) (potentially due to NAT sample dropout), and very strong performance in discriminating between blood samples from 20 cancer types (AUROCavg = 96.21%, 95% CI: [95.62, 96.80]%). Importantly, using metagenomic bins, blood samples from lung adenocarcinoma had the highest AUROC and AUPR combined performance in TCGA (AUROCavg=97.92; AUPRavg=89.99%), which was not the case for either our bacteria- or fungi-centric analyses, suggesting that co-assembly with samples from the CODICES cohort improved concomitant cancer-type specificity.

[0190] All ML analyses using scrambled metadata or shuffled count data on matched CV partitions as negative controls, followed by comparison of expected and actual performance, found statistically significantly better cancer type discrimination than the null model in all cases (Figures 14A-B and 15A). Furthermore, when subsetting to all TCGA blood samples in patients with early-stage (Ia-IIc) cancers and then repeating ML, we found strong pan-cancer discrimination (AUROCavg = 95.23%, 95% CI: [94.05, 96.42]%), particularly including lung adenocarcinoma (AUROCavg = 98.31; AUPravg = 94.90e%; Figure 15B). Similar performance in early-stage and late-stage cancers suggests that it may be difficult to distinguish stage I tumors from stage IV tumors using metagenomic bins, and this was indeed the case (Figure 15C).

[0191] Next, to completely avoid batch correction, we subset all raw metagenomic bin abundances into individual sequencing centers, experimental strategies (WGS or RNA-Seq), and sequencing platforms (Illumina HiSeq) before evaluating 10-fold ML. This provided strong tumor-type discrimination in the Harvard Medical School (HMS) samples (Figure 9A, AUROCavg = 98.27%, AUPRavg = 89.56%) and all other sequencing centers (Figures 16A-16F). Repeating the analysis using the raw data against scrambled metadata and shuffled count data again revealed significantly better pan-cancer, tumor-type discrimination than null in all subsets (Figure 9B; Figures 17A-17F). Applying classical metagenomic analysis further revealed significant cancer-type-specific alpha and beta diversity distributions in all testable primary tumor subsets, with cancer type often explaining more than half of the data variance in metagenomic bin abundances (Figure 9C; Figures 20A-20E; and Figures 21A-21E). To further demonstrate the cancer-type specificity of bins, differential abundance modeling using ANCOM-BC was applied to all raw data subsets, and differential abundance bins were again found in all comparisons (Figures 9D and 9H; representative examples in Figures 22A-22G).

[0192] Repeating the tumor vs. normal ML analysis using raw metagenomic bin subsets confirmed 9 of the 12 tumor types, yet the aggregated comparison still significantly outperformed the null model (Figures 18A-18D).

[0193] We then repeated all blood-based cancer type analyses using one-versus-all-others ML (Figure 9E; Figures 19A, 19C, 19E, and 19G), negative control ML analysis (Figure 9F; and Figures 19B, 19D, 19F, and 19H), alpha diversity (Figures 23A-23E), beta diversity (Figures 24A-24E), and differential abundance modeling (Figures 25A-25E). In all of these analyses, metagenomic bin-based cancer type variation among TCGA blood samples was significant, with the exception of alpha diversity only in a subset from a single sequencing center (Baylor College of Medicine, Figure 23C). Furthermore, two sequencing centers (HMS, Broad Institute) that included blood samples from lung adenocarcinoma consistently demonstrated the strongest diagnostic performance among all other cancer types (Figure 9E; and Figure 19G), again suggesting that CODICES cohort coassembly improved lung cancer detection.

[0194] Additionally, because one-versus-all-others comparisons can inflate the no-information rate (NIR), we performed multiclass ML, in which all cancer types are considered simultaneously using a gradient boosting algorithm. Applying this to all 32 primary tumor types in the batch-corrected data reproduced the one-versus-all-others performance (average pairwise AUROC = 83.19%; Figure 13D), with a mean balanced accuracy of 65.17% compared to the NIR of 10.48% (p < 2.2 × 10-308). Blood-based multiclass ML across 24 cancer types using batch-corrected data performed even better (average pairwise AUROC = 95.50%; Figure 9I), with a mean balanced accuracy of 79.26% compared to the NIR of 8.92% (p < 2.2 × 10-308). Because all TCGA blood samples were WGS, and raw WGS samples had less variation across sequencing centers than disease types using PVCA, we also tested multiclass blood ML using raw metagenomic bin abundances with nearly identical results (average pairwise AUROC = 95.14%, average balanced accuracy = 79.02%; Figure 13E). Overall, metagenomic bins are cancer type-specific.

[0195] Evaluation of de novo assembled metagenomic bins as a database of cancer-associated microbial signatures from non-blood-borne microbial reads: Fecal microbiome data Having demonstrated that metagenomic bins are cancer-type specific when tested against TCGA tissue and blood samples, we analyzed the bins to determine whether they could serve as a database of cancer-associated microbial genomes to which sequencing reads from non-blood sources could be aligned. To this end, we hypothesized that the bins could provide diagnostic utility for colorectal cancer (CRC). Geographically distinct fecal metagenomic CRC cohorts from France (FR) and China (CN) (PMIDs: 25432777, 26408641) were then processed and cross-compared in a subsequent meta-analysis (PMID: 30936547), providing insight into internal cross-validation and external cohort validation performance (Figure 33A).

[0196] Beta diversity revealed significant presence-absence differences in the FR and CN cohorts between subjects with CRC and healthy subjects (Jaccard PERMANOVAs: p = 0.001; Figure 33B). However, as reported, the effect sizes varied for Aitchison-based measures, and there was no significant change in Shannon alpha diversity (PMIDs: 25432777, 26408641). Nevertheless, applying LOOCV ML to each cohort (Figure 33C) revealed higher AUROCs (FR: 88.77%, CN: 85.64%) than those published by the original authors (FR38: 84%; CN37: 83.61%) or the meta-analysis (PMID: 30936547; FR: 85%; CN: 81%). Notably, cross-application of these models without batch correction revealed improved performance over LOOCV (FR-CN: 88.73%; CN-FR: 90.75%; Figure 33C), surpassing an international meta-analysis with an AUC of up to 9%. Using an internal LOOCV and fixing the cutoff yielded specificity and sensitivity of up to 82.0% and 79.2%, respectively (Figure 33C). Because the FR cohort provided stage information, we subsetted the LOOCV and CN cross-tested predictions to early (stages I-II) and late (stages III-IV) samples and found equally strong early-stage performance (92.03-92.62%; Figure 33D). Calculating bootstrapped sensitivity at 92% specificity (PMID: 25432777) showed a ≥15-point increase and consistency in cross-testing between LOOCV and CN (Figure 33E, left). The CN cohort had lower sensitivity at 92% specificity (Figure 33E, right), but its LOOCV and FR cross-test values ​​were similar. Thus, binning applies across diverse sample types (i.e., tissue, blood, and stool), independent cancer types, and geographically diverse cohorts, improving diagnostics.

[0197] Fecal-derived microbiota changes may also be useful for distant lung cancer diagnosis. In the absence of publicly available fecal shotgun data for lung cancer versus controls, we examined the 16S rRNA gene amplicon from Lim et al. (PMID: 34586729) to determine whether it was compatible with metagenomic bins ("Lung-Gut" cohort, Figure 33F). Despite limited 16S rRNA alignment, significant presence / absence (Figure 33G) and Aitchison-based beta diversity differences were present, and previously unreported Shannon alpha diversity was significantly reduced. We then performed LOOCV ML, matching the method to the CRC cohort, and found an AUROC of 86.52% (Figure 33H), or 10% higher than reported (PMID: 34586729). Notably, the sensitivity of 51.6% at 92% specificity approached the CRC results (Figure 33I) and was comparable to the LOOCV results of the CN cohort (Figure 33E, right). Unfortunately, healthy saliva samples from the same cohort were not publicly available for comparison (PMID: 34586729). Nevertheless, metagenomic binning is broadly generalizable for improving the diagnostic performance of multiple cancers, and the data herein demonstrate that aligning sequencing reads (shotgun metagenomic reads and 16S targeted amplicon sequencing) from fecal samples to our pan-cancer metagenomic bins enables the development of colorectal cancer- and lung cancer-specific diagnostic classifiers.

[0198] Complementarity of metagenomic bins with fragmentomics information Human cell-free DNA (cfDNA) is commonly detected alongside metagenomic cell-free DNA, but their diagnostic suitability for cancer detection remains unclear. Therefore, two cfDNA cohorts (PMID: 25646427; 31142840, Figure 29A) were investigated for nine DNA fragment terminal nucleotide frequencies (NTFs; PMID: 36630480) independent of metagenomic assembly and for synergy with metagenomic binning.

[0199] Patients with HCC had significantly different alpha and beta diversity (Figure 29B). Consistent with the methods in the CRC and lung-gut cohorts, subsequent LOOCV ML revealed strong bin-based discrimination (AUROC = 91.74%; Figure 29C, "Bins"), with a synergistic increase in NTFs (AUROC = 96.77%; Figure 29C, "Bins + NTFs"), resulting in an approximately 15%-65.6% increase in median bootstrapped sensitivity at 99% specificity (Figure 29D).

[0200] We then analyzed whole-genome sequencing (WGS) multiplexed cancer plasma samples from Cristiano et al. (hereafter, "Hopkins") (Figure 29E) and found that cancer samples varied significantly in beta and alpha diversity (Figure 29F). Ten-fold cross-validation (CV) for each cancer type using bin abundance versus healthy controls showed that the addition of NTFs improved performance in six of seven cancer types and sensitivity in five of seven cancer types (Figure 29G-J). Aggregating all cancers versus non-cancer controls revealed significant improvement in multiplexed cancer detection combining NTF and bin abundance for all stages, with an overall AUROC of 96.7% (Delong: p ≤ 4.11 × 10-8; Figure 29K-N). In summary, across at least eight cancer types, bins are diagnostically useful and broadly synergistic with human fragmentomics features.

[0201] Evaluation of de novo assembled metagenomic bins in an independent cohort Despite these findings, we also evaluated the diagnostic performance of bins in a plasma dataset (i) independently of the metagenomic assembly and (ii) utilizing non-cancer controls. All 491 base quality filtered, double-host depleted, length-limited (>45 bp), shallow whole-genome sequenced (WGS) plasma samples from Cristiano were realigned. ML comparisons were then performed using raw metagenomic bin abundances for all cancer types versus 260 healthy controls, finding superior performance to null for each of them, including lung cancer (Figure 9J). ML analysis of scrambled and shuffled negative controls showed the expected null results (Figure 9K). A one-versus-all-other cancer type comparison also confirmed superior classification to null for seven of the eight cancer types, with only cholangiocarcinoma (n = 25) failing to reach a sufficient AUPR (Figure 9F).

[0202] Calculation of robust Aitchison beta diversity distances further revealed a clear separation along axis 1 between healthy and pan-cancer samples (Figure 9L; PERMANOVA: F = 4.43, p = 0.001). Subsequent ML comparisons between pan-cancer or grouped stages (Figure 9M) intriguingly reproduced the pattern initially observed in our CODICES cohort data (Figure 27A), where metagenomic bins provided nearly similar diagnostic performance (AUROC range: 86.35–92.86%) among clinical stage I–III samples, followed by the lowest performance in clinical stage IV samples (AUROC 95% CI: [76.16, 79.9]). The mechanism driving this pattern is unclear; it could be artifactually attributable to the lower number of stage IV samples in both the Hopkins (n ​​= 22 samples) and CODICES cohorts, or biologically, it could indicate the translocation of non-cancer-associated microorganisms due to a significant defect in the vascular barrier. Nevertheless, all of these analyses strongly support the cancer type specificity and diagnostic utility of these novel metagenomic bins in several independent cohorts.

[0203] Construction and preliminary evaluation of a multi-omics multi-species test for lung cancer screening To integrate heterogeneous multi-omics data into a single test, we designed a stacked ML strategy (Figure 7A). If positive, we designed a blood test that would trigger a confirmatory follow-up LDCT or PET-CT imaging test (Figure 10A, bottom panel "2").

[0204] Stacked ML training was applied to all LC and healthy samples using bin abundances with and without human information (i.e., proteins and / or NTFs) (overall AUROC = 94.1%; Figure 30A). The combined feature set significantly improved performance over the individual feature sets (Delong's test: q ≤ 1.74 × 10-9). Bins and NTFs alone performed poorer in Stage IV, and 100 iterations of ML with stratified random sampling ruled out sample number bias, but repeating this with the combined feature set yielded the expected trend. Bootstrap sensitivity at 99% specificity demonstrated Stage I detection approaching the median sensitivity of 60% (Figure 30B).

[0205] To confirm that the cross-validated results were not affected by overfitting, these analyses were combined with protein and NTF features (14 features total) using a robust Aitchison PCA-projection (RPCA; PMID:PMID) to convert bin abundances into just three principal component vectors, and similar performance was found (AUROC = 92.3%; Figure 30C). Additionally, 142 LC and 96 healthy subjects were processed in parallel for plasma-derived 16S rRNA amplicons (targeted microbial amplicon sequencing), and the results were independently replicated (AUROC = 93.6%; Figure 30D).

[0206] Because lung cancer screening, by definition, requires asymptomatic subjects with a defined smoking history and age, we subset all screen-eligible LC-bearing patients (n = 95) and age-matched healthy smokers (n = 46) (cases: 67.64 ± 6.94 years; controls: 66.37 ± 6.58 years) as an internal validation cohort. After retraining the stacked ML model on all other LC-bearing and healthy subjects and deriving a 99% specificity-related cutoff, the retrained stacked ML model was applied to the retained internal validation cohort (Figure 30E). For the combined model, the AUROC decreased, but the specificity and sensitivity were similar (specificity: 96.2-100%, sensitivity: 54.7-56.8%), with the bin-only model performing best (AUROC = 91.52%). Subsetting these predictions to stage I-only screen-eligible LC patients, while applying the same cutoff, revealed a consistent sensitivity of 68.2% with an observed specificity of 96.2-100% (Figure 30F). Collectively, these data demonstrate that the combination of metagenomic (human, NTF, and microbial, bin) data with plasma proteins can yield a diagnostic classifier capable of discriminating between subjects with lung cancer and healthy subjects.

[0207] Development and validation of clinical proteometagenomic diagnostics Metagenomic binning significantly improved LD discrimination from LC over WIS overlap biomarkers, but did not provide state-of-the-art diagnostic performance alone or in combination with proteins (Figure 8G). The proteometagenomic classifier was then improved by adding clinically collected and radiography-related metadata. Notably, this diagnostic application could be performed after low-dose computed tomography (LDCT) in high-risk clinical settings, maximizing test sensitivity while eliminating the need to biopsy presumptively nonmalignant nodules (Figure 10A, top). In contrast, the aforementioned low-risk diagnostic relying solely on metagenomic or proteogenomic markers is applied to relatively healthy populations while maximizing specificity, as has been done by others to mitigate false positives (Figure 10A, bottom). These represent two distinct approaches, both of which could be enhanced by metagenomic information.

[0208] Over 400 LC and LD samples in the CODICES cohort had matching clinical metadata, most of which included lesion diameter, shape (e.g., spiculation), solidity, location (e.g., upper lung), and nodule clinical risk scores such as Brock probability (Figure 10B). Notably, the median lesion diameter in this cohort was only 2.5 cm (mean ± SD = 2.95 ± 1.91 cm, n = 431). Integrating these clinical variables with metagenomic bin abundance and CEA and OPN concentrations via stacked ML for all 454 samples significantly increased the discrimination of malignant nodules (AUROC = 90.0%), resulting in a greater accuracy than comparable models built using matching clinical risk scores (Brock, pCA) or PET-CT-derived standardized uptake values ​​(SUV) (AUROC range: 68.9–75.0%; Figure 10C, left). Optionally, adding PET-CT SUV to this clinical proteometagenomics diagnostic synergistically improved diagnostic performance (AUROC = 91.8%). Notably, subsetting these predictions to lesions with a known diameter ≤ 3 cm and a low clinical risk score for malignancy (pCA ≤ 50%) improved the performance of the diagnostic model, but not the performance of equivalent models built using Brock, pCA, or PET-CT SUV (Figure 10C, right).

[0209] The improvement in proteogenomic diagnostic performance with the addition of clinical metadata suggested that proteoamplicon-based strategies may be similarly beneficial. Therefore, we repeated the original LC (n = 142) vs. LD (n = 97) proteoamplicon analysis (Figure 26C) using additional clinical information and found synergistic improvements in performance for both taxonomic organization (AUROC = 91.8%) and functional pathway abundance (AUROC = 91.3%). This surpassed the clinical risk score and PET-CT SUV (Figure 10D). Notably, the shape of these ROC curves (Figures 10C-10D) suggested that metagenomic features could support highly sensitive testing (e.g., 97% sensitivity) while still maintaining sufficient specificity.

[0210] Encouraged by these improvements and intrigued by the cancer type specificity of metagenomic bins in the Hopkins and TCGA cohorts, including between lung cancer histological types, we hypothesized that our diagnostic model might be able to distinguish between lung cancer subtypes in sufficiently large cohorts. If feasible, blood-based discrimination between these two subtypes could provide valuable guidance for clinical management, as their histologies affect therapy differently (e.g., pemetrexed is effective in adenocarcinoma but not squamous cell carcinoma). Because 293 of 454 samples (64.5%) were derived from either lung adenocarcinoma (LUAD) or lung squamous cell carcinoma (LUSC) (Figure 10B), we constructed a LUAD vs. LUSC classifier using the same clinical proteometagenomic features as the diagnostic model (Figure 10E). Notably, despite the small nodule size (median = 2.9 cm, mean = 3.4 cm), the diagnostic model demonstrated strong discrimination (AUROC = 91.5%). However, a comparable model built using only PET-CT SUVs performed significantly worse (AUROC = 68.6%) than a comparable model using only clinical risk scores or protein information, demonstrating pseudorandom performance (AUROC range: 47.9%–59.7%). Such findings suggest that plasma-centric metagenomic features may be able to guide minimally invasive histological subtype determination. Comparing the performance of lung cancer vs. lung disease classifiers with and without NTF information revealed that, in this particular classification scenario, the addition of NTFs did not lead to classifier improvement (Figure 31).

[0211] Application of diagnostic models to a blinded validation cohort After thoroughly exploring the CODICES discovery cohort, we trained a final clinical proteometagenomic diagnostic classifier and stage using all 454 LC and LD samples across various histological types. In addition to metagenomic bins, CEA, and OPN, this classifier included clinical risk scores (Brock, pCA), lesion diameter, lesion spiculation, lesion solidity, upper pulmonary nodule status, emphysema status, and smoking history, when available. For comparability across subsequent samples, we also built equivalent stack models using only clinical risk scores or PET-CT SUV.

[0212] We then acquired a blinded validation cohort from collaborators, including 106 plasma samples from patients with an unknown number of clinical stage I lung cancers and lung diseases of diverse etiologies. DNA was extracted from 400 μL of plasma for independent sequencing runs, including standard positive (mock cluster) and negative blank controls, and non-human reads were aligned to metagenomic bins. Metagenomic and proteomic features were then normalized for inter-run variability, followed by feature standardization to match diagnostic model inputs, including clinical metadata (when available). Notably, the average lesion diameter in this cohort was only 1.91 cm, with multiple lesions as small as 6 and 8 mm (Figure 10F).

[0213] The final diagnostic model and an equivalent model using PET-CT information were then applied to the blinded cohort data to generate predictions in a one-way blinded fashion. Notably, the diagnostic model significantly outperformed the gold-standard PET-CT information (AUROC: 79.1% vs. 64.5%; Delong's test: p = 0.019). Given the shape of the ROC curves from 454 samples in the CODICES cohort (Figure 10C) and the blinded validation cohort (Figure 10G), it was further hypothesized that the model test should provide significantly better specificity at a fixed 97% sensitivity. Indeed, this was also the case when McNemar's test was used to compare model predictions and SUV at a fixed 97% sensitivity, as well as when using Brock and pCA (p < 0.001 for all tests), which had even poorer AUROCs (Brock: AUROC = 0.724; pCA: AUROC = 0.738). Additionally, the sensitivity comparison when fixed at 80% specificity was significantly better than the PET-CT SUV (p = 0.00112) and appeared to match or even outperform the fragmentomics approach among stage I lung cancer. Thus, the model was validated in an independent, blinded validation cohort in clinical stage I lung cancer versus the most challenging conditions of pulmonary disease, while providing state-of-the-art diagnostic performance.

[0214] As described in Example 1, we characterized blood and tissue metagenomes in 17,520 samples from four independent cohorts spanning 34 cancer types, as well as healthy controls and 11 lung diseases of diverse etiologies. Starting with known cancer-associated bacterial and fungal biomarkers, we demonstrated that just 300 plasma-derived microbial biomarkers could provide state-of-the-art lung cancer detection (approximately 50% sensitivity at 99% specificity) in a low-risk screening setting of 808 newly collected treatment-naive individuals across all stages and histological subtypes (Figure 7C). Furthermore, we found that this pared-down set of microbial biomarkers could be synergistically complemented by the orthogonal, host-centric plasma proteomic markers CEA and OPN. Performance outweighs simplicity: BAL-associated metagenomic signatures were also found to provide up to 70% sensitivity at 99% specificity among clinical stage I disease (Figure 7J). Of note, these diagnostic models were constructed independently of ctDNA, fragmentation markers, and epigenetic markers, so the performance presented here likely represents a lower bound on the performance of metagenomics-enhanced cancer diagnostics.

[0215] Notably, both WIS- and BAL-associated biomarkers derived from tissue- or tissue-permeable samples, respectively, were significantly enriched in microorganisms with an aggregate genome coverage of ≥1% in the plasma-derived CODICES cohort (p<1 × 10-58 for both, Fisher's exact test). These data suggest that a significant portion of the plasma metagenome is derived from tumor tissue. The fact that the diagnostic model improved when restricted to subjects with a positive smoking history supports the hypothesis that chronic tissue damage may increase the tissue-derived representation of the concomitant metagenome in plasma. This theory also bears parallels in the gastrointestinal environment of colon cancer, where breakdown of the gut-vascular barrier allows bacterial translocation to the liver, impacting colonic metastasis.

[0216] To address potential concerns regarding mismapping of metagenomic reads to human DNA, we also evaluated the presence and utility of plasma-derived amplicon markers (Figure 7K). Indeed, both taxonomic organization and functional pathways inferred from 16S rRNA data surprisingly demonstrated diagnostic performance similar to shotgun methods, even performing equally well when combined with CEA and OPN (Figure 7K). Furthermore, when shotgun-based diagnostic performance decreased in the LC vs. LD comparison, amplicon-based performance also decreased (Figure 26C), and both similarly increased in the context of additional clinical proteomic features (Figure 10D). Thus, plasma-derived amplicon-based methods may offer a cost-, time-, and quantity-efficient alternative to shotgun-based metagenomic assessment, warranting further investigation, particularly in lung cancer.

[0217] The observed decline in diagnostic performance between LC and age-, sex-, and risk-matched LD controls (Figure 8G) further emphasized the importance of including these sample types in cell-free DNA screening studies to characterize true diagnostic performance in clinically realistic scenarios. This challenge also motivated a metagenomic (simultaneous) assembly between 5,187 whole-genome-sequenced cancer-related blood and tissue samples. Realignment of identified metagenomic bins to 17,185 host-depleted samples, followed by comprehensive statistical and ML analyses, fully demonstrated their cancer-type specificity, diagnostic power, ability to mitigate batch effects in retrospectively collected data, and potential for noninvasively distinguishing histological subtypes of lung cancer in small nodules. The broad generalizability of this metagenomic approach to other retrospectively collected cancer genomics cohorts may be extremely useful for characterizing the functional repertoire of cancer-associated microorganisms.

[0218] Through the development of a multi-omics, multispecies (host and microbial), integrated diagnostic model, we found that a metagenomics-augmented approach could provide superior sensitivity in early-stage disease compared with existing clinical risk scores and PET-CT image intensities (Figure 10C). Specifically, the validation and state-of-the-art performance of the diagnostic model on a highly challenging blinded validation cohort of clinical stage I cancer versus diverse lung diseases (Figures 10F-10H) highlight the broad utility that plasma metagenomics may have for future cancer diagnostics.

[0219] Our study has several caveats. Despite the large sample size of the new retrospective data, all plasma-derived microbial information is inherently low in abundance and difficult to sequence at high coverage. Attempts to calculate aggregate coverage can provide a surrogate for which microorganisms may be present in our cohort, but they do not address the microbial undersampling that occurs in the available data. Similarly, high aggregate coverage cannot distinguish between contaminants and non-contaminants.

[0220] Although we utilized multiple methods to control for possible contamination, including numerous extraction blanks and positive controls, we cannot rule out all false-positive results. Nevertheless, independent replication of diagnostic performance with (i) separately processed, sequenced, and decontaminated amplicon-based approaches, (ii) independently validated subsets of biomarkers (WIS and BAL), (iii) orthogonally developed metagenomic bins, (iv) separate datasets including those extracted from metagenomic assemblies, and (v) blinded validation cohorts all collectively suggest that the extent or rate of contamination does not preclude generalizable, clinically useful conclusions.

[0221] This comprehensive analysis of plasma metagenomes for the early detection of lung cancer provides a generalizable approach to develop or enhance cancer diagnostics using metagenomic information to benefit patients worldwide.

Claims

1. A method for determining a disease in a subject, comprising: receiving a biological sample, electronic medical record information, and one or more radiological images of a subject; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and determining disease in the subject as an output of a predictive model when data derived from one or more nucleic acid molecular sequencing reads, electronic medical record information, and one or more radiological images of the subject are provided as input to the predictive model.

2. 2. The method of claim 1, wherein the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input.

3. 10. The method of claim 1, further comprising identifying one or more protein biomarkers from the biological sample of the subject.

4. 3. The method of claim 2, wherein the predictive model is provided with the one or more protein biomarkers from the biological sample of the subject.

5. 5. The method of claim 4, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.

6. 10. The method of claim 1, wherein the disease comprises cancer or a non-cancerous disease.

7. 10. The method of claim 1, wherein the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof.

8. 10. The method of claim 1, wherein the one or more radiological images comprise images of an X-ray, a computed tomography (CT), a low-dose computed tomography, a magnetic resonance imaging (MRI), an ultrasound, a positron emission tomography, a fluoroscopy, an angiogram, or any combination thereof.

9. 7. The method of claim 6, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.

10. 10. The method of claim 1, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.

11. 11. The method of claim 10, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.

12. 2. The method of claim 1, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.

13. 8. The method of claim 7, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

14. 7. The method of claim 6, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.

15. 7. The method of claim 6, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.

16. 9. The method of claim 8, further comprising calculating one or more features of the one or more radiological images, the one or more features of the one or more radiological images being provided as input to the predictive model.

17. 17. The method of claim 16, wherein the one or more features comprise a Brock cancer probability score, a diameter of the lesion, spiculation of the lesion, solidity of the lesion, or any combination thereof.

18. 10. The method of claim 1, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads.

19. 19. The method of claim 18, wherein the genome database comprises a human genome database.

20. 20. The method of claim 18, wherein the genome database comprises a microbial genome database.

21. 21. The method of claim 20, wherein the microbial genome database comprises a de novo metagenomic assembly.

22. 22. The method of claim 21 , wherein the de novo metagenomic assembly is derived from biological samples representative of one or more health states.

23. 23. The method of claim 22, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.

24. 24. The method of claim 23, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

25. 23. The method of claim 22, wherein the one or more conditions comprise cancer, a pre-cancerous condition, a non-malignant disease condition, or a disease-free condition.

26. 21. The method of claim 20, wherein the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.

27. The method of claim 1 , wherein the predictive model comprises a machine learning model.

28. 10. The method of claim 1, wherein the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a support vector machine, or any combination thereof.

29. 30. The method of claim 28, wherein the machine learning model comprises a machine learning classifier.

30. 28. The method of claim 27, wherein the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.

31. The method of claim 1 , wherein the predictive model is trained using leave-one-out validation.

32. 7. The method of claim 6, wherein the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof.

33. 33. The method of claim 32, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.

34. 10. The method of claim 1, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads.

35. 35. The method of claim 34, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.

36. 2. The method of claim 1, wherein the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

37. 10. The method of claim 1, wherein the sequencing comprises shotgun metagenomic sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.

38. 10. The method of claim 1, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.

39. 39. The method of Claim 38, wherein the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features.

40. 10. The method of claim 1, wherein the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.

41. 20. The method of claim 18, wherein the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.

42. 1. A method comprising: receiving data derived from one or more subject biological samples, electronic medical record information, one or more radiological images, and corresponding diseases; sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and identifying one or more features of the data derived from the one or more nucleic acid molecule sequencing reads, electronic medical record information, and the one or more radiological images corresponding to the disease for the one or more subjects.

43. 43. The method of Claim 42, wherein the one or more nucleic acid molecule sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input.

44. 43. The method of claim 42, wherein identifying comprises aligning the one or more sequencing reads to a genome database.

45. 43. The method of claim 42, further comprising training a predictive model using the data derived from the nucleic acid molecule sequencing reads, electronic medical record information, and the one or more radiological images of the one or more subjects and the one or more features of the corresponding disease.

46. 43. The method of claim 42, wherein the disease comprises cancer or a non-cancerous disease.

47. 43. The method of claim 42, further comprising identifying one or more characteristics of one or more protein biomarkers in the biological sample of the subject.

48. 48. The method of claim 47, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.

49. 43. The method of claim 42, wherein the biological sample comprises a liquid biopsy, a tissue biopsy, or a combination thereof.

50. 43. The method of claim 42, wherein the one or more radiological images comprise images of X-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof.

51. 47. The method of claim 46, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.

52. 43. The method of claim 42, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.

53. 53. The method of Claim 52, wherein said amplicon-based 16S rRNA sequencing sequences the V6 region of said one or more nucleic acid molecules.

54. 43. The method of claim 42, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.

55. 50. The method of claim 49, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

56. 47. The method of claim 46, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.

57. 47. The method of claim 46, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.

58. 43. The method of claim 42, wherein the one or more radiological image features comprise a Brock cancer probability score, a diameter of the lesion, spiculation of the lesion, solidity of the lesion, or any combination thereof.

59. 43. The method of Claim 42, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads.

60. 60. The method of claim 59, wherein the genome database comprises a microbial genome database.

61. 61. The method of claim 60, wherein the microbial genome database comprises a de novo metagenomic assembly.

62. 62. The method of claim 61 , wherein the de novo metagenomic assembly is derived from a biological sample representative of one or more health states.

63. 63. The method of claim 62, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.

64. 64. The method of claim 63, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

65. 63. The method of claim 62, wherein the one or more conditions comprise cancer, a precancerous condition, a non-malignant disease condition, or a disease-free condition.

66. 61. The method of claim 60, wherein the microbial genome database comprises a RefSeq database, a Web of Life database, a Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.

67. 60. The method of claim 59, wherein the genome database comprises a human genome database.

68. 46. ​​The method of claim 45, wherein the predictive model comprises a machine learning model.

69. 46. ​​The method of claim 45, wherein the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a random vector machine, or any combination thereof.

70. 69. The method of claim 68, wherein the machine learning model comprises a machine learning classifier.

71. 69. The method of claim 68, wherein the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.

72. 46. ​​The method of claim 45, wherein the predictive model is trained using leave-one-out validation.

73. 46. ​​The method of claim 45, wherein the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof.

74. 74. The method of claim 73, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.

75. 46. ​​The method of claim 45, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads.

76. 76. The method of claim 75, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.

77. 46. ​​The method of claim 45, wherein the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

78. 43. The method of claim 42, wherein the sequencing comprises shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof.

79. 43. The method of Claim 42, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.

80. 80. The method of Claim 79, wherein the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features.

81. 46. ​​The method of claim 45, wherein the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.

82. 60. The method of claim 59, wherein the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.

83. 1. A computer system configured to determine a disease in a subject, comprising: (a) one or more processors; (b) a non-transitory computer-readable storage medium containing software, the software causing the one or more processors of the computer system to: (i) receiving one or more sequencing reads, electronic medical record information, and one or more images of a subject's biological sample; and (ii) a computer system comprising executable instructions for causing a predictive model to determine disease in the subject as an output when data derived from one or more nucleic acid molecule sequencing reads, electronic medical record information, and one or more radiological images of the subject are provided as input to the predictive model.

84. 84. The computer system of Claim 83, wherein the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input.

85. 84. The computer system of claim 83, wherein the disease comprises cancer or a non-cancerous disease.

86. 84. The computer system of claim 83, wherein the biological sample comprises a tissue biopsy, a liquid biopsy, or a combination thereof.

87. 84. The computer system of claim 83, wherein the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject.

88. 88. The computer system of claim 87, wherein the predictive model is provided with the one or more protein biomarkers from the biological sample of the subject.

89. 88. The computer system of claim 87, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, or a combination thereof.

90. 88. The computer system of claim 87, wherein the predictive model is trained using data derived from the nucleic acid molecule sequencing reads, electronic medical record information, and the one or more radiological images of the one or more subjects and one or more features of the corresponding disease.

91. 89. The computer system of claim 88, wherein the executable instructions comprise identifying one or more characteristics of one or more protein biomarkers in the biological sample of the subject.

92. 88. The computer system of claim 87, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.

93. 88. The computer system of claim 87, wherein the one or more radiological images comprise images of x-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof.

94. 84. The computer system of claim 83, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.

95. 84. The computer system of Claim 83, wherein the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads.

96. 96. The computer system of Claim 95, wherein the amplicon-based 16S rRNA sequencing reads comprise sequencing reads of a V6 region of the one or more nucleic acid molecules.

97. 84. The computer system of Claim 83, wherein the one or more nucleic acid molecule sequencing reads comprise sequencing reads of mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.

98. 87. The computer system of claim 86, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

99. 85. The method of claim 84, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.

100. 85. The computer system of claim 84, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.

101. 84. The computer system of claim 83, wherein the one or more radiological image features comprise a Brock cancer probability score, a diameter of the lesion, a spiculation of the lesion, a solidity of the lesion, or any combination thereof.

102. 84. The computer system of Claim 83, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads.

103. 103. The computer system of claim 102, wherein the genome database comprises a microbial genome database.

104. 104. The computer system of claim 103, wherein the microbial genome database comprises a de novo metagenomic assembly.

105. 105. The computer system of claim 104, wherein the de novo metagenomic assembly is derived from biological samples representative of one or more health conditions.

106. 106. The computer system of claim 105, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.

107. 107. The computer system of claim 106, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

108. 106. The computer system of claim 105, wherein the one or more health conditions comprise a cancer, a pre-cancerous condition, a non-malignant disease condition, or a disease-free health condition.

109. 104. The computer system of claim 103, wherein the microbial genome database comprises the RefSeq database, the Web of Life database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.

110. 103. The computer system of claim 102, wherein the genome database comprises a human genome database.

111. 84. The computer system of claim 83, wherein the predictive model comprises a machine learning model.

112. 84. The computer system of claim 83, wherein the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a random vector machine, or any combination thereof.

113. 112. The computer system of claim 111, wherein the machine learning model comprises a machine learning classifier.

114. 112. The computer system of claim 111, wherein the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.

115. 84. The computer system of claim 83, wherein the predictive model is trained using leave-one-out validation.

116. 85. The computer system of claim 84, wherein the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof.

117. 117. The computer system of claim 116, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.

118. 84. The computer system of Claim 83, wherein the executable instructions further comprise decontaminating the one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads.

119. 119. The computer system of claim 118, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.

120. 84. The computer system of claim 83, wherein the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

121. 84. The computer system of Claim 83, wherein the one or more sequencing reads are generated by shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof.

122. 84. The computer system of Claim 83, wherein the executable instructions further comprise determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.

123. 123. The computer system of Claim 122, wherein the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features.

124. 84. The computer system of claim 83, wherein the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.

125. 103. The computer system of claim 102, wherein the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.

126. A method for determining a disease in a subject, comprising: receiving a biological sample from a subject; sequencing one or more nucleic acid molecules of the biological sample, thereby generating one or more nucleic acid molecule sequencing reads; and determining a disease in the subject as an output of a predictive model when the one or more nucleic acid molecular sequencing reads of the subject are provided to the predictive model, wherein the predictive model is trained using one or more nucleic acid molecular sequencing reads of one or more liquid biological samples and one or more tissue biological samples of one or more subjects and the corresponding diseases.

127. 127. The method of Claim 126, wherein the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input.

128. 127. The method of claim 126, wherein the disease comprises cancer, a non-cancerous disease, or a combination thereof.

129. 127. The method of claim 126, further comprising identifying one or more protein biomarkers from the biological sample of the subject.

130. 130. The method of claim 129, wherein the predictive model is provided with the one or more protein biomarkers from the biological sample of the subject.

131. 130. The method of claim 129, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.

132. 128. The method of claim 127, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.

133. 127. The method of claim 126, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.

134. 134. The method of claim 133, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.

135. 127. The method of Claim 126, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.

136. 127. The method of claim 126, wherein the liquid biopsy comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

137. 128. The method of claim 127, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.

138. 128. The method of claim 127, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.

139. 127. The method of Claim 126, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof, characteristics of the one or more nucleic acid sequencing reads that are provided as input to the predictive model.

140. 140. The method of claim 139, wherein the genome database comprises a microbial genome database.

141. 141. The method of claim 140, wherein the microbial genome database comprises a de novo metagenomic assembly.

142. 142. The method of claim 141, wherein the de novo metagenomic assembly is derived from a biological sample representative of one or more health conditions.

143. 143. The method of claim 142, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.

144. 144. The method of claim 143, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

145. 143. The method of claim 142, wherein the one or more conditions comprise cancer, a precancerous condition, a non-malignant disease condition, or a disease-free condition.

146. 141. The method of claim 140, wherein the microbial genome database comprises the RefSeq database, the Web of Life database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.

147. 140. The method of claim 139, wherein the genome database comprises a human genome database.

148. 127. The method of claim 126, wherein the predictive model comprises a machine learning model.

149. 127. The method of claim 126, wherein the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a random vector machine, or any combination thereof.

150. 149. The method of claim 148, wherein the machine learning model comprises a machine learning classifier.

151. 149. The method of claim 148, wherein the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.

152. 149. The method of claim 148, wherein the predictive model is trained using leave-one-out validation.

153. 149. The method of claim 148, wherein the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof.

154. 154. The method of claim 153, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.

155. 127. The method of Claim 126, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads, wherein the one or more decontaminated nucleic acid molecules are provided as input to the predictive model.

156. 156. The method of claim 155, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.

157. 127. The method of claim 126, wherein the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

158. 127. The method of claim 126, wherein the sequencing comprises shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof.

159. 127. The method of Claim 126, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.

160. 160. The method of Claim 159, wherein the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features.

161. 127. The method of claim 126, wherein the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.

162. 140. The method of claim 139, wherein the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.

163. 1. A method for identifying one or more non-human genomic features, comprising: receiving one or more liquid biosamples, one or more tissue biosamples of one or more subjects and corresponding diseases; sequencing one or more nucleic acid molecules of the one or more liquid biological samples and the one or more tissue biological samples, thereby generating one or more sequencing reads; and identifying from the one or more sequencing reads one or more non-human genomic features corresponding to the disease in the one or more subjects.

164. 164. The method of claim 163, wherein the one or more sequencing reads comprise one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input.

165. 164. The method of Claim 163, wherein identifying comprises aligning or mapping the one or more sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads.

166. 165. The method of claim 164, wherein the genome database comprises a microbial genome database.

167. 167. The method of claim 166, wherein the microbial genome database comprises a de novo metagenomic assembly.

168. 168. The method of claim 167, wherein the de novo metagenomic assembly is derived from a biological sample representative of one or more health conditions.

169. 169. The method of claim 168, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.

170. 170. The method of claim 169, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

171. 169. The method of claim 168, wherein the one or more conditions comprise cancer, a precancerous condition, a non-malignant disease condition, or a disease-free condition.

172. 167. The method of claim 166, wherein the microbial genome database comprises the RefSeq database, the Web of Life database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.

173. 164. The method of claim 163, further comprising training a predictive model using the one or more non-human genomic features and the corresponding diseases of the one or more subjects.

174. 164. The method of claim 163, wherein the disease comprises cancer or a non-cancerous disease.

175. 164. The method of claim 163, further comprising identifying one or more characteristics of one or more protein biomarkers in the one or more liquid biological samples, one or more tissue biological samples, or a combination thereof.

176. 176. The method of claim 175, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.

177. 175. The method of claim 174, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.

178. 164. The method of claim 163, wherein the sequencing comprises amplicon-based 16S rRNA sequencing.

179. 179. The method of claim 178, wherein the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.

180. 164. The method of Claim 163, wherein the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.

181. 164. The method of claim 163, wherein the liquid biological sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

182. 175. The method of claim 174, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.

183. 175. The method of claim 174, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.

184. 165. The method of claim 164, wherein the genome database comprises a human genome database.

185. 167. The method of claim 166, wherein the predictive model comprises a machine learning model.

186. 167. The method of claim 166, wherein the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a random vector machine, or any combination thereof.

187. 186. The method of claim 185, wherein the machine learning model comprises a machine learning classifier.

188. 186. The method of claim 185, wherein the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.

189. 167. The method of claim 166, wherein the predictive model is trained using leave-one-out validation.

190. 167. The method of claim 166, wherein the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof.

191. 191. The method of claim 190, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.

192. 164. The method of claim 163, further comprising decontaminating the one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads.

193. 193. The method of claim 192, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.

194. 167. The method of claim 166, wherein the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

195. 167. The method of claim 166, wherein the sequencing comprises shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof.

196. 164. The method of Claim 163, further comprising determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.

197. 200. The method of claim 196, wherein the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features.

198. 167. The method of claim 166, wherein the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.

199. 165. The method of claim 164, wherein the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.

200. 1. A computer system configured to determine a disease in a subject, comprising: (a) one or more processors; (b) a non-transitory computer-readable storage medium containing software, the software causing the one or more processors of the computer system to: (i) receiving one or more sequencing reads of a subject's biological sample; and (ii) determining a disease in the subject as an output of a predictive model when one or more nucleic acid molecular sequencing reads of the subject are provided to the predictive model, wherein the predictive model is trained using one or more nucleic acid molecular sequencing reads of one or more liquid biological samples and one or more tissue biological samples of one or more subjects and corresponding diseases.

201. The computer system of Claim 126, wherein the one or more nucleic acid molecule sequencing reads include one or more microbial nucleic acid molecule sequencing reads, and the predictive model is provided with the one or more microbial nucleic acid molecule sequencing reads as input.

202. 201. The computer system of claim 200, wherein the disease comprises cancer or a non-cancerous disease.

203. 201. The computer system of claim 200, wherein the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject.

204. 204. The computer system of claim 203, wherein the predictive model is provided with the one or more protein biomarkers from the biological sample of the subject.

205. 204. The computer system of claim 203, wherein the executable instructions comprise identifying one or more characteristics of one or more protein biomarkers in the biological sample of the subject.

206. 204. The computer system of claim 203, wherein the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.

207. 202. The computer system of claim 201, wherein the cancer comprises a tumor mass having a diameter of less than 3 centimeters.

208. 201. The computer system of claim 200, wherein the one or more nucleic acid molecule sequencing reads comprise one or more amplicon-based 16S rRNA sequencing reads.

209. The computer system of claim 208, wherein the amplicon-based 16S rRNA sequencing reads include sequencing reads of the V6 region of the one or more nucleic acid molecules.

210. 201. The computer system of claim 200, wherein the one or more nucleic acid molecule sequencing reads comprise sequencing reads of mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.

211. 201. The computer system of claim 200, wherein the liquid biological sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

212. 202. The method of claim 201, wherein the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof.

213. 202. The computer system of claim 201, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.

214. 201. The computer system of claim 200, further comprising mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads.

215. 215. The computer system of claim 214, wherein the genome database comprises a microbial genome database.

216. 216. The computer system of claim 215, wherein the microbial genome database comprises a de novo metagenomic assembly.

217. 217. The computer system of claim 216, wherein the de novo metagenomic assembly is derived from biological samples representative of one or more health conditions.

218. 218. The computer system of claim 217, wherein the biological sample is a tissue sample, a liquid biopsy sample, or a combination thereof.

219. 219. The computer system of claim 218, wherein the liquid biopsy sample comprises plasma, serum, whole blood, feces, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.

220. 218. The computer system of claim 217, wherein the one or more health conditions comprise a cancer, a precancerous condition, a non-malignant disease condition, or a disease-free health condition.

221. 216. The method of claim 215, wherein the microbial genome database comprises the RefSeq database, the Web of Life database, the Unified Human Gastrointestinal Genome (UHGG) database, or any combination thereof.

222. 215. The computer system of claim 214, wherein the genome database comprises a human genome database.

223. 201. The computer system of claim 200, wherein the predictive model comprises a machine learning model.

224. 201. The computer system of claim 200, wherein the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a random vector machine, or any combination thereof.

225. 224. The computer system of claim 223, wherein the machine learning model comprises a machine learning classifier.

226. 224. The computer system of claim 223, wherein the machine learning model comprises a stack machine learning model, one or more machine learning models, an ensemble machine learning model, or a combination thereof.

227. 201. The computer system of claim 200, wherein the predictive model is trained using leave-one-out validation.

228. 201. The computer system of claim 200, wherein the predictive model is configured to determine the stage of the cancer, the anatomical origin of the cancer, or a combination thereof.

229. 229. The computer system of claim 228, wherein the stage of the cancer is stage I, stage II, stage III, or stage IV.

230. 201. The computer system of claim 200, wherein the executable instructions further comprise decontaminating the one or more nucleic acid molecule sequencing reads to generate one or more decontaminated nucleic acid molecule sequencing reads.

231. 231. The computer system of claim 230, wherein the decontamination comprises in silico decontamination, experimental control decontamination, or a combination thereof.

232. 201. The computer system of claim 200, wherein the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

233. 201. The computer system of claim 200, wherein the one or more sequencing reads are generated by shotgun sequencing, next-generation sequencing, long-read sequencing, or any combination thereof.

234. 201. The computer system of claim 200, wherein the executable instructions further comprise determining one or more characteristics of the one or more nucleic acid molecule sequencing reads.

235. 235. The computer system of claim 234, wherein the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features.

236. 201. The computer system of claim 200, wherein the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.

237. 215. The computer system of claim 214, wherein the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.

Citation Information

Cited By

  • Apparatus and method for detecting multiorgan precancerous signs based on foundation model

    KR102989688B1