Methods for identifying cancer-associated microbial biomarkers

JP2024535736A5Pending Publication Date: 2025-10-10MICRONOMA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024513889
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-03
Filing Date
2022-09-02
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Current methods for diagnosing cancer and predicting response to anti-cancer therapy are limited by the sensitivity and specificity of traditional pathology and liquid biopsy-based assays, which often fail to differentiate between cancer types due to shared genomic aberrations and sensitivity issues with cell-free DNA.

Method used

The use of hybridization-based enrichment sequencing to isolate non-human microbial nucleic acids from human tissue or liquid biopsy samples, followed by machine learning algorithms to identify cancer-associated microbial signatures, allowing for accurate diagnosis and prediction of therapy response.

Benefits of technology

Enhances the sensitivity and specificity of cancer diagnosis and therapy prediction by leveraging microbial signatures, reducing reliance on traditional histology and improving the accuracy of treatment outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods for the identification of cancer-associated microbial signatures and their use in diagnostic and therapeutic stratification are provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 240,434, filed September 3, 2021, the entirety of which is incorporated herein by reference.

[0002] Incorporation by Reference All publications, patents, and patent applications mentioned in this application are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over such conflicting material. Summary of the Invention

[0003] The present disclosure provides methods for identifying cancer-associated microbial signatures and utilizing these identified signatures to accurately diagnose cancer and other non-cancerous conditions, their subtypes, and likelihood of response to anti-cancer therapy using non-human derived nucleic acids from human tissue or liquid biopsy samples. Specifically, the present invention provides methods for identifying the presence and abundance of microbial nucleic acids enriched from tissue or liquid biopsy samples by hybridization-based enrichment, and methods for using the presence or abundance of the microbial nucleic acids to diagnose and classify cancer in human subjects.

[0004] The methods of the invention disclosed herein provide a means to discover microbial features in mammalian genomic datasets derived from hybridization-based enrichment sequencing, and a method to validate the diagnostic or predictive utility of said microbial features. Hybridization-based enrichment, or "targeted enrichment," is a form of targeted sequencing that aims to enrich for genomic regions of interest while simultaneously depleting those regions that are not relevant for a given analysis. The aim is to limit the sequencing effort (and associated costs) to only those regions of the genome that are important for the disease / condition being investigated, a strategy that is cost-effective and allows for high sequencing depth (number of reads spanning a base) and reliable identification of, for example, important genomic mutations. This method is widely used for characterization of cancer tissue and cell-free DNA / RNA (cfDNA / cfRNA) obtained by liquid biopsy. In hybridization-based enrichment, tagged (e.g., biotinylated) oligonucleotide probes with complementarity to the genomic region of interest are mixed with the DNA sample such that nucleotide base pairing can occur between the sequence of the probe and the sequences present in the sample. The tagged probes are then recovered and sequenced. Hybridization probes can also be physically immobilized on a solid surface where they can base pair with solution-phase genomic fragments.

[0005] Numerous hybridization-based enrichment products for use in oncology are widely understood by those skilled in the art. For example, Agilent's "SureSelect Cancer All-In-One" product facilitates the identification of cancer-associated genomic variants. Its "SureSelect Cancer All-In-One Lung Assay" encompasses 20 genes (and all of their known somatic mutations) clinically relevant to non-small cell lung cancer, and its "SureSelect Cancer All-In-One Solid Tumor Assay" profiles 98 genes associated with common solid tumor types, including lung, breast, ovarian, colon, prostate, sarcoma, and skin. Using such kits, it is possible to preferentially enrich for these specific genes and known cancer variants and sequence them while the rest of the genome is depleted from downstream analysis.

[0006] It is important to emphasize that the intention of hybridization-based enrichment in the analysis of cancer samples is to specifically enrich regions of human genome.It has been found that an unexpected but useful by-product of oligonucleotide probe hybridization is a significant level of base pairing with non-human nucleic acid that has sufficient thermodynamic stability to result in the non-human nucleic acid being isolated together with the intended human genomic DNA fragment.It has also been found that this "bystander" enrichment can be shown to be reproducible for a given hybridization probe set, and the associated data derived from targeted sequencing datasets can be used to discover cancer-related microbial features.Given the widespread use of hybridization-based enrichment in cancer genomics and the availability of publicly available targeted sequencing datasets, these data can be a readily available source for the in silico discovery of microbial features with diagnostic utility, as described elsewhere herein.

[0007] Aspects disclosed herein describe a method for identifying microbial signatures for diagnosing cancer in a subject based on the analysis of hybridization-based enrichment sequencing data, comprising: (a) obtaining hybridization capture enriched sequencing reads from a biological sample; (b) filtering the sequencing reads with a genomic database build to isolate non-human sequencing reads; (c) generating taxonomic assignments and their associated abundances for the non-human sequencing reads; (d) identifying and removing contaminating microbial signatures of the taxonomically assigned non-human sequencing reads while retaining other decontaminated microbial signatures, thereby generating a decontaminated cancer-associated microbial signature set; and (e) validating the cancer-associated microbial signature set with known cancer and non-cancer samples to determine microbial signatures with discriminatory power for cancer versus non-cancer. In some embodiments, the biological sample is a tissue, liquid biopsy sample, or any combination thereof. In some embodiments, the subject is a human or non-human mammal. In some embodiments, the hybridization capture enrichment comprises multiplexed oligonucleotide probes targeting mammalian genomic regions. In some embodiments, the hybridization capture enriched sequencing reads comprise a total population of DNA, RNA, cell-free DNA (cfDNA), cell-free RNA (cfRNA), exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the genome database is a human genome database.

[0008] Aspects disclosed herein describe a method for validating identified cancer-associated microbial signatures, the method comprising: (a) hybridization capture-based enrichment of microbial sequences from known cancer and known non-cancer samples; (b) sequencing the captured nucleic acids and analyzing non-human reads to generate a taxonomic abundance table; (c) training a machine learning algorithm with the taxonomic abundance table to generate a trained machine learning model; (d) testing the trained machine learning model to determine its classification performance; and (e) generating model feature output used by the model to discriminate between cancer versus non-cancer conditions.

[0009] Aspects disclosed herein describe a method of creating a diagnostic model for diagnosing cancer in a subject based on non-human feature abundances in a biological sample, the method comprising: (a) obtaining hybridization capture enriched sequencing reads from the biological sample; (b) filtering the sequencing reads using a genomic database to isolate non-human sequencing reads; (c) generating taxonomic assignments and their associated abundances for the non-human sequencing reads; (d) identifying and removing contaminating microbial features of the taxonomically assigned non-human sequencing reads while retaining other decontaminated microbial features, thereby generating a decontaminated cancer-associated microbial feature set; and (e) training a machine learning algorithm using the decontaminated taxonomic abundances to generate a trained diagnostic model. In some embodiments, the biological sample is a tissue, liquid biopsy sample, or any combination thereof, from a subject undergoing anti-cancer therapy. In some embodiments, the subject is a human or non-human mammal. In some embodiments, the hybridization capture enrichment comprises multiplexed oligonucleotide probes targeting mammalian genomic regions. In some embodiments, the hybridization capture enriched sequencing reads comprise an enriched population of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the genome database is a human genome database. In some embodiments, the diagnostic model utilizes taxonomic abundance information from one or more of the following life domains: bacteria, archaea, and / or fungi. In some embodiments, the diagnostic model predicts a subject's response to chemotherapy, immunotherapy, neoadjuvant therapy, or any combination thereof.

[0010] In some embodiments, the diagnostic model diagnoses one or more of acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasms diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, or uveal melanoma. In some embodiments, the diagnostic model identifies and removes certain non-human features as contaminants, referred to as noise, while selectively retaining other non-human features, referred to as signals. In some embodiments, the liquid biopsy includes, but is not limited to, one or more of plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, or exhaled breath condensate. In some embodiments, the filtering includes computational filtering of the sequencing reads with bowtie2, Kraken programs, or any combination thereof.

[0011] Another aspect of the disclosure provided herein describes a method for identifying microbial signatures for determining disease in a subject, the method comprising: (a) exposing a biological sample of the subject to one or more probes, the one or more probes non-specifically binding to one or more nucleic acid molecules of the biological sample; (b) obtaining a first set of sequencing reads of one or more nucleic acid molecules bound to the one or more probes; (c) identifying a second set of sequencing reads within the first set of sequencing reads, the second set of sequencing reads comprising non-human sequencing reads obtained through non-specific hybridization; and (d) identifying one or more microbial signatures for determining disease in the subject from the second set of sequencing reads. In some embodiments, the biological sample is a tissue, a liquid biopsy sample, or any combination thereof. In some embodiments, the method further comprises generating a taxonomic assignment and abundance for the second set of sequencing reads. In some embodiments, the method further comprises removing one or more contaminant microbial signatures of the taxonomic assignment and abundance, thereby generating one or more decontaminated microbial signatures. In some embodiments, the subject comprises a human or non-human mammalian subject, hi some embodiments, the disease comprises cancer, a non-cancer disease, or a combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma of the cervix and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell renal carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof.In some embodiments, the one or more microbial features are from viruses, bacteria, fungi, archaea, or any combination of non-mammalian domains of life thereof. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions. In some embodiments, the first and second sequencing read sets comprise enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the identifying of step (c) comprises comparing the second sequencing read set to a genome database. In some embodiments, the genome database is a human genome database. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. In some embodiments, the method further comprises validating the microbial signature of the cancer-associated microbial signature, which includes (a) hybridization-based enrichment of microbial sequences from known cancer and known non-cancer samples; (b) sequencing the captured nucleic acids and analyzing the non-human reads to generate a taxonomic abundance table; (c) training a machine learning algorithm using the taxonomic abundance table to generate a trained machine learning model; (d) testing the trained machine learning model to determine its classification performance; and (e) generating an output of model features used by the model to distinguish between cancer and non-cancer conditions. In some embodiments, the hybridization capture-based enrichment includes multiplexed oligonucleotide probes targeting microbial genomic regions. In some cases, identifying the second set of sequencing reads includes filtering the first set of sequencing reads using bowtie2, Kraken, or a combination of these programs.

[0012] Another aspect of the disclosure provided herein describes a method of validating a microbial feature indicative of a disease in a subject, comprising: (a) receiving a first one or more microbial feature sets of a first biological sample from a first subject having a disease determined by non-specific interaction of a first one or more probe sets with one or more nucleic acid molecules of the first biological sample; (b) training a predictive model using the first one or more microbial feature sets of the first biological sample and the disease of the first subject, thereby generating a trained predictive model; (c) receiving a second one or more microbial feature sets of a second biological sample from a second subject having the disease; and (d) validating the first one or more microbial feature sets by comparing a predicted disease provided by the trained predictive model with a disease of the second subject, where a predicted disease provided by the trained predictive model is generated when the second one or more microbial feature sets are provided as inputs to the trained predictive model. In some embodiments, the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the first and second subjects comprise human or non-human mammalian subjects. In some embodiments, the first one or more microbial feature sets comprise taxonomic assignments and abundances of the first microbial sequencing read set, and the second one or more microbial feature sets comprise taxonomic assignments and abundances of the second microbial sequencing read set. In some embodiments, the disease of the first subject or the disease of the second subject comprises cancer, a non-cancerous disease, or a combination thereof. In some embodiments, the method further comprises removing one or more contaminant microbial features from the first one or more microbial feature sets, the second one or more microbial feature sets, or a combination thereof. In some embodiments, removing one or more contaminant microbial features is accomplished by in silico decontamination, experimental control, or a combination thereof.In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma of the cervix and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell renal carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the one or more microbial features are derived from viruses, bacteria, fungi, archaea, or any combination thereof. In some embodiments, the first one or more probe sets or the second one or more probe sets include multiplexed oligonucleotide probes targeting mammalian genomic regions. In some embodiments, the first one or more microbial feature sets and the second one or more microbial feature sets include enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the first one or more microbial feature sets or the second one or more microbial feature sets are determined by sequencing one or more nucleic acid molecules bound to the first one or more probe sets or the second one or more probe sets, thereby generating one or more sequencing reads, mapping the one or more sequencing reads to a genome database to identify one or more non-human sequencing reads, and determining the first one or more microbial feature sets or the second one or more microbial feature sets from the one or more non-human sequencing reads. In some embodiments, the first one or more probe sets or the second one or more probe sets comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. In some embodiments, the one or more microbial characteristics of the second biological sample are determined by sequencing enriched or non-enriched microbial nucleic acid molecules of the second biological sample.In some embodiments, the enriched microbial nucleic acid molecules are generated by exposing one or more nucleic acid molecules of the second biological sample to a second one or more probe sets, where the second one or more probe sets non-specifically bind to the one or more microbial nucleic acid molecules of the second biological sample.

[0013] Another aspect of the disclosure provided herein describes a method for training a predictive model using microbial features, the method comprising: (a) exposing a biological sample of a first subject having a first disease to one or more probes, where the one or more probes non-specifically bind to one or more nucleic acid molecules of the biological sample; (b) sequencing the one or more nucleic acid molecules bound to the one or more probes, thereby generating one or more sequencing reads; (c) mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads; and (d) generating a predictive model for predicting a second disease of a second subject, where the predictive model is trained using one or more microbial features of the one or more non-human sequencing reads and the first disease of the first subject. In some embodiments, the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. In some embodiments, the biological sample is obtained from a subject undergoing anti-cancer therapy. In some embodiments, the one or more microbial features are taxonomic assignments and abundances of the one or more non-human sequencing reads. In some embodiments, the method further comprises removing one or more contaminant microbial features from the one or more microbial features prior to training the predictive model. In some embodiments, removing one or more contaminant microbial features is accomplished by in silico decontamination, experimental control, or a combination thereof. In some embodiments, the first subject and the second subject comprise a human or non-human mammalian subject. In some embodiments, the one or more nucleic acids comprise one or more human nucleic acid molecules, non-human nucleic acid molecules, or a combination thereof. In some embodiments, the non-human nucleic acid molecules are derived from viruses, bacteria, fungi, archaea, or any combination thereof. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian nucleic acid molecules. In some embodiments, the one or more sequencing reads comprise an enriched population of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof.In some embodiments, the genome database is a human genome database. In some embodiments, the predictive model is configured to predict a subject's response to chemotherapy, immunotherapy, neoadjuvant therapy, or any combination of these therapies administered to treat the disease. In some embodiments, the first disease and the second disease include cancer, a non-cancerous disease, or a combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary renal cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, or uveal melanoma. In some embodiments, the predictive model is configured to identify and remove one or more contaminating microbial features while selectively retaining one or more non-contaminating microbial features. In some embodiments, the liquid biopsy sample comprises serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, identifying comprises computationally filtering one or more sequencing reads using bowtie2, Kraken, or a combination of these programs. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the machine learning model comprises one or more machine learning models, or an ensemble of machine learning models. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules.

[0014] Aspects of the disclosure provided herein describe a method comprising exposing a biological sample of a subject having a disease to one or more probes, where the one or more probes non-specifically bind to one or more nucleic acid molecules of the biological sample; identifying one or more sequencing reads of the one or more nucleic acid molecules bound to the one or more probes; mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more sequencing reads; and identifying one or more microbial features of the one or more non-human sequencing reads to classify the disease of the subject. In some embodiments, the biological sample comprises tissue, liquid biopsy, or any combination of such samples. In some embodiments, the one or more microbial features comprise taxonomic assignments and abundances of the non-human sequencing reads. In some embodiments, the method further comprises removing one or more contaminant microbial features of taxonomic assignments and abundances, thereby generating one or more decontaminated microbial features. In some embodiments, the subject comprises a human or non-human mammalian subject. In some embodiments, the disease comprises cancer, a non-cancerous disease, or a combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma of the cervix and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell renal carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the one or more microbial features are from viruses, bacteria, fungi, archaea, or any combination of these non-mammalian domains of life. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes targeted to mammalian genomic regions.In some embodiments, the one or more sequencing reads include sequencing reads of enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the genome database includes a human genome database. In some embodiments, the one or more probes include multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. In some embodiments, the one or more probes include multiplexed oligonucleotide probes that target mammalian nucleic acid molecules. In some embodiments, mapping includes filtering the one or more sequencing reads using bowtie2, Kraken, or a combination of these programs.

[0015] Aspects of the disclosure provided herein describe a system comprising one or more processors and a non-transitory computer-readable storage medium comprising software, the software comprising executable instructions that, as a result of execution, cause the one or more processors of the computer system to receive one or more nucleic acid molecule sequencing reads of a biological sample of a subject, the subject having a disease, the one or more nucleic acid molecule sequencing reads being obtained from one or more nucleic acid molecules enriched by one or more probes exposed to the biological sample of the subject; map the one or more nucleic acid molecule sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more nucleic acid molecule sequencing reads; and identify one or more microbial features of the one or more non-human sequencing reads to classify the disease of the subject. In some embodiments, the biological sample comprises a tissue, a liquid biopsy, or any combination of these samples. In some embodiments, the one or more microbial features comprise a taxonomic assignment and abundance of the one or more non-human sequencing reads. In some embodiments, the method further comprises removing one or more contaminant microbial signatures of taxonomic assignment and abundance, thereby generating one or more decontaminated microbial signatures. In some embodiments, removing one or more contaminant microbial signatures is accomplished by in silico decontamination, experimental control, or a combination thereof. In some embodiments, the subject comprises a human or non-human mammalian subject. In some embodiments, the disease comprises cancer, a non-cancerous disease, or a combination thereof.In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma of the cervix and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell renal carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, uveal melanoma, or any combination thereof. In some embodiments, the one or more microbial features are derived from viruses, bacteria, fungi, archaea, or any combination of non-mammalian life domains thereof. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions. In some embodiments, the one or more nucleic acid molecule sequencing reads comprise sequencing reads of enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the one or more probes comprise multiplexed oligonucleotide probes that non-specifically link to one or more microbial nucleic acid molecules. In some embodiments, mapping the one or more nucleic acid molecule sequencing reads comprises filtering the one or more nucleic acid molecule sequencing reads using bowtie2, Kraken, or a combination of these programs. In some embodiments, the software further comprises generating a predictive model, the predictive model being trained using the one or more microbial features and the disease of interest. In some embodiments, the predictive model comprises one or more machine learning models. In some embodiments, the predictive model comprises an ensemble of one or more machine learning models. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.In some embodiments, the predictive model is configured to predict a subject's response to chemotherapy, immunotherapy, neoadjuvant therapy, or any combination of these therapies administered to treat a disease.

[0016] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Brief description of the drawings]

[0017] [Figure 1A] 1 shows an exemplary microbial feature discovery scheme that incorporates feature validation of health- and cancer-associated microbial signatures to generate diagnostic models, as described in some embodiments herein. [Figure 1B] As described in some embodiments herein, an exemplary microbial feature discovery scheme incorporating feature validation of health- and cancer-associated microbial signatures to generate diagnostic models is shown. An exemplary method for validating the discovered microbial features of FIG. 1A to obtain diagnostic models that distinguish between healthy, cancerous, and non-cancerous conditions utilizing the microbial features of FIG. 1A is shown. [Figure 1C] As described in some embodiments herein, an exemplary microbial feature discovery scheme is shown that incorporates feature validation of health- and cancer-associated microbial signatures to generate diagnostic models. An exemplary method is shown for identifying microbial features associated with a subject's response to an anti-cancer therapy and generating a treatment response predictive machine learning model that utilizes those features. [Figure 2A] 1 shows an example of microbial feature discovery derived from a hybridization-based enrichment sequencing dataset, as described in some embodiments herein. 2 shows microbial reads present in a dataset of hybridization-based enrichment sequencing data. [Figure 2B]1 shows an example of microbial signature discovery derived from a hybridization-based enrichment sequencing dataset, as described in some embodiments herein. 2 shows the most abundant genera identified in hybridization-based enriched colorectal cancer cfDNA. [Figure 3A] FIG. 1 shows performance receiver operating characteristic (ROC) data for predictive models for predicting colorectal cancer based on bacterial abundance features of biological samples enriched with hybridization-based probes, as described in several embodiments herein. [Figure 3B] FIG. 1 shows performance receiver operating characteristic (ROC) data for predictive models for predicting colorectal cancer based on bacterial abundance features of biological samples enriched with hybridization-based probes, as described in several embodiments herein. [Figure 3C] FIG. 1 shows performance receiver operating characteristic (ROC) data for predictive models for predicting colorectal cancer based on bacterial abundance features of biological samples enriched with hybridization-based probes, as described in several embodiments herein. [Figure 4] 1 shows a schematic diagram of a computer system configured to implement the methods of the present disclosure as described in some embodiments herein. [Diagram 5] 1 shows a flow diagram for a method of validating one or more microbial characteristics as described in some embodiments herein. [Figure 6] 1 shows a flow diagram for a method of identifying one or more microbial characteristics as described in some embodiments herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] In some embodiments, the present invention provides a method for identifying one or more cancer-associated microbial features and utilizing these identified features to accurately diagnose cancer and other non-cancer conditions, their subtypes, and the likelihood of responding to anti-cancer therapies simply using non-human derived nucleic acids from a biological sample, where the biological sample may include human tissue or liquid biopsy samples. In some embodiments, this is accomplished by identifying microbial nucleic acids isolated by hybridization-based enrichment of mammalian genomic regions and then testing the usefulness of their microbial taxonomic abundance to distinguish subjects with and without cancer. In some embodiments, the identified microbial features and their presence or abundance in a subject's biological sample can be used to assign a probability that: (1) an individual has cancer, (2) an individual has cancer from a particular body site, (3) an individual has a particular type of cancer, and / or (4) a cancer that may or may not be diagnosed at the time is more or less likely to respond to a particular cancer therapy. Other uses of such methods are reasonably conceivable and readily feasible to those skilled in the art.

[0019] In some embodiments, the invention disclosed herein uses nucleic acids of non-human origin to diagnose a condition (i.e., cancer, non-cancerous disease, and / or disorder). In some embodiments, the disclosed invention may provide better clinical outcomes compared to typical pathology reports because it does not need to include one or more observed tissue architectures, cellular atypia, or other subjective measures traditionally used to diagnose cancer. In some embodiments, the disclosed methods may provide high sensitivity by focusing on microbial sources rather than modified human (i.e., cancerous) sources, which are often modified at very low frequencies in the background of "normal" human sources. In some embodiments, the methods disclosed herein may achieve such outcomes with either solid tissue or blood derived biological samples, the latter requiring minimal sample preparation and being minimally invasive. In some embodiments, liquid biopsy-based assays may overcome challenges posed by circulating tumor DNA (ctDNA) assays, which often suffer from sensitivity issues due to cell-free DNA (cfDNA) derived from non-malignant human cells. In some embodiments, liquid biopsy-based microbial assays may distinguish cancer types, which typically cannot be achieved with ctDNA assays, since the most common cancer genomic abnormalities are shared between cancer types (e.g., TP53 mutations, KRAS mutations). In some embodiments, as described elsewhere herein, methods may constrain the size of the signature, which may be predicted by those skilled in the art (e.g., regularized machine learning), and microbial assays may be made clinically available, for example, using multiplexed quantitative polymerase chain reaction (qPCR), and targeted assay panels for multiplexed amplicon sequencing, next-generation sequencing (NGS), or any combination thereof.

[0020] In some embodiments, the methods of the invention disclosed herein may include (a) analyzing a hybridization-based enrichment sequencing dataset and (b) identifying disease-associated microbial signatures present in the dataset. In some embodiments, the sequencing method may include next-generation sequencing or long-read sequencing (e.g., nanopore sequencing) or a combination thereof. In some embodiments, the targeted sequencing dataset 103 may result from isolating genomic regions of interest from a total nucleic acid sample from a subject with cancer 102 using a nucleic acid molecular capture probe, such as a DNA or RNA hybridization capture probe 101, as shown in FIG. 1A. In some embodiments, the microbial nucleic acids present in the hybridization probe sequencing dataset may be identified through taxonomic assignment 108, and human sequencing reads are computationally filtered from the total raw sequencing reads 103 via alignment to a human reference genome 104 using bowtie2 and / or Kraken or their equivalents. In some embodiments, the resulting non-human reads 105 may be taxonomically classified using bowtie2 or Kraken with a reference microbial database such as Web of Life. In some embodiments, the taxonomically assigned microbial reads 106 may be processed through decontamination 107 to remove sequences derived from common microbial contaminants resulting in a decontaminated cancer-associated microbial signature 109. In some embodiments, the decontaminated cancer-associated microbial signature 109 may serve as the basis for microbial-specific assays 110 intended to demonstrate the presence of these microorganisms in a subject's biological sample. In some embodiments, these microbial-specific assays 110 may include hybridization-based enrichment probes targeting genomic regions of the identified microbial taxa 109. In some embodiments, the microbial-specific assays 110 may include multiplex PCR assays to facilitate multiplexed amplicon sequencing.

[0021] In some embodiments, the methods disclosed herein may include a method of identifying one or more microbial signatures 600, as seen in Figure 6. In some cases, the method includes exposing a biological sample of a subject having a disease to one or more probes, where the one or more probes non-specifically bind to one or more nucleic acid molecules of the biological sample 602, identifying one or more sequencing reads of the one or more nucleic acid molecules bound to the one or more probes 604, mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more sequencing reads 606, and identifying one or more microbial signatures of the one or more non-human sequencing reads to classify the disease of the subject 608.

[0022] In some cases, decontamination may include in silico decontamination and / or experimental control decontamination. In some examples, decontamination may increase the area under the curve of the receiver operating characteristic curve of the predictive model by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60, at least 70%, at least 80%, at least 90%, or at least 95% compared to a predictive model trained with non-decontaminated microbial features. In some examples, in silico decontamination may include comparing individual microbial abundances across one or more biological samples of various analyte (e.g., nucleic acid molecule) concentrations. One or more contaminating microorganisms may be identified by a fractional abundance of microbial reads that is inversely proportional to the analyte concentration of one or more biological samples. For example, at lower analyte concentrations, the contaminating microorganisms will have a higher fractional read abundance compared to the overall abundance of microbial nucleic acids. In some examples, such decontamination methods may include (i) measuring a plurality of analyte concentrations from one or more biological samples of a subject; (ii) sequencing a plurality of nucleic acids at a plurality of dilutions to generate a plurality of nucleic acid sequences; (iii) mapping the plurality of nucleic acid sequencing reads to a microbial genome database, thereby generating a plurality of microbial nucleic acid reads for the plurality of dilutions; (iv) identifying a contaminating microbial organism from the plurality of microbial nucleic acid reads, wherein the contaminating microbial organism is present at a partial abundance that is inversely proportional to the plurality of dilutions across the one or more biological samples; and (v) removing the contaminating microbial organism signature from the microbial organism signature dataset to train a predictive model, as described elsewhere herein.

[0023] In some examples, the experimental control decontamination may include identifying the presence of microbial contaminants from nucleic acid molecules of the biological sample. In some cases, the experimental control decontamination may include identifying such microbial contaminants from one or more negative control samples (e.g., empty sample collection containers, vials, dishes, sealable containers, swabs, reagent-only vials, etc.). In some cases, the microbial contaminants may be removed from the identified microbial features prior to training the predictive model, as described elsewhere herein. In some cases, the microorganisms and their corresponding microbial nucleic acids are removed if they are identified in proportionally more negative control samples than in the biological sample. In some cases, the microorganisms and their corresponding microbial nucleic acids are removed based on a statistical test, such as the Fisher exact test, that accounts for the difference in the proportionality of the presence of microbial nucleic acids between the negative control and the biological sample. In some cases, the experimental control decontamination method may include (i) obtaining one or more negative control containers or chambers or reagents used to transport and / or store and / or process one or more biological samples; (ii) sequencing nucleic acid molecules in the one or more negative control containers, thereby generating a plurality of negative control sequencing reads; (iii) mapping the plurality of negative control sequencing reads to a microbial genome database, thereby generating a plurality of microbial nucleic acid molecule reads; and (iv) removing the plurality of negative control microbial nucleic acid molecule reads from the microbial nucleic acid molecule reads of the one or more biological samples prior to training a predictive model using one or more microbial features of the microbial nucleic acid molecule reads.

[0024] In some embodiments, the microbial signature 109 associated with cancer, a non-cancerous disease, disorder, or any combination thereof may be validated for use in cancer diagnosis by analyzing known non-cancer subjects 111 (which may include healthy subjects and / or subjects with non-cancer symptoms) and cancer subjects 112 using the microbial specific assay 110 of FIG. 1A as shown in FIG. 1B. In some embodiments, the microbial specific assay may include a sequencing-based assay to generate one or more sequencing reads of the hybridization enriched nucleic acid molecules of the biological sample 114. In some embodiments, the sequencing method may include next-generation sequencing or long-read sequencing (e.g., nanopore sequencing) or a combination thereof. In some embodiments, the sequencing reads may be processed through a taxonomic assignment pipeline 108 to result in a taxonomic abundance table that may be used to train a machine learning algorithm 115 to generate a trained diagnostic model 116. In some embodiments, the diagnostic model may be a regularized machine learning model. In some embodiments, the trained machine learning model algorithm may include linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k nearest neighbors (kNN), k-means, random forest algorithm models, or any combination thereof, as described elsewhere herein. In some embodiments, the identified microbial features 117 for diagnostic performance may be determined and used to justify the inclusion or exclusion of certain microbial features 109 from subsequent analyses, thereby facilitating the redesign of the microbial-specific assay 110 and validating the use of some (or all) of the microbial features 109 originally identified through the analysis of the human genome-directed hybridization-based enrichment sequencing dataset 103.

[0025] In some embodiments, the machine learning model 116 may be trained to be able to predict a subject's response to an anti-cancer therapy, as shown in FIG. 1C. In some embodiments, a hybridization-based enrichment sequencing dataset 103 derived from a cancer subject 118 undergoing therapy is processed through a taxonomic assignment pipeline 108 to obtain a taxonomic abundance table of treatment response-associated microorganisms. The taxonomic abundance table can be used to train a machine learning algorithm 115 to generate a trained diagnostic model 116. In some embodiments, the diagnostic model may be a regularized machine learning model. In some embodiments, the trained machine learning model algorithm may include linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbors (kNN), k-means, random forest algorithm models, or any combination thereof, as described elsewhere herein. In some embodiments, a microbial signature 120 identified to predict a response to a particular anti-cancer therapy may be identified.

[0026] Embodiments disclosed herein provide a method for identifying cancer-associated microbial signatures (FIG. 1A), comprising: (a) obtaining a human genome-oriented hybridization-based enrichment dataset 103; (b) computationally removing human sequencing reads from the dataset and generating taxonomic assignments 108 of the remaining non-human reads to obtain taxonomically identified cancer-associated microorganisms 109; (c) validating the presence of the identified cancer-associated microorganisms 109; and (d) evaluating the diagnostic value of the cancer-associated microorganisms (FIG. 1B).

[0027] Embodiments disclosed herein provide a method 500 of validating one or more microbial features, as shown in Figure 5. In some cases, the method includes receiving 502 a first one or more microbial feature sets of a first biological sample from a first subject having a disease determined by non-specific interaction of a first one or more probe sets with one or more nucleic acid molecules of the first biological sample, training a predictive model using the first one or more microbial feature sets of the first biological sample and the disease of the first subject, thereby generating a trained predictive model 504, receiving 506 a second one or more microbial feature sets of a second biological sample of a second subject having the disease, and validating 508 the first one or more microbial feature sets by comparing a predicted disease provided by the trained predictive model with a disease of the second subject, where a predicted disease provided by the trained predictive model is generated when the second one or more microbial feature sets are provided as inputs to the trained predictive model.

[0028] An embodiment disclosed herein provides a method of training a predictive model (FIG. 1C), comprising: (a) providing one or more sequenced microbial abundances 119 of one or more subjects as a training dataset; (b) providing one or more sequenced microbial abundances 119 of one or more subjects as a test set; (c) training the predictive model with a sample ratio of 60:40 training samples to validation samples, respectively; and (d) evaluating the predictive accuracy of the predictive model.

[0029] In some embodiments, predictions made by the trained predictive model may include machine learning signatures indicative of a therapy-responsive subject or machine learning derived signatures indicative of a therapy-non-responsive subject. In some embodiments, the trained predictive model may identify and remove one or more microbial or non-microbial nucleic acids classified as noise while selectively retaining one or more other microbial or non-microbial sequences, referred to as signals, through one or more decontamination methods, as described elsewhere herein.

[0030] In some embodiments, the microbial signature 109 may be validated for use in determining a disease state using an in silico approach. In some cases, a method of validating the microbial signature 109 for determining a disease state in silico may include (a) training a predictive model using microbial signatures of one or more subjects having known one or more disease states, thereby generating a trained predictive model, whereby the microbial signatures of the one or more subjects are determined by non-specific binding of one or more probes to one or more nucleic acid molecules of a biological sample of the one or more subjects, and (b) validating the microbial signature by comparing the disease state output of the trained predictive model when the trained predictive model is provided with a database of the microbial signatures of the one or more subjects and corresponding disease states. In some cases, the predictive model may include a machine learning model and / or algorithm. In some examples, the machine learning model may include one or more machine learning models, and / or an ensemble of machine learning models. In some cases, the database of microbial signatures of the one or more subjects may include one or more microbial genome segments. In some cases, the microbial signature may include the abundance of the corresponding microorganism represented by the one or more microbial genome segments. In some cases, the disease state may include healthy, cancerous, non-cancerous. In some cases, the cancer may include acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary renal cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, or uveal melanoma.

[0031] In some cases, the one or more genes may include about 1 gene to about 600 genes. In some cases, the one or more genes may include about 1 gene to about 5 genes, about 1 gene to about 15 genes, about 1 gene to about 25 genes, about 1 gene to about 50 genes, about 1 gene to about 100 genes, about 1 gene to about 150 genes, about 1 gene to about 200 genes, about 1 gene to about 300 genes, about 1 gene to about 400 genes, about 1 gene to about 500 genes, about 1 gene to about 600 genes, about 5 genes to about 15 genes, about ... about 5 genes to about 25 genes, about 5 genes to about 50 genes, about 5 genes to about 100 genes, about 5 genes to about 150 genes, about 5 genes to about 200 genes, about 5 genes to about 300 genes, about 5 genes to about 400 genes, about 5 genes to about 500 genes, about 5 genes to about 600 genes, about 15 genes to about 25 genes, about 15 genes to about 50 genes, about 15 genes to about 100 genes, about 15 genes to about 150 genes, about 15 genes about 200 genes, about 15 genes to about 300 genes, about 15 genes to about 400 genes, about 15 genes to about 500 genes, about 15 genes to about 600 genes, about 25 genes to about 50 genes, about 25 genes to about 100 genes, about 25 genes to about 150 genes, about 25 genes to about 200 genes, about 25 genes to about 300 genes, about 25 genes to about 400 genes, about 25 genes to about 500 genes, about 25 genes to about 600 genes genes, about 50 genes to about 100 genes, about 50 genes to about 150 genes, about 50 genes to about 200 genes, about 50 genes to about 300 genes, about 50 genes to about 400 genes, about 50 genes to about 500 genes, about 50 genes to about 600 genes, about 100 genes to about 150 genes, about 100 genes to about 200 genes, about 100 genes to about 300 genes, about 100 genes to about 400 genes, about 100 genes to about 500 genes,It may contain about 100 genes to about 600 genes, about 150 genes to about 200 genes, about 150 genes to about 300 genes, about 150 genes to about 400 genes, about 150 genes to about 500 genes, about 150 genes to about 600 genes, about 200 genes to about 300 genes, about 200 genes to about 400 genes, about 200 genes to about 500 genes, about 200 genes to about 600 genes, about 300 genes to about 400 genes, about 300 genes to about 500 genes, about 300 genes to about 600 genes, about 400 genes to about 500 genes, about 400 genes to about 600 genes, or about 500 genes to about 600 genes. In some cases, the one or more genes may include about 1 gene, about 5 genes, about 15 genes, about 25 genes, about 50 genes, about 100 genes, about 150 genes, about 200 genes, about 300 genes, about 400 genes, about 500 genes, or about 600 genes. In some cases, the one or more genes may include at least about 1 gene, about 5 genes, about 15 genes, about 25 genes, about 50 genes, about 100 genes, about 150 genes, about 200 genes, about 300 genes, about 400 genes, or about 500 genes. In some cases, the one or more genes may include up to about 5 genes, about 15 genes, about 25 genes, about 50 genes, about 100 genes, about 150 genes, about 200 genes, about 300 genes, about 400 genes, about 500 genes, or about 600 genes.

[0032] In some cases, the abundance of the corresponding microorganisms may include between about 1 microorganism and about 100 microorganisms. In some cases, the abundance of the corresponding microorganisms may include between about 1 microorganism and about 10 microorganisms, between about 1 microorganism and about 20 microorganisms, between about 1 microorganism and about 30 microorganisms, between about 1 microorganism and about 40 microorganisms, between about 1 microorganism and about 50 microorganisms, between about 1 microorganism and about 60 microorganisms, between about 1 microorganism and about 70 microorganisms, between about 1 microorganism and about 80 microorganisms, between about 1 microorganism and about 90 microorganisms, between about 1 microorganism and about 100 microorganisms, between about 10 microorganisms and about 20 microorganisms, between about 10 microorganisms and about 30 microorganisms, between about 10 microorganisms and about 200 microorganisms, between about 10 microorganisms and about 30 microorganisms, between about 10 microorganisms and about 40 microorganisms, between about 1 microorganism and about 50 microorganisms, between about 1 microorganism and about 60 microorganisms, between about 1 microorganism and about 70 microorganisms, between about 1 microorganism and about 80 microorganisms, between about 1 microorganism and about 90 microorganisms, between about 1 microorganism and about 100 microorganisms, between about 10 microorganisms and about 20 microorganisms, between about 10 microorganisms and about 30 microorganisms, between about 10 microorganisms and about 40 microorganisms, between about 10 microorganisms and about 50 microorganisms, between about 10 microorganisms and about 60 microorganism from about 10 microorganisms to about 40 microorganisms, from about 10 microorganisms to about 50 microorganisms, from about 10 microorganisms to about 60 microorganisms, from about 10 microorganisms to about 70 microorganisms, from about 10 microorganisms to about 80 microorganisms, from about 10 microorganisms to about 90 microorganisms, from about 10 microorganisms to about 100 microorganisms, from about 20 microorganisms to about 30 microorganisms, from about 20 microorganisms to about 40 microorganisms, from about 20 microorganisms to about 50 microorganisms, from about 20 microorganisms to about 60 microorganisms, from about 20 microorganisms to about 70 microorganisms, from about 20 microorganisms to about 80 microorganisms, from about from about 100 microorganisms to about 90 microorganisms, from about 20 microorganisms to about 100 microorganisms, from about 30 microorganisms to about 40 microorganisms, from about 30 microorganisms to about 50 microorganisms, from about 30 microorganisms to about 60 microorganisms, from about 30 microorganisms to about 70 microorganisms, from about 30 microorganisms to about 80 microorganisms, from about 30 microorganisms to about 90 microorganisms, from about 30 microorganisms to about 100 microorganisms, from about 40 microorganisms to about 50 microorganisms, from about 40 microorganisms to about 60 microorganisms, from about 40 microorganisms to about 70 microorganisms, from about 40 microorganisms to about 80 microorganisms, 0 microorganisms to about 90 microorganisms, about 40 microorganisms to about 100 microorganisms, about 50 microorganisms to about 60 microorganisms, about 50 microorganisms to about 70 microorganisms, about 50 microorganisms to about 80 microorganisms, about 50 microorganisms to about 90 microorganisms, about 50 microorganisms to about 100 microorganisms, about 60 microorganisms to about 70 microorganisms, about 60 microorganisms to about 80 microorganisms, about 60 microorganisms to about 90 microorganisms, about 60 microorganisms to about 100 microorganisms, about 70 microorganisms to about 80 microorganisms, about 70 microorganisms to about 90 microorganisms,It may include about 70 microorganisms to about 100 microorganisms, about 80 microorganisms to about 90 microorganisms, about 80 microorganisms to about 100 microorganisms, or about 90 microorganisms to about 100 microorganisms. In some cases, the corresponding microbial abundance may include about 1 microorganism, about 10 microorganisms, about 20 microorganisms, about 30 microorganisms, about 40 microorganisms, about 50 microorganisms, about 60 microorganisms, about 70 microorganisms, about 80 microorganisms, about 90 microorganisms, or about 100 microorganisms. In some cases, the corresponding microbial abundance may include at least about 1 microorganism, about 10 microorganisms, about 20 microorganisms, about 30 microorganisms, about 40 microorganisms, about 50 microorganisms, about 60 microorganisms, about 70 microorganisms, about 80 microorganisms, or about 90 microorganisms. In some cases, the corresponding abundance of microorganisms may include at most about 10 microorganisms, about 20 microorganisms, about 30 microorganisms, about 40 microorganisms, about 50 microorganisms, about 60 microorganisms, about 70 microorganisms, about 80 microorganisms, about 90 microorganisms, or about 100 microorganisms.

[0033] Although the above steps represent each of the methods or sets of operations according to the embodiments, one of ordinary skill in the art will recognize many variations based on the teachings provided herein. Steps may be performed in different orders. Steps may be added or omitted. Some of the steps may include sub-steps. Many of the steps may be repeated as often as is beneficial.

[0034] As described elsewhere herein, one or more of the steps of each of the methods or sets of operations may be performed using one or more of the circuitry described herein, e.g., a processor or logic circuitry such as programmable array logic for a field programmable gate array, and / or using a computer system. The circuitry may be programmed to provide one or more of the steps of each of the methods or sets of operations, and the program may include program instructions stored on a computer readable memory or programmed steps of a logic circuitry, e.g., programmable array logic or a field programmable gate array.

[0035] Predictive Model The disclosed methods and systems may utilize or access external capabilities of artificial intelligence, predictive models, and / or machine learning techniques to identify one or more microbial features of a hybridization-enriched biological sample. In some cases, the microbial features determined from a subject's hybridization-enriched biological sample may predict one or more subjects' cancer and / or non-cancerous diseases. In some cases, the features may be used to train one or more predictive models described elsewhere herein. These features may be used to accurately predict diseases, such as cancer, non-cancerous diseases, disorders, or any combination thereof. Using such predictive capabilities, health care providers (e.g., physicians) may make informed and accurate risk-based decisions, thereby improving the quality of care and monitoring provided to patients with cancer, non-cancerous diseases, disorders, or any combination thereof.

[0036] The methods and systems of the present disclosure may analyze the presence and / or abundance of microorganisms (e.g., abundance of microorganisms of a particular genus and / or taxonomy) of a biological sample enriched by a hybridization probe, where the hybridization probe may non-specifically bind to microbial nucleic acids, as described elsewhere. The presence and / or abundance of microorganisms may then be used to determine one or more microbial and / or non-microbial features that may predict cancer and / or non-cancerous disease in one or more subjects. In some cases, the methods and systems described elsewhere herein may train a predictive model using one or more microbial and / or non-microbial features indicative of cancer and / or non-cancerous disease in a subject. In some cases, the trained predictive model may then be used to generate a likelihood (e.g., a prediction) of cancer and / or non-cancerous disease for one or more subjects that are different from the one or more subjects utilized to train the predictive model. The trained predictive model may include an artificial intelligence-based model, such as a machine learning-based classifier, configured to process one or more microbial nucleic acid molecule sequencing reads obtained from the hybridization enriched biological sample to generate a likelihood that the subject has a disease or disorder. The model may be trained using the presence or abundance of microorganisms in hybridization enriched biological samples from one or more patient cohorts, such as cancer patients, patients with a non-cancerous disease, patients without disease and without cancer, cancer patients undergoing treatment for cancer, patients undergoing treatment for a non-cancerous disease, or any combination thereof. In some cases, the predictive model may be trained to provide a treatment prediction for treating cancer in one or more patients that are not part of the training dataset of the predictive model. Such a predictive model may output a treatment recommendation for one or more patients that are not part of the training dataset when provided with an input of the presence and abundance of one or more microorganisms in the patient's hybridization enriched biological sample.

[0037] The predictive model may include one or more predictive models. The model may include one or more machine learning algorithms. Examples of machine learning algorithms may include support vector machines (SVM), naive Bayes classification, random forests, neural networks (such as deep neural networks (DNN)), recurrent neural networks (RNN), deep RNN, long short-term memory (LSTM) recurrent neural networks (RNN), gated recurrent units (GRU), gradient boosting machines, random forests, or other supervised learning algorithms or unsupervised machine learning, statistics, linear regression, k-nearest neighbors, k-means, decision trees, logistic regression, or any combination thereof. The model may be used for classification or regression. The model may also involve the estimation of an ensemble model consisting of multiple predictive models, and may utilize techniques such as gradient boosting in the construction of gradient boosting decision trees. The model may be trained using one or more training datasets including one or more microbial features, patient data, such as the patient's medical history, the patient's family's medical history, patient vitals (e.g., blood pressure, pulse, temperature, oxygen saturation), or any combination thereof.

[0038] The predictive model may include any number of machine learning algorithms. In some embodiments, the random forest machine learning algorithm may be an ensemble of bagged decision trees. The ensemble may be at least about 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 500, 1000 or more bagged decision trees. The ensemble may be up to about 1000, 500, 250, 200, 180, 160, 140, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2 or fewer bagged decision trees. The ensemble can be about 1-1000, 1-500, 1-200, 1-100, or 1-10 bagged decision trees.

[0039] In some embodiments, the machine learning algorithm may have various parameters, which may be, for example, a learning rate, a mini-batch size, a number of epochs to train, momentum, learning weight decay, or neural network layers.

[0040] In some embodiments, the learning rate may be between about 0.00001 and 0.1.

[0041] In some embodiments, the mini-batch size can be about 16-128.

[0042] In some embodiments, the neural network may include neural network layers. The neural network may have at least about 2 to 1000 or more neural network layers.

[0043] In some embodiments, the number of epochs for training may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 500, 1000, 10000, or more.

[0044] In some embodiments, the momentum can be at least about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or more. In some embodiments, the momentum can be up to about 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, or less.

[0045] In some embodiments, the learning weight decay may be at least about 0.00001, 0.0001, 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, or more. In some embodiments, the learning weight decay may be up to about 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, 0.0001, 0.00001, or less.

[0046] In some embodiments, the machine learning algorithm may use a loss function, which may be, for example, a regression loss, a mean absolute error, a mean biased error, a hinge loss, an Adam optimizer, and / or a cross entropy.

[0047] In some embodiments, the parameters of the machine learning algorithm may be tuned with the aid of a human and / or a computer system.

[0048] In some embodiments, the machine learning algorithm may prioritize certain features. The machine learning algorithm may prioritize features that may be more relevant to detecting cancer, non-cancerous diseases, disorders, or any combination thereof. If a feature is classified more frequently than another feature in determining cancer, non-cancerous diseases, and / or disorders, the feature may be more relevant to detecting cancer, non-cancerous diseases, and / or disorders. In some cases, features may be prioritized using a weighting system. In some cases, features may be prioritized on a probability statistic based on the frequency and / or amount of occurrence of the feature. The machine learning algorithm may prioritize features with the help of a human and / or a computer system.

[0049] In some cases, a machine learning algorithm may prioritize certain features to reduce computational cost, conserve processing power, conserve processing time, improve reliability, or reduce random access memory usage, etc.

[0050] The training dataset may be generated, for example, from one or more cohorts of patients with a diagnosis of a common cancer, non-cancerous disease, or disorder. The training dataset may include one or more microbial features in the form of microbial presence and / or abundance of hybridization enriched biological samples of one or more subjects. The features may include one or more subject cancer diagnoses corresponding to the microbial features. In some cases, the features may include patient information such as the patient's age, the patient's medical history, other medical conditions, current or past medications, clinical risk scores, and time since last observation. For example, a set of features collected from a given patient at a given time point may collectively function as a signature that may be indicative of the patient's health state or status at that given time point.

[0051] The label may include, for example, a clinical outcome, such as the presence, absence, diagnosis, and / or prognosis of a cancer, a non-cancer disease, disorder, or a combination thereof, in a subject (e.g., a patient). Clinical outcomes may include treatment efficacy (e.g., whether the subject is a positive or negative responder to a cancer and / or disease-based treatment).

[0052] The input features may be structured by aggregating the data into bins, or alternatively by using one-hot encoding. The input may also include feature values ​​or vectors derived from previous inputs such as cross-correlation.

[0053] The training dataset may be constructed from one or more microbial presence and / or abundance features in the hybridization enriched biological sample, or a combination of one or more microbial presence and / or abundance features with one or more somatic nucleic acid molecules of the hybridization enriched biological sample that are indicative of cancer, a non-cancerous disease, disorder, or any combination thereof.

[0054] The model may process the input features to generate output values ​​including one or more classifications, one or more predictions, or a combination thereof. For example, such classifications or predictions may include a binary classification of the presence or absence of cancer, the presence of a non-cancerous disease, the presence of a disorder, or any combination of those classifications of the subject. In some cases, one or more predictive models and / or machine learning algorithms may classify the subject between a group of categorical labels (e.g., "no cancer, non-cancer disease and / or disorder", "definite cancer, non-cancer disease and / or disorder", and "possible cancer, non-cancer disease and / or disorder") and a score indicating the likelihood (e.g., relative likelihood or probability) of developing a particular cancer, non-cancer disease, and / or disorder, a "risk factor" for the presence of cancer, non-cancer disease and / or disorder, the likelihood of the patient's death, and a confidence interval for any numerical prediction. Various machine learning techniques may be cascaded such that the output of the machine learning techniques may also be used as input features to subsequent layers or subsections of the model.

[0055] To train a model (e.g., by determining model weights and correlations) for generating real-time classifications or predictions, the model can be trained using a training dataset and / or one or more training features described elsewhere herein. Such a dataset and / or features can be large enough to generate statistically significant classifications or predictions. For example, a dataset can include a database of data including the presence and / or abundance of fungi, viruses, archaea, bacteria, or any combination of these microorganisms in one or more subject biological samples.

[0056] The dataset may be divided into subsets (e.g., discrete or overlapping), such as a training dataset, a development dataset, and a test dataset. For example, the dataset may be divided into a training dataset that comprises 80% of the dataset, a development dataset that comprises 10% of the dataset, and a test dataset that comprises 10% of the dataset. The training dataset may comprise about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The development dataset may comprise about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The test dataset may comprise about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, leave-one-out cross-validation may be used. The training set (e.g., training data set) may be selected by random sampling of a set of data corresponding to one or more patient cohorts to ensure sampling independence. Alternatively, the training set (e.g., training data set) may be selected by proportional sampling of a set of data corresponding to one or more patient cohorts to ensure sampling independence.

[0057] To improve the accuracy of model predictions and reduce overfitting of the model, the dataset may be expanded to increase the number of samples in the training set. For example, data expansion may include rearranging the order of observations in the training records. To handle datasets with missing observations, methods for imputing missing data such as forward imputation, backward imputation, linear interpolation, and multitask Gaussian processes may be used. The dataset may be filtered or batch corrected to remove or reduce confounding factors. For example, within the database, a subset of patients may be excluded.

[0058] The model may include one or more neural networks, such as a neural network, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), or a deep RNN. The recurrent neural network may include units that may be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, the model may include an algorithmic architecture including a neural network having an input feature set, such as microbial features, vital measurements, a patient's medical history, a patient's demographics, or any combination thereof, as described elsewhere herein. Neural network techniques such as dropout or regularization may be used during training of the model to prevent overfitting. The neural network may include multiple sub-networks, each of which is configured to generate classifications or predictions of different types of output information, which may be combined to form an overall output of the neural network. The machine learning model may alternatively utilize statistical or related algorithms, including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, and ensemble and gradient boosted variants thereof.

[0059] If the model generates a classification or prediction of cancer, a non-cancerous disease, a disorder, or a combination thereof, a notification (e.g., an alert or alarm) may be generated and transmitted to a healthcare provider, such as a doctor, a nurse, or other member of the patient's care team in a hospital. The notification may be transmitted via an automated phone call, a short message service (SMS), a multimedia message service (MMS) message, an email, and / or an alert in a dashboard. The notification may include output information such as a prediction of cancer, a non-cancerous disease, and / or a disorder; a predicted probability of the cancer, a non-cancerous disease, and / or a disorder; an expected time to onset of the cancer, a non-cancerous disease, and / or a disorder; a confidence interval of the probability or time, a recommended course of treatment for the cancer, a non-cancerous disease, and / or a disorder, or any combination of that information.

[0060] Different performance metrics may be generated to validate the performance of the model. For example, the area under the receiver operating characteristic curve (AUROC) may be used to determine the diagnostic, prognostic, screening, or any combination of their capabilities of the model. For example, the model may use adjustable classification thresholds so that the specificity and sensitivity are adjustable, and the receiver operating characteristic curve (ROC) may be used to identify different operating points that correspond to different values ​​of specificity and sensitivity.

[0061] In some cases, such as when the dataset is not large enough, cross-validation may be performed to assess the robustness of the model across different training and testing datasets.

[0062] The following definitions may be used to calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), area under the precision-recall curve (AUPR), AUROC, or the like. A "false positive" may refer to an outcome in which a positive outcome or result is generated incorrectly or prematurely (e.g., before or without the actual onset of a cancer, non-cancer disease and / or disorder). A "true positive" may refer to an outcome in which a positive outcome or result is correctly generated when a patient has a cancer, non-cancer disease and / or disorder (e.g., the patient exhibits symptoms of a cancer, non-cancer disease and / or disorder or the patient's records indicate a cancer, non-cancer disease and / or disorder). A "false negative" may refer to an outcome in which a negative outcome or result is generated but the patient has a cancer, non-cancer disease and / or disorder (e.g., the patient exhibits symptoms of a cancer, non-cancer disease and / or disorder or the patient's records indicate a cancer, non-cancer disease and / or disorder). A "true negative" can refer to an outcome in which a negative outcome or result is produced (e.g., prior to or without the actual onset of cancer, a non-cancerous disease and / or disorder).

[0063] The model may be trained until certain predefined conditions for accuracy or performance are met, such as having a minimum desired value corresponding to a diagnostic accuracy measure. For example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of occurrence of cancer, non-cancerous disease and / or disorder in a subject. As another example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of worsening or recurrence of a cancer, non-cancerous disease and / or disorder for which a subject has previously been treated. Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, AUPR, and AUROC, which correspond to diagnostic accuracy for detecting or predicting cancer, non-cancerous disease and / or disorder.

[0064] For example, such a predetermined condition can be that the sensitivity for predicting cancer, a non-cancerous disease and / or disorder includes values, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0065] As another example, such a predetermined condition can be a specificity for predicting cancer, a non-cancerous disease and / or disorder comprising a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0066] As another example, such a predetermined condition can be a positive predictive value (PPV) for predicting cancer, a non-cancerous disease and / or disorder including, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0067] As another example, such a predetermined condition can be a negative predictive value (NPV) for predicting cancer, a non-cancerous disease and / or disorder that includes a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0068] As another example, such a predetermined condition may be that the area under the curve (AUC) (AUROC) of a receiver operating characteristic curve (ROC) predicting cancer, non-cancerous disease and / or disorder comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0069] As another example, such a predetermined condition may be that the area under the precision-recall curve (AUPR) predicting cancer, non-cancer diseases and / or disorders includes a value of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0070] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0071] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0072] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with a positive predictive value (PPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0073] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases and / or disorders with a negative predictive value (NPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.

[0074] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with an area under the curve (AUC) (AUROC) of a receiver operating characteristic curve (ROC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0075] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases and / or disorders with an area under the precision-recall curve (AUPR) of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0076] The training dataset may be collected from training subjects (e.g., humans), each having a diagnosis status indicating that they have been diagnosed with a biological condition or have not been diagnosed with cancer, a non-cancerous disease and / or disorder.

[0077] In some embodiments, the model is a neural network or a convolutional neural network, see Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.

[0078] In some embodiments, independent component analysis (ICA) is used to de-dimensionalize the data as described in Lee, T.-W. (1998): Independent component analysis: Theory and applications, Boston, Mass: Kluwer Academic Publishers, ISBN 0-7923-8261-7, and Hyvarinen, A.; Karhunen, J.; Oja, E. (2001): Independent Component Analysis, New York: Wiley, ISBN 978-0-471-40540-5, which are incorporated by reference in their entireties.

[0079] In some embodiments, principal component analysis (PCA) is used to de-dimensionalize the data as described in Jolliffe, IT (2002). Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag. doi:10.1007 / b98835. ISBN 978-0-387-95442-4, which is incorporated herein by reference in its entirety.

[0080] SVM is a method used by Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge, Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5 thAnnual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVM separates a given set of binary labeled data using a hyperplane that is maximally far away from the labeled data. When linear separation is not possible, SVMs can work in conjunction with "kernel" techniques that automatically achieve a nonlinear mapping to the feature space: the hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.

[0081] Decision trees are reviewed by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Tree-based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that may be used is Classification and Regression Trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and Random Forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York. pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests-Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.

[0082] Clustering (e.g., unsupervised and supervised clustering model algorithms) are described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter, "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a data set. To identify natural groupings, two problems are addressed. First, determine how to measure the similarity (or dissimilarity) between two samples. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than samples in the other cluster. Second, determine a mechanism for splitting the data into clusters using the similarity measure. Similarity measures are discussed in section 6.7 of Duda 1973, where it is stated that one way to begin a clustering study is to define a distance function and compute a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, then the distance between reference entities in the same cluster will be significantly smaller than the distance between reference entities in different clusters. However, as stated on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x,x') can be used to compare two vectors x and x'. Conventionally, s(x,x') is a symmetric function whose value is large if x and x' are "similar" in some way. An example of a non-metric similarity function s(x,x') is provided on page 218 of Duda 1973. Once a method for measuring "similarity" or "dissimilarity" between points in a data set has been selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. A partition of the data set that extremizes a criterion function is used to cluster the data. See Duda 1973, p. 217.Criterion functions are discussed in section 6.8 of Duda 1973. More recently, Duda et al., Pattern Classification, 2nd edition, John Wiley & Sons, Inc. New York, was published. Clustering is described in detail on pages 537-563. Further information on clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is incorporated herein by reference. Certain exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum of squares algorithms), k-means clustering, fuzzy k-means clustering algorithms, and Jarvis-Patrick clustering. In some embodiments, the clustering comprises unsupervised clustering, where no preconceptions are imposed as to which clusters should be formed when a training set is clustered.

[0083] Regression models such as multi-category logit models are described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, Inc., New York, Chapter 8, which is incorporated herein by reference in its entirety. In some embodiments, the model utilizes the regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, which is incorporated herein by reference in its entirety. In some embodiments, gradient boosting models are used, for example, for the classification algorithms described herein, and these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019) "Gradient Boosting". Hands-On Machine Learning with R. Chapman & Hall. pp. 221-245. ISBN 978-1-138-49568-5., which is incorporated herein by reference in its entirety. In some embodiments, ensemble modeling techniques are used, which are described in the implementation of classification models herein and in Zhou Zhihua (2012). Ensemble Methods: Foundations and Algorithms. Chapman and Hall / CRC. ISBN 978-1-439-83003-1, which is incorporated herein by reference in its entirety.

[0084] In some embodiments, the machine learning analysis is performed by a device executing one or more programs (e.g., one or more programs stored in non-persistent memory or in persistent memory) that include instructions for performing the data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., a processing core) and a memory (e.g., one or more programs stored in non-persistent memory or in persistent memory) that includes instructions for performing the data analysis.

[0085] system The present disclosure provides a computer system programmed to implement the methods of the present disclosure. Figure 4 shows a computer system 400 programmed or otherwise configured to predict cancer, non-cancerous disease, or any combination thereof, train a predictive model, generate a recommended treatment, or perform any combination of the methods described elsewhere herein. The computer system 400 may be a user's electronic device or a computer system located remotely to the electronic device. The electronic device may be a mobile electronic device.

[0086] The computer system 400 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 406, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 400 also includes memory or memory locations 404 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 402 (e.g., hard disk), a communication interface 408 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 410, such as cache, other memory, data storage, and / or electronic display adapters. The memory 404, the storage unit 402, the interface 408, and the peripheral devices 410 communicate with the CPU 406 through a communication bus (solid lines), such as a motherboard. The storage unit 402 may be a data storage unit (or data repository) for storing data. The computer system 400 may be operatively coupled to a computer network ("network") 412 with the aid of the communication interface 408. The network 412 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, the network 412 is a telecommunications and / or data network. The network 412 may include one or more computer servers, which may enable distributed computing, such as cloud computing. The network 412 may implement a peer-to-peer network, in some cases with the aid of the computer system 400, which may enable devices coupled to the computer system 400 to act as clients or servers.

[0087] CPU 406 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 404. The instructions can be directed to CPU 406, which can then program or otherwise configure CPU 406 to implement the methods of the present disclosure described elsewhere herein. Examples of operations performed by CPU 406 can include fetch, decode, execute, and writeback.

[0088] The CPU 406 may be part of a circuit, such as an integrated circuit. One or more of the other components of the system 400 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0089] The storage unit 402 can store files such as drivers, libraries, and saved programs. The storage unit 402 can store user data, such as user preferences and user programs. In some cases, the computer system 400 can include one or more additional data storage units that are external to the computer system 400, such as located on a remote server that communicates with the computer system 400 through an intranet or the Internet.

[0090] Computer system 400 can communicate with one or more remote computer systems through network 412. For example, computer system 400 can communicate with a remote computer system of a user. Examples of remote computer systems may include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 400 via network 412.

[0091] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored on electronic storage locations of computer system 400, such as, for example, on memory 404 or electronic storage unit 402. Machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by processor 406. In some cases, the code may be retrieved from storage unit 402 and stored in memory 404 for ready access by processor 406. In some circumstances, electronic storage unit 402 may be excluded and machine executable instructions are stored on memory 404.

[0092] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or as-compiled manner.

[0093] In some embodiments, as described elsewhere herein, the system may include one or more processors and a non-transitory computer-readable storage medium including software, the software including executable instructions that, when executed, cause the one or more processors of the computer system to receive one or more nucleic acid molecule sequencing reads of a subject's biological sample, where the subject has a disease, and the one or more nucleic acid molecule sequencing reads are obtained from one or more nucleic acid molecules enriched by one or more probes exposed to the subject's biological sample; map the one or more nucleic acid molecule sequencing reads to a genomic database, thereby identifying one or more non-human sequencing reads of the one or more nucleic acid molecule sequencing reads; and identify one or more microbial features of the one or more non-human sequencing reads to classify the disease in the subject.

[0094] Aspects of the systems and methods provided herein, such as computer system 400, can be embodied in programming. Various aspects of the technology can be considered as "products" or "articles of manufacture" that are typically in the form of machine (or processor) executable code and / or associated data executed or embodied in some type of machine-readable medium. The machine executable code can be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of the tangible memory of a computer, processor, etc., or their associated modules (various semiconductor memories, tape drives, disk drives, etc.), which may provide non-transitory storage at any time for software programming. All or portions of the software may be communicated from time to time over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from a management server or host computer to a computer platform of an application server. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical terrestrial communication networks, and on various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered as media that carry software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0095] Thus, a machine-readable medium such as a computer executable code may take many forms, including but not limited to a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer(s) that may be used to implement, for example, the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire and fiber optics, including the wires that make up a bus in a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, such as those generated during radio frequency (RF) and infrared (IR) data communications, or sound or light waves. Thus, common forms of computer readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium having a pattern of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave carrying data or instructions, a cable or link carrying such a carrier wave, or any other medium from which programming code and / or data may be read by a computer. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0096] The computer system 400 may include or be in communication with an electronic display 414 that includes a user interface (UI) 416, for example, for providing a display for visualizing prediction results or an interface for training a predictive model. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0097] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in the practice of the present invention. It is therefore contemplated that the present invention also encompasses any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

[0098] definition Unless otherwise defined, all terminology, notation, and other technical and scientific terms or glossaries used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms having a commonly understood meaning are defined herein for clarity and / or ease of reference, and the inclusion of such definitions herein should not necessarily be construed as representing a substantial difference from what is commonly understood in the art.

[0099] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Thus, the description of a range should be considered to specifically disclose all possible subranges as well as individual numerical values ​​within that range. For example, the description of a range such as 1-6 should be considered to specifically disclose subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc., as well as individual numerical values ​​within that range, such as 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0100] As used in this specification and claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a sample" includes multiple samples (including mixtures thereof).

[0101] The terms "determining," "measuring," "evaluating," "assessing," "assaying," and "analyzing" are often used interchangeably herein to refer to forms of measurement. These terms include determining whether an element is present or not (e.g., detecting). These terms can include quantitative, qualitative, or quantitative and qualitative determinations. Evaluating can be relative or absolute. "Detecting the presence of" can include determining the amount of something present in addition to determining whether it is present or absent, depending on the context.

[0102] The terms "subject," "individual," or "patient" are often used interchangeably herein. A "subject" can be a biological entity that contains expressed genetic material. The biological entity can be a plant, an animal, or a microorganism (including, for example, bacteria, viruses, fungi, and protozoa). A subject can be tissues, cells, and their progeny of a biological entity obtained in vivo or cultured in vitro. A subject can be a mammal. A mammal can be a human. A subject can be diagnosed or suspected of being at high risk for a disease. In some cases, a subject is not necessarily diagnosed or suspected of being at high risk for a disease.

[0103] The term "hybridization-based enrichment" is used to describe the use of oligonucleotide probes having nucleobase pairing complementarity to regions of the genome to specifically bind, via Watson-Crick base pairing interactions, and thereby isolate genomic DNA or RNA fragments from a sample by their association with the oligonucleotide probes.

[0104] The term "taxonomic abundance" is used to describe the number of sequencing reads that can be assigned to an identified microbial taxon in each sample.

[0105] The term "in vivo" is used to describe events that take place inside a subject's body.

[0106] The term "ex vivo" is used to describe events that occur outside of a subject's body. An ex vivo assay is not performed on a subject. Rather, it is performed on a sample separate from the subject. An example of an ex vivo assay that is performed on a sample is an "in vitro" assay.

[0107] The term "in vitro" is used to describe events that occur contained within a container for holding laboratory reagents such that the substance is separate from the biological source from which it is obtained. In vitro assays can include cell-based assays that use live or dead cells. In vitro assays can also include cell-free assays that do not use intact cells.

[0108] As used herein, the term "about" of a number refers to that number plus or minus 10%. The term "about" of a range refers to that range of minus 10% of its minimum value and plus 10% of its maximum value.

[0109] The use of absolute or sequential terms, such as "will," "will not," "shall," "shall not," "must," "must not," "first," "initially," "next," "subsequently," "before," "after," "lastly," and "finally" are not intended to limit the scope of the embodiments disclosed herein, but are exemplary.

[0110] Any systems, methods, software, compositions, and platforms described herein are modular and not limited to sequential steps, and thus, terms such as "first" and "second" do not necessarily imply a priority, order of importance, or order of actions.

[0111] As used herein, the term "treatment" or "treating" is used in reference to a pharmaceutical or other intervention regimen to obtain a beneficial or desired result in a recipient. Beneficial or desired results include, but are not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit may refer to the eradication or amelioration of the condition or underlying disease being treated. Therapeutic benefit may also be achieved by eradication or amelioration of one or more of the physiological symptoms associated with the underlying disease, such that an improvement is observed in the subject, even if the subject is still suffering from the underlying disease. Prophylactic effects include delaying, preventing, or eliminating the appearance of a disease or condition, delaying or eliminating the onset of symptoms of a disease or condition, slowing, halting, or reversing the progression of a disease or condition, or any combination thereof. For prophylactic benefit, a subject at risk of developing a particular disease or reporting one or more of the physiological symptoms of a disease may be treated even if a diagnosis of the disease has not been made.

[0112] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0113] Embodiment Numbered embodiment 1 includes a method of identifying a microbial signature for determining a disease of a subject, the method comprising: exposing a biological sample of the subject to one or more probes, the one or more probes non-specifically binding to one or more nucleic acid molecules of the biological sample; obtaining a first set of sequencing reads of one or more nucleic acid molecules bound to the one or more probes; identifying a second set of sequencing reads within the first set of sequencing reads, the second set of sequencing reads comprising non-human sequencing reads obtained through non-specific hybridization; and identifying one or more microbial signatures for determining a disease of the subject from the second set of sequencing reads. Numbered embodiment 2 includes the method of embodiment 1, the biological sample comprises a tissue, a liquid biopsy, or a combination of these samples. Numbered embodiment 3 includes the method of embodiment 1 or embodiment 2, further comprising generating a taxonomic assignment and abundance for the second set of sequencing reads. Numbered embodiment 4 includes the method of any one of embodiments 1-3, further comprising removing one or more contaminant microbial characteristics of taxonomic assignment and abundance, thereby generating one or more decontaminated microbial characteristics. Numbered embodiment 5 includes the method of any one of embodiments 1-4, wherein the subject comprises a human or non-human mammalian subject. Numbered embodiment 6 includes the method of any one of embodiments 1-5, wherein the disease comprises cancer, a non-cancerous disease, or a combination thereof.Numbered embodiment 7 includes the method of any one of embodiments 1-6, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma of the cervix and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasms diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, sarcoma of the uterine carcinoma, endometrial carcinoma of the uterine body, uveal melanoma, or any combination thereof. Numbered embodiment 8 includes the method of any one of embodiments 1-7, wherein the one or more microbial features are from viruses, bacteria, fungi, archaea, or any combination of non-mammalian life domains thereof. Numbered embodiment 9 includes the method of any one of embodiments 1-8, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions. Numbered embodiment 10 includes the method of any one of embodiments 1-9, wherein the first and second sequencing read sets comprise enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 11 includes the method of any one of embodiments 1-10, wherein the identifying of step (c) comprises comparing the second sequencing read set to a genome database. Numbered embodiment 12 includes the method of any one of embodiments 1-11, wherein the genome database is a human genome database. Numbered embodiment 13 includes the method of any one of embodiments 1 to 12, wherein one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. Numbered embodiment 14 includes the method of any one of embodiments 1 to 13, wherein one or more probes comprise multiplexed oligonucleotide probes that target mammalian genomic regions.Numbered embodiment 15 includes the method of any one of embodiments 1 to 14, wherein identifying the second sequencing read set includes filtering the first sequencing read set using bowtie2, Kraken, or a combination of these programs.

[0114] Numbered embodiment 16 includes a method of validating a microbial signature, comprising: receiving a first one or more microbial feature sets of a first biological sample from a first subject having a disease determined by non-specific interaction of the first one or more probe sets with one or more nucleic acid molecules of the first biological sample; training a predictive model using the first one or more microbial feature sets of the first biological sample and the disease of the first subject, thereby generating a trained predictive model; receiving a second one or more microbial feature sets of a second biological sample from a second subject having the disease; and validating the first one or more microbial feature sets by comparing a predicted disease provided by the trained predictive model with a disease of the second subject, wherein a predicted disease provided by the trained predictive model is generated when the second one or more microbial feature sets are provided as inputs to the trained predictive model. Numbered embodiment 17 includes the method of embodiment 16, wherein the biological sample comprises a tissue, a liquid biopsy, or a combination of such samples. Numbered embodiment 18 includes the method of embodiment 16 or embodiment 17, wherein the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. Numbered embodiment 19 includes the method of any one of embodiments 16-18, wherein the first and second subjects comprise human or non-human mammalian subjects. Numbered embodiment 20 includes the method of any one of embodiments 16-19, wherein the first set of one or more microbial features comprises a taxonomic assignment and abundance of a first set of microbial sequencing reads, and the second set of one or more microbial features comprises a taxonomic assignment and abundance of a second set of microbial sequencing reads. Numbered embodiment 21 includes the method of any one of embodiments 16-20, further comprising removing one or more contaminant microbial features from the first set of one or more microbial features, the second set of one or more microbial features, or a combination thereof. Numbered embodiment 22 includes the method of any one of embodiments 16 to 21, wherein removing one or more contaminant microbial characteristics is accomplished by in silico decontamination, experimental control, or a combination thereof.Numbered embodiment 23 includes the method of any one of embodiments 16-22, wherein the first subject and the second subject comprise a human or non-human mammalian subject. Numbered embodiment 24 includes the method of any one of embodiments 16-23, wherein the disease of the first subject or the disease of the second subject comprises cancer, a non-cancerous disease, or a combination thereof. Numbered embodiment 25 includes the method of any one of embodiments 16-24, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary renal cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasms diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. Numbered embodiment 26 includes the method of any one of embodiments 16-25, wherein the one or more microbial features are from a virus, a bacterium, a fungus, an archaea, or any combination thereof. Numbered embodiment 27 includes the method of any one of embodiments 16-26, wherein the first one or more probe sets or the second one or more probe sets comprise multiplexed oligonucleotide probes target mammalian genomic regions. Numbered embodiment 28 includes the method of any one of embodiments 16-27, wherein the first one or more microbial feature sets and the second one or more microbial feature sets comprise enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof.Numbered embodiment 29 includes the method of any one of embodiments 16-28, wherein the first one or more microbial feature sets or the second one or more microbial feature sets are determined by sequencing one or more nucleic acid molecules bound to the first one or more probe sets or the second one or more probe sets, thereby generating one or more sequencing reads, mapping the one or more sequencing reads to a genome database to identify one or more non-human sequencing reads, and determining the first one or more microbial feature sets or the second one or more microbial feature sets from the one or more non-human sequencing reads. Numbered embodiment 30 includes the method of any one of embodiments 16-29, wherein the first one or more probe sets or the second one or more probe sets comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. Numbered embodiment 31 includes the method of any one of embodiments 16-30, wherein the one or more microbial features of the second biological sample are determined by sequencing enriched or non-enriched microbial nucleic acid molecules of the second biological sample. Numbered embodiment 32 includes the method of any one of embodiments 16 to 31, wherein the enriched microbial nucleic acid molecules are generated by exposing one or more nucleic acid molecules of the second biological sample to a second one or more probe sets, wherein the second one or more probe sets non-specifically bind to the one or more microbial nucleic acid molecules of the second biological sample.

[0115] Numbered embodiment 33 includes a method comprising exposing a biological sample of a first subject having a first disease to one or more probes, where the one or more probes non-specifically bind to one or more nucleic acid molecules of the biological sample; sequencing the one or more nucleic acid molecules bound to the one or more probes, thereby generating one or more sequencing reads; mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads; and generating a predictive model for predicting a second disease of a second subject, where the predictive model is trained using one or more microbial signatures of the one or more non-human sequencing reads and the first disease of the first subject. Numbered embodiment 34 includes a method of embodiment 33, where the biological sample includes a tissue, a liquid biopsy, or any combination of those samples. Numbered embodiment 35 includes a method of embodiment 33 or embodiment 34, where the one or more microbial signatures include taxonomic assignments and abundances of the one or more non-human sequencing reads. Numbered embodiment 36 includes the method of any one of embodiments 33-35, further comprising removing one or more contaminant microbial signatures from the one or more microbial signatures prior to training the predictive model. Numbered embodiment 37 includes the method of any one of embodiments 33-36, wherein removing one or more contaminant microbial signatures is accomplished by in silico decontamination, experimental control, or a combination thereof. Numbered embodiment 38 includes the method of any one of embodiments 33-37, wherein the first subject and the second subject comprise a human or non-human mammalian subject. Numbered embodiment 39 includes the method of any one of embodiments 33-38, wherein the one or more nucleic acids comprise one or more human nucleic acid molecules, non-human nucleic acid molecules, or a combination thereof. Numbered embodiment 40 includes the method of any one of embodiments 33-39, wherein the one or more nucleic acids comprise one or more human nucleic acid molecules, non-human nucleic acid molecules, or a combination thereof, wherein the non-human nucleic acid molecules are derived from a virus, a bacterium, a fungus, an archaea, or any combination thereof.Numbered embodiment 41 includes the method of any one of embodiments 33-40, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian nucleic acid molecules. Numbered embodiment 42 includes the method of any one of embodiments 33-41, wherein the one or more sequencing reads comprise sequencing reads of enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 43 includes the method of any one of embodiments 33-42, wherein the genomic database is a human genome database. Numbered embodiment 44 includes the method of any one of embodiments 33-43, wherein the predictive model is configured to predict a subject's response to chemotherapy, immunotherapy, neoadjuvant therapy, or any combination of those therapies administered to treat the disease. Numbered embodiment 45 includes the method of any one of embodiments 33-44, wherein the first disease and the second disease comprise cancer, non-cancerous disease, or a combination thereof. Numbered embodiment 46 includes the method of any one of embodiments 33-45, wherein the cancer comprises acute myeloid leukemia, adrenal cortical carcinoma, urothelial carcinoma of the bladder, brain low-grade glioma, invasive carcinoma of the breast, squamous cell carcinoma of the cervix and adenocarcinoma of the cervix, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe cell of the kidney, clear cell carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, cutaneous melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thymoma, thyroid carcinoma, uterine carcinosarcoma, endometrial carcinoma of the uterine corpus, or uveal melanoma. Numbered embodiment 47 includes the method of any one of embodiments 33-46, wherein the predictive model is configured to identify and remove one or more contaminant microbial characteristics while selectively retaining one or more non-contaminant microbial characteristics. Numbered embodiment 48 includes the method of any one of embodiments 33-47, wherein the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.Numbered embodiment 49 includes the method of any one of embodiments 33-48, wherein identifying comprises computationally filtering one or more sequencing reads using bowtie2, Kraken, or a combination of these programs. Numbered embodiment 50 includes the method of any one of embodiments 33-49, wherein the predictive model comprises a machine learning model. Numbered embodiment 51 includes the method of any one of embodiments 33-50, wherein the machine learning model comprises one or more machine learning models, or an ensemble of machine learning models. Numbered embodiment 52 includes the method of any one of embodiments 33-51, wherein one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules.

[0116] Numbered embodiment 53 includes a method comprising exposing a biological sample of a subject having a disease to one or more probes, where the one or more probes non-specifically bind to one or more nucleic acid molecules of the biological sample; identifying one or more sequencing reads of the one or more nucleic acid molecules bound to the one or more probes; mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more sequencing reads; and identifying one or more microbial features of the one or more non-human sequencing reads to classify the disease of the subject. Numbered embodiment 54 includes the method of embodiment 53, where the biological sample comprises a tissue, a liquid biopsy, or any combination of such samples. Numbered embodiment 55 includes the method of embodiment 53 or embodiment 54, where the one or more microbial features comprise a taxonomic assignment and abundance of the non-human sequencing reads. Numbered embodiment 56 includes the method of any one of embodiments 53 to 55, further comprising removing one or more contaminant microbial features of the taxonomic assignment and abundance, thereby generating one or more decontaminated microbial features. Numbered embodiment 57 includes the method of any one of embodiments 53 to 56, wherein the subject comprises a human or non-human mammalian subject. Numbered embodiment 58 includes the method of any one of embodiments 53 to 57, wherein the disease comprises cancer, a non-cancer disease, or a combination thereof. Numbered embodiment 59 includes the method of any one of embodiments 53-58, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary renal cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasms diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.Numbered embodiment 60 includes the method of any one of embodiments 53 to 59, wherein the one or more microbial features are from viruses, bacteria, fungi, archaea, or any combination of non-mammalian life domains thereof. Numbered embodiment 61 includes the method of any one of embodiments 53 to 60, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions. Numbered embodiment 62 includes the method of any one of embodiments 53 to 61, wherein the one or more sequencing reads comprise sequencing reads of enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 63 includes the method of any one of embodiments 53 to 62, wherein the genome database comprises a human genome database. Numbered embodiment 64 includes the method of any one of embodiments 53 to 63, wherein the one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. Numbered embodiment 65 includes the method of any one of embodiments 53 to 64, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian nucleic acid molecules. Numbered embodiment 66 includes the method of any one of embodiments 52 to 65, wherein the mapping comprises filtering the one or more sequencing reads with bowtie2, Kraken, or a combination of these programs.

[0117] Numbered embodiment 67 includes a system including one or more processors and a non-transitory computer readable storage medium including software, the software including executable instructions that, as a result of execution, cause the one or more processors of the computer system to receive one or more nucleic acid molecule sequencing reads of a biological sample of a subject, the subject having a disease, the one or more nucleic acid molecule sequencing reads being obtained from one or more nucleic acid molecules enriched by one or more probes exposed to the biological sample of the subject, map the one or more nucleic acid molecule sequencing reads to a human genome database, thereby identifying one or more non-human sequencing reads of the one or more nucleic acid molecule sequencing reads, and identifying one or more microbial features of the one or more non-human sequencing reads to classify the disease of the subject. Numbered embodiment 68 includes the system of embodiment 67, wherein the biological sample includes tissue, liquid biopsy, or any combination of such samples. Numbered embodiment 69 includes the system of embodiment 67 or embodiment 68, wherein the one or more microbial features include taxonomic assignment and abundance of the one or more non-human sequencing reads. Numbered embodiment 70 includes the system of any one of embodiments 67-69, further comprising removing one or more contaminant microbial signatures of taxonomic assignment and abundance, thereby generating one or more decontaminated microbial signatures. Numbered embodiment 71 includes the system of any one of embodiments 67-70, wherein removing one or more contaminant microbial signatures is accomplished by in silico decontamination, experimental control, or a combination thereof. Numbered embodiment 72 includes the system of any one of embodiments 67-71, wherein the subject comprises a human or non-human mammalian subject. Numbered embodiment 73 includes the system of any one of embodiments 67-72, wherein the disease comprises cancer, a non-cancer disease, or a combination thereof.Numbered embodiment 74 includes the system of any one of embodiments 67-73, wherein the cancer includes acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary renal cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. Numbered embodiment 75 includes the system of any one of embodiments 67-74, wherein the one or more microbial features are from viruses, bacteria, fungi, archaea, or any combination of non-mammalian life domains thereof. Numbered embodiment 76 includes the system of any one of embodiments 67-75, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions. Numbered embodiment 77 includes the system of any one of embodiments 67-76, wherein the one or more nucleic acid molecule sequencing reads comprise sequencing reads of enriched populations of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 78 includes the system of any one of embodiments 67-77, wherein the one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules. Numbered embodiment 79 includes the system of any one of embodiments 67-78, wherein mapping the one or more nucleic acid molecular sequencing reads includes filtering the one or more nucleic acid molecular sequencing reads with bowtie2, Kraken, or a combination of these programs. Numbered embodiment 80 includes the system of any one of embodiments 67-79, wherein the software further includes generating a predictive model, the predictive model being trained using the one or more microbial signatures and the disease of interest.Numbered embodiment 81 includes the system of any one of embodiments 67-80, wherein the predictive model comprises one or more machine learning models. Numbered embodiment 82 includes the system of any one of embodiments 67-81, wherein the predictive model comprises an ensemble of one or more machine learning models. Numbered embodiment 83 includes the system of any one of embodiments 67-82, wherein the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. Numbered embodiment 84 includes the system of any one of embodiments 67-83, wherein the predictive model is configured to predict a subject's response to chemotherapy, immunotherapy, neoadjuvant therapy, or any combination of those therapies administered to treat the disease. EXAMPLES

[0118] Example 1: Non-specific hybridization of microorganisms in enriched biological samples: Incubation of biological samples with probes targeting gene segments indicative of colon cancer progression showed non-specific hybridization of cell-free microbial DNA. Biological samples (cell-free DNA) from 11 colon cancer patients were exposed to hybridization probes targeting 226 genes involved in CRC progression. As shown in Figure 2A, the nucleic acid molecules enriched by the hybridization probes were sequenced to generate both human and non-human sequencing reads (raw sequencing data from publicly available sources: Clonal evolution and resistance to EGFR blockade in the blood of colorectal cancer patients. Nature medicine, 21(7), PMID 26151329; https: / / www.ncbi.nlm.nih.gov / bioproject / 285189). The sequencing reads were then mapped to a human genome library to remove human somatic nucleic acid molecules. The results of the reads before and after human filtering and / or mapping are shown in Figure 2A. The remaining sequencing reads were then mapped to a reference microbial database (web of life) to determine the genus classification of the sequencing reads, of which the top 20 most abundant genera are shown in Figure 2B. From Figure 2B, the associated genera of the microorganisms present and the total reads of the identified genera can be seen. From this example, it can be seen that microbial nucleic acid molecules bind non-specifically to hybridization probes intended to enrich samples of human somatic nucleic acid molecules (e.g., cell-free DNA, cell-free RNA, DNA, RNA, etc.). Furthermore, it was observed that the microbial enrichment in the targeted hybridization probes was 10 times lower than in a typical shotgun metagenomic dataset at the same read depth. Therefore, we investigated whether this smaller set of enriched genera is biologically relevant, as shown in Example 2.

[0119] Example 2: Training and validation of predictive models using non-specifically enriched microbial features To determine whether the microbial genera identified in Example 1 are associated with the presence of colorectal cancer (CRC) (e.g., the diagnostic, prognostic, and / or screening capabilities of the microbial genera), a predictive model was trained and validated for the top 20 abundant genera in Figure 2B.

[0120] Cell-free DNA biosamples from 241 healthy and 26 colon cancer patients were analyzed by low-pass whole genome sequencing (~20 million reads / sample, publicly available sequencing data from PMID 31142840, https: / / ega-archive.org / datasets / EGAD00001005339). The resulting sequencing reads were in silico filtered to remove human reads. The resulting non-human reads were taxonomically assigned as described herein, and the sample-specific genera and associated abundances were used to train a cancer vs. health classifier that was intentionally constrained to use only the abundances of the 20 genera listed in Figure 2B. The receiver operating characteristic curve of the resulting trained predictive model and the corresponding area under the curve can be seen in Figure 3A. Notably, the top 20 microbial genera features used to train the predictive model showed an area under the curve of 0.987, indicating that the top 20 microbial features may serve as suitable diagnostic indicators for determining the presence of colon cancer in patients. The feature importance of the top 20 microbial genera used to train the predictive model can be seen in Figure 3B.

[0121] Example 3: Comparison of diagnostic power of non-specific enrichment microbial signatures across cancer types The 20 microbial features used to generate the predictive model described in Example 2 were analyzed to determine whether they could also provide cancer type diagnosis, prognosis, screening, or any combination of these capabilities. Publicly available cell-free DNA sequencing data (low-pass whole genome sequencing data from PMID 31142840) from seven cancer types (colon, bile duct, breast, stomach, lung, ovarian, and pancreatic cancer) were processed to remove human sequencing reads. The resulting non-human reads were taxonomically assigned as described herein, and sample-specific genera and associated abundances were used to train a colon cancer vs. other cancer classifier that was intentionally constrained to use only the abundances of the 20 genera listed in Figure 2B. Two sets of predictive models were generated, the first set of predictive models trained on the top 20 microbial features in Figure 2B, and the second set of predictive models trained on all taxonomically assigned microbial features of the DNA sequencing data without the mapped microbial cells. Figure 3C shows the resulting performance of the machine learning model area under the curve for each predictive model trained on microbial cell-free DNA sequencing data for a particular cancer type. From Figure 3C, when distinguishing between colorectal cancer and different cancer types, predictive models trained on the top 20 microbial features performed with an average area under the curve of 0.8 or more. Despite using only 20 features, these models performed surprisingly well compared to predictive models trained on all taxonomically assigned microbial features (3,107 features, of which an average of 692 microbial features were used in the "all" feature model), which showed an average area under the receiver operating characteristic curve of more than 0.88, as seen in Figure 3C. From these results, it can be understood that microbial cell-free nucleic acid molecules enriched and identified through non-specific interactions with mammalian targeted hybridization enrichment probes provide diagnostic capabilities to distinguish cancer types.

Claims

1. 1. A method for identifying a microbial signature for determining disease in a subject, the method comprising: (a) exposing the subject's biological sample to one or more probes, wherein the one or more probes non-specifically bind to one or more nucleic acid molecules in the biological sample; (b) obtaining a first sequencing read set of the one or more nucleic acid molecules bound to the one or more probes; (c) identifying a second sequencing read set within the first sequencing read set, wherein the second sequencing read set comprises non-human sequencing reads obtained through non-specific hybridization; and (d) identifying one or more microbial signatures for determining the disease in the subject from the second sequencing read set.

2. The method of claim 1, further comprising generating taxonomic assignments and abundances for the second sequencing read set.

3. The method of claim 1, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions.

4. The method described in claim 1, wherein identifying step (c) includes comparing the second sequencing read set with a genome database.

5. The method of claim 1, wherein the one or more probes comprise multiplexed oligonucleotide probes that non-specifically bind to one or more microbial nucleic acid molecules.

6. A method for verifying a microbial signature, comprising: (a) receiving a first one or more microbial feature sets of a first biological sample from a first subject having a disease determined by non-specific interaction of one or more probes with one or more nucleic acid molecules of the first biological sample; (b) training a predictive model using the first one or more microbial feature sets of the first biological sample and the disease of the first subject, thereby generating a trained predictive model; (c) receiving a second set of one or more microbial signatures of a second biological sample from a second subject with the disease; and (d) validating the first one or more microbial feature sets by comparing a predicted disease provided by the trained predictive model with the disease of the second subject, wherein the predicted disease provided by the trained predictive model is produced when the second one or more microbial feature sets are provided as input to the trained predictive model.

7. The method of claim 6, wherein the one or more probes, or the second one or more probe sets, comprise multiplexed oligonucleotide probes targeting mammalian genomic regions.

8. The first set of one or more microbial characteristics or the second set of one or more microbial characteristics: (a) sequencing one or more nucleic acid molecules bound to a first one or more probe sets or a second one or more probe sets, thereby generating one or more sequencing reads; (b) mapping the one or more sequencing reads to a human genome database to identify one or more non-human sequencing reads; and (c) determining a first one or more microbial feature sets or a second one or more microbial feature sets from the one or more non-human sequencing reads. (a) exposing a biological sample of a first subject having a first disease to one or more probes, wherein the one or more probes non-specifically bind to one or more nucleic acid molecules of the biological sample; (b) sequencing the one or more nucleic acid molecules bound to the one or more probes, thereby generating one or more sequencing reads; (c) mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads; and (d) generating a predictive model for predicting a second disease in a second subject, wherein the predictive model is trained using one or more microbial features of the one or more non-human sequencing reads and the first disease in the first subject.

10. The method described in claim 9, wherein the predictive model is configured to predict a subject's response to chemotherapy, immunotherapy, neoadjuvant therapy, or any combination of these therapies administered to treat a disease.

11. The method described in claim 9, wherein the predictive model includes a machine learning model.

12. (a) exposing a biological sample from a subject having a disease to one or more probes, wherein the one or more probes non-specifically bind to one or more nucleic acid molecules in the biological sample; (b) identifying one or more sequencing reads of the one or more nucleic acid molecules bound to the one or more probes; and (c) mapping the one or more sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more sequencing reads; and (d) identifying one or more microbial signatures in the one or more non-human sequencing reads to classify disease in the subject.

13. The method of claim 12, wherein the one or more probes comprise multiplexed oligonucleotide probes targeting mammalian genomic regions.

14. The method of claim 12, wherein the one or more probes comprise a multiplexed oligonucleotide probe targeting a mammalian nucleic acid molecule. (a) one or more processors; (b) a non-transitory computer-readable storage medium containing software, the software causing the one or more processors of the computer system to: (i) receiving one or more nucleic acid molecule sequencing reads of a biological sample from a subject, wherein the subject has a disease, and the one or more nucleic acid molecule sequencing reads are obtained from one or more nucleic acid molecules enriched by one or more probes exposed to the biological sample from the subject; (ii) mapping the one or more nucleic acid molecule sequencing reads to a genome database, thereby identifying one or more non-human sequencing reads of the one or more nucleic acid molecule sequencing reads; (iii) identifying one or more microbial signatures in the one or more non-human sequencing reads to classify a disease in the subject.