Meta-epigenomics-based disease diagnostics
Patent Information
- Application Number
- JP2024520848
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-08
- Filing Date
- 2022-10-07
- Publication Date
- 2025-10-15
AI Technical Summary
Existing disease diagnosis methods, particularly in liquid biopsies, are limited by the exclusion of non-mammalian epigenetic data, which can provide valuable disease-specific signatures, leading to reduced sensitivity and specificity in cancer detection.
A method that enriches and integrates cross-kingdom epigenetic data from mammalian, bacterial, fungal, and viral nucleic acids in tissue or liquid biopsy samples, using affinity targeting and sequencing to generate combined metaepigenomic signatures for disease diagnosis.
Enhances diagnostic sensitivity and specificity by combining mammalian and non-mammalian epigenetic information, allowing for accurate disease classification and prediction of cancer types, stages, and therapy responses.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 253,655, filed October 8, 2021, which is incorporated herein by reference. Summary of the Invention
[0002] The present disclosure provides methods for identifying disease-associated meta-epigenomic biomarkers and using these biomarkers to accurately diagnose specific diseases from tissue or liquid biopsy samples. Specifically, the present invention provides methods for enriching and integrating cross-kingdom epigenetic data from mammalian, bacterial, fungal, archaeal, and viral kingdoms from tissue or liquid biopsy samples, and methods for using this combined dataset to diagnose and classify diseases in mammalian subjects.
[0003] The inventive methods disclosed herein provide a means to discover disease diagnostic biomarkers from cross-kingdom nucleic acid analysis, where the biomarkers are specifically derived from epigenetic features contained within a mixed (i.e., multi-kingdom) population of nucleic acids. These epigenetic features can be, for example, common features shared by two or more taxonomic kingdoms, or they can be taxonomically distinct, non-overlapping epigenetic features that are analyzed independently and then combined to provide a cross-kingdom diagnostic signature.
[0004] Human DNA methylation-based biomarkers have long been the subject of academic and clinical research (see, for example, DNA Methylation and Complex Human Disease, Michael Neidhart, 2016, ISBN: 978-0-12-420194-1) and have been incorporated into several commercially available diagnostic assays that utilize the presence or absence of 5-methylcytosine (5mC) modified DNA disease-characteristic. For example, the only blood-based liquid biopsy assay with FDA approval for cancer diagnosis is Epigenomics' Epi proColon colon cancer screening assay. This is a PCR assay for the qualitative detection of methylated Septin9 ctDNA isolated from 3.5 ml of patient plasma (methylation of specific CpG motifs in the promoter region of the SEPT9_v2 transcript has been associated with colon cancer but not healthy tissue). Specifically, the Epigenomics assay utilizes bisulfite treatment of isolated cfDNA and methylation-specific primers to detect the presence of methylated Septin9. More recently, Grail Inc. has used differential DNA methylation of genomic CpG sites to distinguish between different cancers and between cancer and non-cancer samples. GRAIL has set an ambitious goal to accurately screen for over 50 unique cancer types from a single sample through targeted bisulfite sequencing of methylation patterns of cell-free circulating tumor DNA (ctDNA). DNA methylation-based biomarkers have been investigated in many disease areas, but may prove particularly useful in liquid biopsy-based cancer diagnostic methods as a means to determine which ctDNA fragments are truly tumor-derived. While most driver mutations in cancer genes (e.g., TP53, KRAS) are common among cancers regardless of their tissue of origin, CpG methylation profiles are highly specific to the tissue and tumor from which they originate, potentially enabling more accurate diagnosis of cancer.In addition, there are 28 million CpG sites across the human genome whose methylation status (methylated vs. unmethylated) may contain cancer-specific signatures, but canonical ctDNA mutations are copy number / genome restricted, thus limiting the sensitivity of detection. In these and other analyses utilizing mammalian DNA modifications, it is important to emphasize that these epigenetic analyses are performed with the deliberate exclusion of nucleic acid data from non-mammalian sources, which may concomitantly reveal disease-specific presence or abundance.
[0005] Similarly, while it is understood that microbial genomes carry epigenetic information in the form of heritable but enzymatically reversible chemical modifications of the genome's underlying polynucleotide sequence, the differences between mammalian and microbial DNA methylation have been used in the past as a means to separate prokaryotic DNA from mammalian DNA to improve the diagnostic sensitivity of assays focused on selected prokaryotic targets. For example, Schmidt et al. (US8288115 B2) teach the use of specific proteins (Toll-like receptor 9 (TLR9) and CpG-binding protein (CGBP)) to enrich for unmethylated prokaryotic DNA from samples containing both mammalian and non-mammalian DNA. Because unmethylated CpG sites are 20-fold more abundant in prokaryotic DNA than in mammalian DNA, physical enrichment of unmethylated CpG-containing DNA helps to limit the amount of mammalian DNA present in downstream molecular assays, specifically, assays using Schmidt et al.'s PCR-based analysis.
[0006] Similarly, Forsyth (US8927218 B2) teaches the use of specific microbial DNA methylation motifs and catalytically inactive restriction enzymes that can bind but not hydrolyze methylation-specific antibodies to enrich for prokaryotic sequences from a complex mixture of nucleic acids. Again, the intention is to physically separate prokaryotic sequences from non-prokaryotic sequences, thereby improving the detection limit in downstream analyses focused on the detection of selected prokaryotes.
[0007] Zhou et al. (WO2020 / 198664; PCT US2020 / 025425) teach methods for preparing sequencing libraries from cell-free DNA to facilitate "genomic and epigenomic profiling of the microbiome," but again, the objective is to separate mammalian from non-mammalian nucleic acid molecules such that most downstream sequencing reads are of microbial origin. Furthermore, while the method of Zhou et al. provides a means for preparing sequencing libraries that may be suitable for microbial epigenomic analysis, it does not teach the mode of epigenomic analysis or the epigenetic features that are analyzed.
[0008] In contrast to the aforementioned technical fields, where epigenetic features of exclusively mammalian or non-mammalian origin (but not both) are the subject of analysis, the method of the present invention utilizes and combines epigenetic data from taxonomically diverse life forms represented within a nucleic acid sample. Since microorganisms are increasingly implicated in mammalian disease processes and disease-specific mammalian epigenetic features have proven to be a robust source of diagnostic biomarkers, we reasoned that combining epigenetic content from both mammalian and microbial taxonomic sources within a nucleic acid sample would allow the creation of sensitive and specific "metaepigenomic" diagnostic signatures. In this way, we make a major departure from all existing technologies and create a new method for identifying disease diagnostic biomarkers.
[0009] Aspects disclosed herein provide a method for creating a diagnostic model for diagnosing disease in a subject based on a combination of mammalian and non-mammalian epigenetic information contained in a nucleic acid sample, the method comprising: (a) enriching one or more mammalian and non-mammalian nucleic acid molecules by affinity targeting of epigenetic features shared by both the one or more mammalian and non-mammalian nucleic acid molecules; (b) sequencing the enriched nucleic acid composition to generate sequencing reads; (c) filtering the sequencing reads with a genomic database build to isolate non-mammalian sequencing reads and generate a mammalian alignment file; and (d) analyzing the mammalian alignment file to generate a mammalian feature abundance table. (e) analyzing the non-mammalian sequencing reads to generate a non-mammalian feature abundance table; (f) combining the mammalian and non-mammalian feature abundance tables to generate a combined meta-epigenomic machine learning feature set; (g) training and testing a predictive model on the meta-epigenomic feature set to generate a trained predictive model; and (h) using the output of the trained predictive model to provide a diagnosis of the presence or absence of disease in the subject. In some embodiments, the nucleic acid sample may be derived from a tissue, a liquid biopsy sample, or any combination thereof. In some embodiments, the subject may include a human or a non-human mammal. In some embodiments, the nucleic acid may include a total population of DNA, RNA, cell-free DNA, cell-free RNA, ethosomal DNA, exosomal RNA, or any combination thereof.
[0010] In some embodiments, affinity targeting may include enriching epigenetic features of the shared nucleic acid. In some embodiments, the epigenetic features of the shared nucleic acid may include methylated CpG dinucleotide pairs. In some embodiments, the epigenetic features of the shared nucleic acid may include unmethylated CpG dinucleotide pairs. In some embodiments, the epigenetic features of the shared nucleic acid may include modified nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, and N6-methyladenine.
[0011] In some embodiments, the affinity targeting may include a specific affinity reagent. In some embodiments, the specific affinity reagent may include streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, or recombinant epigenetic proteins. In some embodiments, the recombinant epigenetic proteins may include epigenetic readers, writers, erasers, or any combination thereof. In some embodiments, the epigenetic readers may include recombinant methyl-CpG binding proteins Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or recombinant methyl-binding domains derived therefrom. In some embodiments, the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. In some embodiments, the epigenetic writers and erasers may be catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers may comprise an epitope tag. In some embodiments, the epitope tag may comprise an N-terminal or C-terminal 6× histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. In some embodiments, the molecular recognition motif may comprise birA or a sortase motif. In some embodiments, the nucleic acid composition may be concentrated by a solid support, which may comprise a complementary antibody covalently bound to the epitope tag. In some embodiments, the specific affinity reagent may comprise a region for recognizing and binding to the epigenetic feature. In some embodiments, the affinity targeting may comprise incubating the nucleic acid sample with a solid support comprising a plurality of immobilized affinity agents.In some embodiments, the plurality of immobilized affinity agents may comprise regions that will bind to epigenetic features. In some embodiments, the solid support may comprise magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymers, or any combination thereof. In some embodiments, the genome database may be a human genome database.
[0012] In some embodiments, the mammalian feature abundance table may include mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. In some embodiments, the mammalian feature abundance table may include mammalian functional gene and biochemical pathway abundance tables. In some embodiments, the non-mammalian feature abundance table may include microbial taxonomic assignments and a number of sequencing reads associated therewith. In some embodiments, the non-mammalian feature abundance table may include non-mammalian functional gene and biochemical pathway abundance tables. In some embodiments, the output of the trained predictive model may include an analysis of a combination of a mammalian feature set and a non-mammalian feature set. In some embodiments, the trained predictive model may be trained on a set of mammalian and non-mammalian epigenome abundances where feature abundances are known to be present or absent in a disease of interest. In some embodiments, the diagnostic model may utilize epigenome abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. In some embodiments, the diagnostic model may diagnose a category or tissue-specific location of a disease. In some embodiments, the diagnostic model may be used to diagnose one or more types of cancer in a subject. In some embodiments, the diagnostic model can be used to diagnose one or more subtypes of cancer in a subject. In some embodiments, the diagnostic model can be used to predict the stage of cancer in a subject and / or predict the prognosis of cancer in a subject. In some embodiments, the diagnostic model can be used to predict a subject's cancer therapy response. In some embodiments, the diagnostic model can be utilized to select the optimal therapy for a particular subject. In some embodiments, the diagnostic model can be utilized to longitudinally model the course of one or more cancers' response to therapy and then adjust the treatment regimen.
[0013] In some embodiments, the diagnostic model is capable of diagnosing one or more of the following: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, or uveal melanoma. In some embodiments, the diagnostic model can identify and remove certain non-human features as contaminants, called noise, while selectively retaining other non-human features, called signals. In some embodiments, the diagnostic model can be used to diagnose systemic lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), or sarcoidosis. In some embodiments, the liquid biopsy sample can include, but is not limited to, one or more of the following: plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, or exhaled breath condensate.
[0014] Embodiments disclosed herein provide a method for generating a diagnostic model for diagnosing disease in a subject based on a combination of mammalian and non-mammalian epigenetic information contained in a nucleic acid sample, the method comprising: (a) enriching one or more mammalian nucleic acid molecules by affinity targeting of epigenetic features present in the one or more mammalian nucleic acid molecules; (b) enriching one or more non-mammalian nucleic acid molecules by affinity targeting of epigenetic features present in the one or more non-mammalian nucleic acid molecules; (c) sequencing the enriched mammalian nucleic acid composition to generate sequencing reads; (d) sequencing the enriched non-mammalian nucleic acid composition to generate sequencing reads; and (e) substituting the mammalian sequencing reads into a genome database. (f) aligning the non-mammalian sequencing reads with the genome database construct to isolate the non-mammalian sequencing reads; (g) analyzing the mammalian alignment file to generate a mammalian feature abundance table; (h) analyzing the non-mammalian sequencing reads to generate a non-mammalian feature abundance table; (i) combining the mammalian and non-mammalian feature abundance tables to generate a combined meta-epigenomic machine learning feature set; (j) training and testing a predictive model on the meta-epigenomic feature set to generate a trained predictive model; and (k) using the output of the trained predictive model to provide a diagnosis of the presence or absence of a disease in the subject. In some embodiments, the nucleic acid sample may be derived from a tissue, a liquid biopsy sample, or any combination thereof. In some embodiments, the subject may include a human or a non-human mammal. In some embodiments, the nucleic acid may include a total population of DNA, RNA, cell-free DNA, cell-free RNA, ethosomal DNA, exosomal RNA, or any combination thereof.
[0015] In some embodiments, affinity targeting may include enriching mammalian and non-mammalian nucleic acid epigenetic features. In some embodiments, mammalian nucleic acid epigenetic features may include modified nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxycytosine, N4-acetylcytosine, and N6-methyladenine. In some embodiments, non-mammalian nucleic acid epigenetic features may include modified nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, 4-methylcytosine, N4-acetylcytosine, N6-methyladenine. In some embodiments, non-mammalian nucleic acid epigenetic features may include phosphorothioate-linked nucleotides.
[0016] In some embodiments, the affinity targeting may include a specific affinity reagent. In some embodiments, the specific affinity reagent may include streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, or recombinant epigenetic proteins. In some embodiments, the recombinant epigenetic proteins may include epigenetic readers, writers, erasers, or any combination thereof. In some embodiments, the epigenetic readers may include recombinant methyl-CpG binding proteins Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, DnaA, SeqA, MutHLS, Lrp, OxyR, Fur, HdfR, or recombinant methyl-binding domains derived therefrom. In some embodiments, the epigenetic reader may comprise recombinant zinc finger CXXC domain-containing proteins KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or recombinant CXXC domains derived therefrom. In some embodiments, the epigenetic reader may comprise microbial proteins Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or recombinant methyl-binding domains derived therefrom. In some embodiments, the epigenetic writers and erasers may be catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers may comprise epitope tags. In some embodiments, the epitope tag may comprise an N-terminal or C-terminal 6x histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. In some embodiments, the molecular recognition motif may comprise a birA or a sortase motif. In some embodiments, the nucleic acid composition may be concentrated by a solid support, which may comprise a complementary antibody covalently attached to the epitope tag.In some embodiments, the specific affinity reagent may include a region for recognizing and binding to the epigenetic feature. In some embodiments, the affinity targeting may include incubating the nucleic acid sample with a solid support comprising a plurality of immobilized affinity agents. In some embodiments, the plurality of immobilized affinity agents may include a region capable of binding to the epigenetic feature. In some embodiments, the solid support may include magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymers, or any combination thereof. In some embodiments, the genome database may be a human genome database.
[0017] In some embodiments, the mammalian feature abundance table may include mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. In some embodiments, the mammalian feature abundance table may include mammalian functional gene and biochemical pathway abundance tables. In some embodiments, the non-mammalian feature abundance table may include non-mammalian taxonomic assignments and a number of sequencing reads associated therewith. In some embodiments, the non-mammalian feature abundance table may include non-mammalian functional gene and biochemical pathway abundance tables. In some embodiments, the output of the trained predictive model may include an analysis of a combined set of mammalian and non-mammalian features. In some embodiments, the trained predictive model may be trained on a set of mammalian and non-mammalian epigenome abundances where feature abundances are known to be present or absent in a disease of interest. In some embodiments, the diagnostic model may utilize epigenome abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. In some embodiments, the diagnostic model may diagnose a category or tissue-specific location of a disease. In some embodiments, the diagnostic model may be used to diagnose one or more types of cancer in a subject. In some embodiments, the diagnostic model can be used to diagnose one or more subtypes of cancer in a subject. In some embodiments, the diagnostic model can be used to predict the stage of cancer in a subject and / or predict the prognosis of cancer in a subject. In some embodiments, the diagnostic model can be used to predict a subject's cancer therapy response. In some embodiments, the diagnostic model can be utilized to select the optimal therapy for a particular subject. In some embodiments, the diagnostic model can be utilized to longitudinally model the course of one or more cancers' response to therapy and then adjust the treatment regimen.
[0018] In some embodiments, the diagnostic model is capable of diagnosing one or more of the following: acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, or uveal melanoma. In some embodiments, the diagnostic model can identify and remove certain non-human features as contaminants, called noise, while selectively retaining other non-human features, called signals. In some embodiments, the diagnostic model can be used to diagnose systemic lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), or sarcoidosis. In some embodiments, the liquid biopsy sample can include, but is not limited to, one or more of the following: plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, or exhaled breath condensate.
[0019] Aspects of the disclosure provided herein include a method of creating a Feature Set for a disease of one or more subjects, the method comprising: (a) providing one or more mammalian and non-mammalian nucleic acid molecules of a biological sample of one or more subjects having a disease; (b) enriching one or more mammalian and non-mammalian nucleic acid molecules of the biological sample of the one or more subjects by affinity targeting of epigenetic features common to the one or more mammalian and non-mammalian nucleic acid molecules; (c) sequencing the enriched one or more mammalian and non-mammalian nucleic acid molecules to generate one or more mammalian and non-mammalian sequencing reads; (d) filtering the mammalian and non-mammalian sequencing reads to isolate non-mammalian sequencing reads, thereby generating a mammalian feature abundance; (e) analyzing the non-mammalian sequencing reads to generate a non-mammalian feature abundance; and (f) creating a Feature Set by combining the mammalian and non-mammalian feature abundances with the disease of the one or more subjects. In some embodiments, the epigenetic features include nucleic acid epigenetic features. In some embodiments, the biological sample comprises a tissue, a liquid biopsy sample, or any combination thereof. In some embodiments, the one or more subjects are human or non-human mammals. In some embodiments, the mammalian and non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, affinity targeting of nucleic acid epigenetic features comprises enriching for nucleic acid epigenetic features. In some embodiments, the shared nucleic acid epigenetic features comprise methylated CpG dinucleotide pairs, unmethylated CpG dinucleotide pairs, or any combination thereof. In some embodiments, the nucleic acid epigenetic features comprise the nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof.
[0020] In some embodiments, the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature. In some embodiments, the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, recombinant epigenetic proteins, or any combination thereof. In some embodiments, the recombinant epigenetic proteins comprise epigenetic readers, writers, erasers, or any combination thereof. In some embodiments, the epigenetic readers comprise recombinant methyl-CpG binding proteins Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or recombinant methyl-binding domains derived therefrom. In some embodiments, the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. In some embodiments, the epigenetic writers and erasers are catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers comprise an epitope tag. In some embodiments, the epitope tag comprises an N-terminal or C-terminal 6× histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. In some embodiments, the molecular recognition motif comprises a birA or a sortase motif. In some embodiments, the method further comprises enriching the mammalian and non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody to the epitope tag. In some embodiments, the immobilized complementary antibody is immobilized to the solid support by passive, electrostatic, covalent, or any combination thereof forces. In some embodiments, the specific affinity reagent comprises a region for recognizing and binding to an epigenetic feature.In some embodiments, affinity targeting comprises incubating mammalian and non-mammalian nucleic acid molecules with a solid support comprising a plurality of immobilized affinity agents, in some embodiments, the plurality of immobilized affinity agents are immobilized to the solid support by passive, electrostatic, covalent, or any combination thereof forces.
[0021] In some embodiments, the plurality of immobilized affinity agents comprises a region that will bind to an epigenetic feature, hi some embodiments, the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, a pH-sensitive polymer, or any combination thereof.
[0022] In some embodiments, the filtering comprises filtering the mammalian and non-mammalian sequencing reads against a genomic database, in some embodiments, the genomic database is a human genomic database.
[0023] In some embodiments, the mammalian feature abundance comprises a mammalian genomic coordinate or annotated genomic locus and a number of sequencing reads associated therewith. In some embodiments, the mammalian feature abundance comprises a mammalian functional gene and biochemical pathway abundance table. In some embodiments, the non-mammalian feature abundance comprises a non-mammalian taxonomic assignment and a number of sequencing reads associated therewith. In some embodiments, the non-mammalian feature abundance comprises a non-mammalian functional gene and biochemical pathway abundance table. In some embodiments, the liquid biopsy sample includes, but is not limited to, one or more of the following: plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, or exhaled breath condensate.
[0024] Aspects of the disclosure provided herein include, in some embodiments, a method of using the output of a predictive model to determine disease in a subject, the method comprising: (a) enriching one or more mammalian and non-mammalian nucleic acid molecules of a biological sample of a first set of subjects having a first disease and a second set of subjects having a second disease by affinity targeting of an epigenetic feature common to the one or more mammalian and non-mammalian nucleic acid molecules of the first and second sets of subjects; (b) sequencing the enriched one or more mammalian and non-mammalian nucleic acid molecules of the first and second sets of subjects to generate one or more mammalian and non-mammalian sequencing reads; and (c) sequencing the enriched one or more mammalian and non-mammalian nucleic acid molecules of the first and second sets of subjects to generate one or more mammalian and non-mammalian sequencing reads. (d) filtering the first and second sets of mammalian and non-mammalian sequencing reads to isolate non-mammalian sequencing reads, thereby generating a first and second set of mammalian feature abundances; (e) analyzing the first and second sets of non-mammalian sequencing reads to generate a first and second set of non-mammalian feature abundances; (e) training a predictive model on the first set of mammalian and non-mammalian feature abundances and the first disease of the first set of subjects, thereby generating a trained predictive model; and (f) receiving an output of the second disease of the second set of subjects using the second set of mammalian and non-mammalian feature abundances as inputs to the trained predictive model. In some embodiments, the first or second set of subjects includes one or more subjects. In some embodiments, the genomic database is a human genomic database. In some embodiments, the non-mammalian nucleic acid molecule includes a non-mammalian nucleic acid molecule. In some embodiments, the biological sample is derived from a tissue, a liquid biopsy sample, or any combination thereof. In some embodiments, the first or second set of subjects is a human or non-human mammal. In some embodiments, the first or second set of mammalian and non-mammalian nucleic acid molecules comprises a total population of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, affinity targeting comprises enriching the first and second sets of mammalian and non-mammalian nucleic acid epigenetic signatures.
[0025] In some embodiments, the first and second sets of mammalian and non-mammalian nucleic acid epigenetic features include modified nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, 5-carboxycytosine, N4-acetylcytosine, and N6-methyladenine. In some embodiments, the first and second sets of mammalian and non-mammalian nucleic acid epigenetic features include modified nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, 4-methylcytosine, N4-acetylcytosine, N6-methyladenine. In some embodiments, the first and second sets of mammalian and non-mammalian nucleic acid epigenetic features include phosphorothioate-linked nucleotides. In some embodiments, the affinity targeting includes a specific affinity reagent. In some embodiments, the specific affinity reagent includes streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, or recombinant epigenetic proteins.
[0026] In some embodiments, the recombinant epigenetic protein comprises an epigenetic reader, writer, eraser, or any combination thereof. In some embodiments, the epigenetic reader comprises a recombinant methyl-CpG binding protein Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, DnaA, SeqA, MutHLS, Lrp, OxyR, Fur, HdfR, or a recombinant methyl-binding domain derived therefrom. In some embodiments, the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. In some embodiments, the epigenetic reader comprises microbial proteins Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or recombinant methyl-binding domains derived therefrom. In some embodiments, the epigenetic writers and erasers are catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers comprise an epitope tag. In some embodiments, the epitope tag comprises an N-terminal or C-terminal 6× histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. In some embodiments, the molecular recognition motif comprises a birA or a sortase motif.
[0027] In some embodiments, the method further comprises concentrating the first or second mammalian or non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody to the epitope tag. In some embodiments, the complementary antibody is immobilized to the solid support by passive, electrostatic, covalent, or any combination thereof. In some embodiments, the specific affinity reagent comprises a region for recognizing and binding to the first or second mammalian or non-mammalian epigenetic feature. In some embodiments, the affinity targeting comprises incubating the first or second set of mammalian or non-mammalian nucleic acid molecules with a solid support comprising a plurality of immobilized affinity agents. In some embodiments, the affinity agents are immobilized by electrostatic, passive, covalent, or any combination thereof. In some embodiments, the plurality of immobilized affinity agents comprises a region that binds to the first or second set of mammalian or non-mammalian epigenetic features. In some embodiments, the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymer, or any combination thereof.
[0028] In some embodiments, the first or second set of mammalian feature abundances comprises mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. In some embodiments, the first or second set of mammalian feature abundances comprises mammalian functional gene and biochemical pathway abundance tables. In some embodiments, the first or second set of non-mammalian feature abundances comprises non-mammalian taxonomic assignments and a number of sequencing reads associated therewith. In some embodiments, the non-mammalian feature abundances comprise non-mammalian functional gene and biochemical pathway abundance tables. In some embodiments, the output of the trained predictive model comprises an analysis of the combined first and second sets of mammalian and non-mammalian feature abundances. In some embodiments, the input of the trained predictive model comprises epigenomic abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. In some embodiments, the first or second disease comprises a disease category or tissue-specific location. In some embodiments, the first or second disease further comprises one or more types of cancer, one or more subtypes of cancer, cancer stage, cancer prognosis, or any combination thereof.
[0029] In some embodiments, the trained predictive model is used to predict the cancer therapy response of the second set of subjects.In some embodiments, the trained predictive model is utilized to select the optimal therapy for the second set of subjects.In some embodiments, the trained predictive model is utilized to longitudinally model the course of the response of one or more cancers of the second set of subjects to therapy, and then adjust the treatment regimen.
[0030] In some embodiments, the first or second disease further comprises one or more of the following: acute myeloid leukemia, adrenal cortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, or uveal melanoma.
[0031] In some embodiments, the predictive model is configured to remove contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. In some embodiments, the first or second disease further comprises lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof. In some embodiments, the liquid biopsy sample comprises one or more of the following: plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, or exhaled breath condensate.
[0032] Aspects of the disclosure provided herein include, in some embodiments, a method of determining a disease in a subject, the method comprising: providing a biological sample from a subject; enriching one or more nucleic acid molecules from the biological sample by affinity targeting of epigenetic features common to the one or more nucleic acid molecules; sequencing the enriched one or more nucleic acid molecules to generate one or more nucleic acid molecule sequencing reads; and determining the disease in the subject as an output of a predictive model when the enriched one or more nucleic acid molecules are provided as input to the predictive model. In some embodiments, the one or more nucleic acid molecules comprise one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the non-cancerous disease comprises a non-cancerous disease of lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof. In some embodiments, the method further comprises filtering the one or more nucleic acid molecular sequencing reads to identify one or more non-mammalian sequencing reads and one or more mammalian sequencing reads. In some embodiments, the epigenetic feature comprises a nucleic acid epigenetic feature. In some embodiments, the epigenetic feature comprises a mammalian nucleic acid epigenetic feature or a non-mammalian nucleic acid epigenetic feature. In some embodiments, the non-mammalian nucleic acid epigenetic feature comprises phosphorothioate linked nucleotides.In some embodiments, the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. In some embodiments, the subject is a human or a non-human mammal. In some embodiments, the one or more mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the one or more non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the affinity targeting of the nucleic acid epigenetic feature comprises enriching the nucleic acid epigenetic feature. In some embodiments, the nucleic acid epigenetic feature comprises methylated CpG dinucleotide pairs, unmethylated CpG dinucleotide pairs, or a combination thereof. In some embodiments, the nucleic acid epigenetic feature comprises the nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof. In some embodiments, the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature. In some embodiments, the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, recombinant epigenetic proteins, or any combination thereof. In some embodiments, the recombinant epigenetic proteins comprise epigenetic readers, writers, erasers, or any combination thereof. In some embodiments, the epigenetic readers comprise recombinant methyl-CpG binding proteins Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or recombinant methyl-binding domains derived therefrom.In some embodiments, the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. In some embodiments, the epigenetic reader comprises a microbial protein Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or a recombinant methyl-binding domain derived therefrom. In some embodiments, the epigenetic writers and erasers are catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers comprise epitope tags. In some embodiments, the epitope tag comprises an N-terminal or C-terminal 6x histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. In some embodiments, the molecular recognition motif comprises a birA or a sortase motif. In some embodiments, the method further comprises enriching one or more mammalian and non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody to the epitope tag. In some embodiments, the specific affinity reagent comprises a region for recognizing and binding to the epigenetic feature. In some embodiments, the affinity targeting may comprise incubating the biological sample with a solid support comprising a plurality of immobilized affinity agents. In some embodiments, the plurality of immobilized affinity agents comprises a region that will bind to the epigenetic feature. In some embodiments, the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymer, or any combination thereof. In some embodiments, the filtering comprises filtering the one or more mammalian sequencing reads and the one or more non-mammalian sequencing reads against a genomic database, in some embodiments, the genomic database is a human genomic database.In some embodiments, the predictive model is trained on one or more mammalian, one or more non-mammalian, or combinations thereof features determined from one or more nucleic acid molecules of a biological sample of one or more subjects and a corresponding disease of one or more subjects. In some embodiments, the one or more mammalian features include mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. In some embodiments, the one or more mammalian features include mammalian functional gene and biochemical pathway abundances. In some embodiments, the one or more non-mammalian features include microbial taxonomic assignments and a number of sequencing reads associated therewith. In some embodiments, the one or more non-mammalian features include microbial functional gene and biochemical pathway abundances. In some embodiments, the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, when the biological sample is enriched for one or more nucleic acid molecules, the accuracy of the predictive model for determining disease is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% compared to when the biological sample is not enriched for one or more nucleic acid molecules. In some embodiments, the predictive model comprises an area under the curve of at least about 0.70, at least about 0.80, at least about 0.85, at least about 0.90, or at least about 0.95 when determining disease in a subject. In some embodiments, the output of the trained predictive model comprises an analysis of a combination of one or more mammalian feature abundances and one or more non-mammalian feature abundances. In some embodiments, the input of the trained predictive model comprises epigenome abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. In some embodiments, the predictive model is further trained on tissue-specific locations of disease. In some embodiments, the predictive model is further trained on cancer type, subtype, stage, prognosis, or any combination thereof.In some embodiments, the predictive model outputs a cancer type, subtype, stage, prognosis, or any combination thereof when provided with nucleic acid sequencing reads of a biological sample of a subject. In some embodiments, the predictive model outputs a cancer therapy response of the subject. In some embodiments, the trained predictive model outputs a therapy for the subject that results in at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% reduction in the cancer area of the subject. In some embodiments, the trained predictive model outputs a longitudinal model of the cancer of the subject in response to a therapy, an adjustment to a therapy to treat the cancer of the subject, or a combination thereof. In some embodiments, the predictive model removes contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. In some embodiments, concentrating reduces the total of one or more nucleic acid molecules by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99%.
[0033] Aspects of the disclosure provided herein include, in some embodiments, a method of training a predictive model, comprising providing a biological sample of one or more subjects having a disease, enriching the biological sample of the one or more subjects by affinity targeting of epigenetic features common to one or more nucleic acid molecules of the biological sample, sequencing the enriched one or more nucleic acid molecules to generate one or more nucleic acid molecule sequencing reads, and training a predictive model with the one or more nucleic acid molecule sequencing reads and one or more features of the disease of the one or more subjects. In some embodiments, the epigenetic features comprise mammalian epigenetic features or non-mammalian epigenetic features. In some embodiments, the one or more features comprise one or more disease features. In some embodiments, the trained predictive model determines the disease of another one or more subjects different from the one or more subjects when another nucleic acid sequencing read of the biological sample of the one or more subjects is provided to the trained predictive model. In some embodiments, the one or more nucleic acid molecules comprise one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. In some embodiments, the method further comprises filtering the one or more nucleic acid sequencing reads to identify one or more non-mammalian sequencing reads, one or more mammalian sequencing reads, or a combination thereof. In some embodiments, the epigenetic signature comprises a nucleic acid epigenetic signature. In some embodiments, the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. In some embodiments, the one or more subjects are human or non-human mammals. In some embodiments, the one or more mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the one or more non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, affinity targeting the nucleic acid epigenetic signature comprises enriching the nucleic acid epigenetic signature.In some embodiments, the nucleic acid epigenetic feature comprises a methylated CpG dinucleotide pair, an unmethylated CpG dinucleotide pair, or a combination thereof. In some embodiments, the nucleic acid epigenetic feature comprises the nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof. In some embodiments, the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature. In some embodiments, the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, recombinant epigenetic proteins, or any combination thereof. In some embodiments, the recombinant epigenetic proteins comprise epigenetic readers, writers, erasers, or any combination thereof. In some embodiments, the epigenetic reader comprises recombinant methyl-CpG binding proteins Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or recombinant methyl-binding domains derived therefrom. In some embodiments, the epigenetic reader comprises recombinant zinc finger CXXC domain-containing proteins KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or recombinant CXXC domains derived therefrom. In some embodiments, the epigenetic writers and erasers are catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers comprise epitope tags. In some embodiments, the epitope tag comprises an N-terminal or C-terminal 6x histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof, hi some embodiments, the molecular recognition motif comprises a birA or a sortase motif.In some embodiments, the method further comprises enriching the one or more mammalian nucleic acid molecules and the one or more non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody to the epitope tag. In some embodiments, the specific affinity reagent comprises a region for recognizing and binding to the epigenetic feature. In some embodiments, the affinity targeting may comprise incubating the biological sample with a solid support comprising a plurality of immobilized affinity agents. In some embodiments, the plurality of immobilized affinity agents comprises a region that will bind to the epigenetic feature. In some embodiments, the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymers, or any combination thereof. In some embodiments, the filtering comprises filtering the one or more mammalian and non-mammalian sequencing reads against a genome database. In some embodiments, the genome database is a human genome database. In some embodiments, the one or more features comprise one or more mammalian features, one or more non-mammalian features, or a combination thereof. In some embodiments, the one or more mammalian features comprise a mammalian genome coordinate or annotated genomic locus and a number of sequencing reads associated therewith. In some embodiments, the one or more mammalian features include mammalian functional gene and biochemical pathway abundances. In some embodiments, the one or more non-mammalian features include a microbial taxonomic assignment and a number of sequencing reads associated therewith. In some embodiments, the one or more non-mammalian features include microbial functional gene and biochemical pathway abundances. In some embodiments, the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the disease includes cancer or a non-cancerous disease.In some embodiments, when the biological sample is enriched for one or more nucleic acid molecules, the accuracy of the predictive model for determining disease is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% compared to when the biological sample is not enriched for one or more nucleic acid molecules. In some embodiments, the predictive model comprises an area under the curve of at least about 0.70, at least about 0.80, at least about 0.85, at least about 0.90, or at least about 0.95 when determining disease in one or more additional subjects. In some embodiments, the non-mammalian nucleic acid epigenetic signature comprises phosphorothioate-linked nucleotides. In some embodiments, the epigenetic reader comprises microbial protein Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or recombinant methyl-binding domains derived therefrom. In some embodiments, the output of the trained predictive model includes an analysis of a combination of one or more mammalian feature abundances and one or more non-mammalian feature abundances. In some embodiments, the input of the trained predictive model includes epigenomic abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. In some embodiments, the predictive model is further trained on tissue-specific location of disease. In some embodiments, the predictive model is further trained on cancer type, subtype, stage, prognosis, or any combination thereof. In some embodiments, the predictive model outputs a cancer type, subtype, stage, prognosis, or any combination thereof when provided with nucleic acid sequencing reads of one or more other subjects of the biological sample. In some embodiments, the trained predictive model outputs a cancer therapy response of one or more other subjects.In some embodiments, the trained model outputs a therapy for one or more other subjects that results in at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% reduction in the cancer area of the one or more other subjects. In some embodiments, the trained predictive model outputs a longitudinal model of the cancer in one or more other subjects in response to a therapy, an adjustment to a therapy to treat the cancer in the subject, or a combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the predictive model is configured to remove contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. In some embodiments, the non-cancerous disease comprises a non-cancerous disease of lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof. In some embodiments, enriching reduces the total of one or more nucleic acid molecules by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99%.
[0034] In some embodiments, aspects of the disclosure provided herein include a computer system for determining a disease in a subject, the computer system including one or more processors and a non-transitory computer-readable storage medium including software, the software including executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) receive one or more nucleic acid molecule sequencing reads of one or more nucleic acid molecules of a biological sample of a subject, the one or more nucleic acid molecules being enriched by affinity targeting of epigenetic features common to the one or more nucleic acid molecules; and (ii) determine the disease in the subject as an output of a predictive model when the one or more nucleic acid molecule sequencing reads are provided to a predictive model. In some embodiments, the one or more nucleic acid molecules include one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. In some embodiments, the disease includes cancer or a non-cancerous disease. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the non-cancerous disease comprises a non-cancerous disease of lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof. In some embodiments, the executable instructions further comprise filtering the one or more nucleic acid molecular sequencing reads to identify one or more non-mammalian sequencing reads and one or more mammalian sequencing reads. In some embodiments, the epigenetic signature comprises a nucleic acid epigenetic signature.In some embodiments, the epigenetic feature comprises a mammalian nucleic acid epigenetic feature or a non-mammalian nucleic acid epigenetic feature. In some embodiments, the non-mammalian nucleic acid epigenetic feature comprises phosphorothioate linked nucleotides. In some embodiments, the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. In some embodiments, the subject is a human or a non-human mammal. In some embodiments, the one or more mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the one or more non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, affinity targeting of the nucleic acid epigenetic feature comprises enriching the nucleic acid epigenetic feature. In some embodiments, the nucleic acid epigenetic feature comprises a methylated CpG dinucleotide pair, an unmethylated CpG dinucleotide pair, or a combination thereof. In some embodiments, the nucleic acid epigenetic feature comprises the nucleobases 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof. In some embodiments, the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature. In some embodiments, the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibodies, aptamers, recombinant epigenetic proteins, or any combination thereof. In some embodiments, the recombinant epigenetic proteins comprise epigenetic readers, writers, erasers, or any combination thereof. In some embodiments, the epigenetic readers comprise recombinant methyl-CpG binding proteins Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or recombinant methyl-binding domains derived therefrom.In some embodiments, the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. In some embodiments, the epigenetic reader comprises a microbial protein Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or a recombinant methyl-binding domain derived therefrom. In some embodiments, the epigenetic writers and erasers are catalytically inactive. In some embodiments, the epigenetic readers, writers, and erasers comprise epitope tags. In some embodiments, the epitope tag comprises an N-terminal or C-terminal 6x histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. In some embodiments, the molecular recognition motif comprises a birA or a sortase motif. In some embodiments, the executable instructions further comprise enriching one or more mammalian and non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody to the epitope tag. In some embodiments, the specific affinity reagent comprises a region for recognizing and binding to the epigenetic feature. In some embodiments, the affinity targeting comprises incubating the biological sample with a solid support comprising a plurality of immobilized affinity agents. In some embodiments, the plurality of immobilized affinity agents comprises a region that will bind to the epigenetic feature. In some embodiments, the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, a pH-sensitive polymer, or any combination thereof. In some embodiments, the filtering comprises filtering the one or more mammalian sequencing reads and the one or more non-mammalian sequencing reads against a genomic database, in some embodiments, the genomic database is a human genomic database.In some embodiments, the predictive model is trained on one or more mammalian, one or more non-mammalian, or combinations thereof features determined from one or more nucleic acid molecules of a biological sample of one or more subjects and a corresponding disease of one or more subjects. In some embodiments, the one or more mammalian features include mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. In some embodiments, the one or more mammalian features include mammalian functional gene and biochemical pathway abundances. In some embodiments, the one or more non-mammalian features include microbial taxonomic assignments and a number of sequencing reads associated therewith. In some embodiments, the one or more non-mammalian features include microbial functional gene and biochemical pathway abundances. In some embodiments, the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, when the biological sample is enriched for one or more nucleic acid molecules, the accuracy of the predictive model for determining disease is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% compared to when the biological sample is not enriched for one or more nucleic acid molecules. In some embodiments, the predictive model comprises an area under the curve of at least about 0.70, at least about 0.80, at least about 0.85, at least about 0.90, or at least about 0.95 when determining disease in a subject. In some embodiments, the output of the trained predictive model comprises an analysis of a combination of one or more mammalian feature abundances and one or more non-mammalian feature abundances. In some embodiments, the input of the trained predictive model comprises epigenome abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. In some embodiments, the predictive model is further trained on tissue-specific locations of disease. In some embodiments, the predictive model is further trained on cancer type, subtype, stage, prognosis, or any combination thereof.In some embodiments, the predictive model outputs a cancer type, subtype, stage, prognosis, or any combination thereof when provided with nucleic acid sequencing reads of a biological sample of a subject. In some embodiments, the predictive model outputs a cancer therapy response of the subject. In some embodiments, the trained predictive model outputs a therapy for the subject that results in at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% reduction in the cancer area of the subject. In some embodiments, the trained predictive model outputs a longitudinal model of the cancer of the subject in response to a therapy, an adjustment to a therapy to treat the cancer of the subject, or a combination thereof. In some embodiments, the predictive model removes contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. In some embodiments, the enriched nucleic acid comprises a reduction of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99% of one or more nucleic acid molecules prior to enrichment.
[0035] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief description of the drawings]
[0036] [Figure 1A] FIG. 1 shows a flow diagram of a meta-epigenomic workflow for generating disease classifications based on epigenetic features present within the mammalian, bacterial, archaeal, fungal, and viral domains of life, as described in some embodiments herein. [Figure 1B]FIG. 1 shows a flow diagram of a meta-epigenomic workflow for generating disease classifications based on epigenetic features present within the mammalian, bacterial, archaeal, fungal, and viral domains of life, as described in some embodiments herein. [Figure 1C] FIG. 1 shows a flow diagram of a meta-epigenomic workflow for generating disease classifications based on epigenetic features present within the mammalian, bacterial, archaeal, fungal, and viral domains of life, as described in some embodiments herein. [Figure 1D] FIG. 1 shows a flow diagram of a meta-epigenomic workflow for generating disease classifications based on epigenetic features present within the mammalian, bacterial, archaeal, fungal, and viral domains of life, as described in some embodiments herein. [Figure 1E] FIG. 1 shows a flow diagram of a meta-epigenomic workflow for generating disease classifications based on epigenetic features present within the mammalian, bacterial, archaeal, fungal, and viral domains of life, as described in some embodiments herein. [Figure 1F] FIG. 1 shows a flow diagram of a meta-epigenomic workflow for generating disease classifications based on epigenetic features present within the mammalian, bacterial, archaeal, fungal, and viral domains of life, as described in some embodiments herein. [Figure 2A] Exemplary mammalian nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 2A shows 5-methylcytosine (5mC). Figure 2B shows 5-hydroxymethylcytosine (5hmC). Figure 2C shows 5-formylcytosine (5fC). Figure 2D shows 5-carboxycytosine (5caC). Figure 2E shows N4-acetylcytosine (N4AcC), as described in some embodiments herein. [Figure 2B]Exemplary mammalian nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 2A shows 5-methylcytosine (5mC). Figure 2B shows 5-hydroxymethylcytosine (5hmC). Figure 2C shows 5-formylcytosine (5fC). Figure 2D shows 5-carboxycytosine (5caC). Figure 2E shows N4-acetylcytosine (N4AcC), as described in some embodiments herein. [Figure 2C] Exemplary mammalian nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 2A shows 5-methylcytosine (5mC). Figure 2B shows 5-hydroxymethylcytosine (5hmC). Figure 2C shows 5-formylcytosine (5fC). Figure 2D shows 5-carboxycytosine (5caC). Figure 2E shows N4-acetylcytosine (N4AcC), as described in some embodiments herein. [Figure 2D] Exemplary mammalian nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 2A shows 5-methylcytosine (5mC). Figure 2B shows 5-hydroxymethylcytosine (5hmC). Figure 2C shows 5-formylcytosine (5fC). Figure 2D shows 5-carboxycytosine (5caC). Figure 2E shows N4-acetylcytosine (N4AcC), as described in some embodiments herein. [Figure 2E] Exemplary mammalian nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 2A shows 5-methylcytosine (5mC). Figure 2B shows 5-hydroxymethylcytosine (5hmC). Figure 2C shows 5-formylcytosine (5fC). Figure 2D shows 5-carboxycytosine (5caC). Figure 2E shows N4-acetylcytosine (N4AcC), as described in some embodiments herein. [Figure 3A]Exemplary microbial nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 3A shows 6-methyladenine (6mA). Figure 3B shows 5-methylcytosine (5mC). Figure 3C shows 4-methylcytosine (4mC). Figure 3D shows N4-acetylcytosine (N4AcC). Figure 3E shows 5-hydroxymethylcytosine (5hmC), as described in some embodiments herein. [Figure 3B] Exemplary microbial nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 3A shows 6-methyladenine (6mA). Figure 3B shows 5-methylcytosine (5mC). Figure 3C shows 4-methylcytosine (4mC). Figure 3D shows N4-acetylcytosine (N4AcC). Figure 3E shows 5-hydroxymethylcytosine (5hmC), as described in some embodiments herein. [Figure 3C] Exemplary microbial nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 3A shows 6-methyladenine (6mA). Figure 3B shows 5-methylcytosine (5mC). Figure 3C shows 4-methylcytosine (4mC). Figure 3D shows N4-acetylcytosine (N4AcC). Figure 3E shows 5-hydroxymethylcytosine (5hmC), as described in some embodiments herein. [Figure 3D] Exemplary microbial nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 3A shows 6-methyladenine (6mA). Figure 3B shows 5-methylcytosine (5mC). Figure 3C shows 4-methylcytosine (4mC). Figure 3D shows N4-acetylcytosine (N4AcC). Figure 3E shows 5-hydroxymethylcytosine (5hmC), as described in some embodiments herein. [Figure 3E]Exemplary microbial nucleic acid modifications utilized in meta-epigenomic analysis according to the methods of the present invention are shown. Figure 3A shows 6-methyladenine (6mA). Figure 3B shows 5-methylcytosine (5mC). Figure 3C shows 4-methylcytosine (4mC). Figure 3D shows N4-acetylcytosine (N4AcC). Figure 3E shows 5-hydroxymethylcytosine (5hmC), as described in some embodiments herein. [Figure 4] 1 shows bacterial and archaeal phosphorothioate modifications utilized for meta-epigenome analysis by the methods of the present invention, as described in some embodiments herein. [Figure 5A] As described in some embodiments herein, experimental data are presented regarding the discovery of microbial epigenetic biomarkers and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine, an epigenetic feature previously considered to be a mammalian epigenetic feature. [Figure 5B] As described in some embodiments herein, experimental data are presented regarding the discovery of microbial epigenetic biomarkers and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine, an epigenetic feature previously considered to be a mammalian epigenetic feature. [Figure 5C] As described in some embodiments herein, experimental data are presented regarding the discovery of microbial epigenetic biomarkers and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine, an epigenetic feature previously considered to be a mammalian epigenetic feature. [Figure 5D] As described in some embodiments herein, experimental data are presented regarding the discovery of microbial epigenetic biomarkers and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine, an epigenetic feature previously considered to be a mammalian epigenetic feature. [Figure 5E] As described in some embodiments herein, experimental data are presented regarding the discovery of microbial epigenetic biomarkers and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine, an epigenetic feature previously considered to be a mammalian epigenetic feature. [Figure 5F] As described in some embodiments herein, experimental data are presented regarding the discovery of microbial epigenetic biomarkers and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine, an epigenetic feature previously considered to be a mammalian epigenetic feature. [Figure 6A] 1 presents experimental data for microbial epigenetic biomarker discovery and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine-based enrichment of microbial nucleic acids, as described in some embodiments herein. [Figure 6B] 1 presents experimental data for microbial epigenetic biomarker discovery and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine-based enrichment of microbial nucleic acids, as described in some embodiments herein. [Figure 6C] 1 presents experimental data for microbial epigenetic biomarker discovery and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine-based enrichment of microbial nucleic acids, as described in some embodiments herein. [Figure 6D] 1 presents experimental data for microbial epigenetic biomarker discovery and derived cancer diagnostic models utilizing 5-hydroxymethylcytosine-based enrichment of microbial nucleic acids, as described in some embodiments herein. [Figure 7] 1 illustrates a diagram of a system configured to perform, implement, and / or execute methods described in embodiments herein and elsewhere herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0037] Aspects of the disclosure provided herein may include methods of creating diagnostic models for diagnosing disease in a subject based on a combination of mammalian and non-mammalian epigenetic information (referred to herein as "metaepigenomic" information, data, features, signatures, or biomarkers) contained in a nucleic acid sample. In some cases, the non-mammalian epigenetic information may include bacterial, fungal, archaeal, viral, or any combination thereof epigenetic information. This may be accomplished in some embodiments by identifying both mammalian and non-mammalian nucleic acid molecules isolated via antibody- or non-antibody protein-based enrichment of genomic regions with one or more specific epigenetic marks, and then testing the utility of those enriched nucleic acids to distinguish between subjects with disease and those without the disease. In some embodiments, the identified meta-epigenomic biomarkers and their presence or abundance in a subject's sample can be used to assign a particular probability that (1) an individual has a particular disease, (2) an individual has a benign or malignant mass in a particular body site, (3) an individual has a particular type of benign or malignant mass, and / or (4) a disease is more or less likely to respond to a particular therapy. Other uses of such methods are reasonably conceivable and readily implementable by one of skill in the art.
[0038] The invention disclosed herein, in some embodiments, can use meta-epigenomic biomarkers derived from nucleic acids of mammalian and non-mammalian origin to diagnose a condition (i.e., cancer). In some embodiments, the disclosed invention may provide better clinical outcomes compared to typical pathology reports, as it does not require the inclusion of one or more of observed tissue architecture, cellular atypia, or other subjective measures traditionally used to diagnose cancer. In some embodiments, the disclosed methods can provide high sensitivity by utilizing sequence information drawn from all possible genomes in a sample, rather than limiting the analysis to cancer genomes that are often modified at very low frequencies in the background of "normal" human sources. In some embodiments, the methods disclosed herein can achieve such outcomes with either solid tissue or blood-derived samples, the latter requiring minimal sample preparation and being minimally invasive. In some embodiments, liquid biopsy-based assays can overcome challenges posed by circulating tumor DNA (ctDNA) assays, which often suffer from sensitivity issues due to cell-free DNA (cfDNA) derived from non-malignant human cells. In some embodiments, liquid biopsy-based meta-epigenomic assays can distinguish between cancer types that ctDNA assays typically cannot achieve because the most common cancer genomic abnormalities (e.g., TP53 mutations, KRAS mutations) are shared between cancer types. In some embodiments, the methods described can limit the size of the signatures that are predicted by those skilled in the art (e.g., regularized machine learning), and meta-epigenomic assays may become clinically available by using targeted assay panels for, for example, multiplex quantitative polymerase chain reaction (qPCR) and multiplex amplicon sequencing.
[0039] In some embodiments, the methods of the invention disclosed herein may include a method for generating a feature set for a disease of one or more subjects, as seen in Figure 1A. In some cases, the method may include (a) providing one or more mammalian and non-mammalian nucleic acid molecules of a biological sample of one or more subjects having a disease (e.g., cancer or a non-cancerous disease) 101, (b) isolating a total unfractionated nucleic acid composition 102, (c) enriching one or more mammalian and non-mammalian nucleic acid molecules of the biological sample of the one or more subjects via targeting common epigenetic features 103, (d) sequencing the enriched one or more mammalian and non-mammalian nucleic acid molecules 104, (e) filtering the enriched one or more mammalian nuclear sequencing reads 105, and (f) receiving one or more non-mammalian sequencing reads 108 from the results of the filtered one or more mammalian nuclear sequencing reads. (g) generating taxonomic or pathway assignments for the one or more enriched non-mammalian sequencing reads, thereby generating non-mammalian feature abundances 109; (h) decontaminating the one or more non-mammalian feature abundances 110; (i) aligning the enriched one or more mammalian nucleic acid sequencing reads, thereby generating a mammalian alignment file 106; (j) selecting mammalian feature abundances of the one or more enriched mammalian nucleic acid sequencing reads in the mammalian alignment file 107; and (k) creating a feature set for a disease of the one or more subjects by combining the one or more mammalian and non-mammalian feature abundances and the disease of the one or more subjects into a feature set. In some cases, the feature set may include a meta-epigenomic machine learning feature set 111. In some cases, the method may further include identifying disease-associated nucleic acid sequences of mammalian or non-mammalian nucleic acid molecules having epigenetic features in the dataset. In some cases, the identification of disease-associated nucleic acid sequences may be from subjects who have a disease (e.g., cancer or a non-cancerous disease) or from subjects who are healthy. In some cases, the disease state may include cancer, diabetes, etc., or any disease or disorder discussed elsewhere herein.In some embodiments, the enriched sequencing dataset may be obtained using next generation sequencing, long read sequencing (e.g., nanopore sequencing), or a combination thereof. In some embodiments, the enriched sequencing dataset 104 is obtained from affinity targeting of epigenetic features common to both mammalian and non-mammalian nucleic acid molecules by antibodies or non-antibody protein-based agents specific for the common epigenetic features 103, thereby isolating genomic regions of interest from a nucleic acid sample 102 from a biological sample containing nucleic acid sequences of mammalian and non-mammalian origin 101, as shown in FIG. 1A. In some embodiments, meta-epigenomic features present in the enriched population of nucleic acids 103 may be identified through a meta-epigenomic computational workflow 112, and the enriched mammalian sequencing reads may be computationally filtered 105 from all raw sequencing reads 104 via alignment to a mammalian reference genome using bowtie2 or Kraken or their equivalents to generate a mammalian alignment file. In some embodiments, the mammalian alignment files can be processed through an analysis pipeline 107 (such as MethylAction or MEDIPS) to identify genomic regions enriched via affinity targeting 103 of selected epigenetic features, thereby generating an output of selected mammalian features. In some embodiments, the resulting non-mammalian reads 108 can be taxonomically classified using bowtie2 or Kraken with a reference microbial database such as Web of Life 109. In some embodiments, the abundance of non-mammalian genes with specific epigenetic marks can be confirmed using the Web of Life Toolkit App (WolTka) or any equivalent thereof 109. In some embodiments, the identified non-mammalian reads 109 can be processed through a decontamination pipeline 110 to remove sequences derived from common non-mammalian contaminants to generate decontaminated non-mammalian features.In some embodiments, the decontaminated non-mammalian features 110 can be combined with the output of the mammalian analysis pipeline 107 to generate a meta-epigenomic feature set 111 that can serve as a training feature set for a predictive model.
[0040] In some embodiments, the disclosure provided herein may include a method for preparing separate mammalian and non-mammalian epigenomic analyses through sample partitioning and parallel isolation of nucleic acids based on distinct epigenetic features present in mammalian and non-mammalian domains, as seen in FIG. 1B. In some cases, the method may include (a) providing a biological sample including one or more mammalian and non-mammalian nucleic acid compositions 101, (b) isolating unfractionated nucleic acid compositions 102, (c) partitioning the isolated unfractionated nucleic acid compositions into one or more aliquots 113, (d) concentrating the mammalian and non-mammalian nucleic acid compositions of the one or more aliquots, thereby generating enriched mammalian and non-mammalian nucleic acid compositions (114, 155), and (e) converting the enriched mammalian and non-mammalian nucleic acid compositions into a feature set for a disease 112. In some cases, converting the enriched mammalian and non-mammalian nucleic acid molecular compositions into a feature set may include inputting the enriched sequencing reads into a meta-epigenome computational workflow 112 with a step of filtering mammalian reads 105. In some cases, samples of mammalian and non-mammalian nucleic acid molecules 102 may be physically divided 113 to facilitate separate analysis of mammalian and non-mammalian (microbial) epigenetic features. In some embodiments, mammalian epigenetic features 114 may be enriched by affinity targeting of epigenetic features with antibodies or non-antibody protein-based agents specific for the epigenetic features. In some embodiments, the distribution of epigenetic features across the mammalian genome may be ascertained by a specific sequencing method that may or may not utilize a first enrichment step, such as bisulfite sequencing, reduced representation bisulfite sequencing, oxidative bisulfite sequencing, ACE-seq, enzymatic methyl-seq (EM-seq), nanopore sequencing, or equivalents thereof. In some embodiments, non-mammalian epigenetic features 115 may be enriched by affinity targeting of epigenetic features with antibodies or non-antibody protein-based agents specific for the epigenetic features.In some embodiments, the distribution of epigenetic features across the non-mammalian genome in a sample can be ascertained by a specific sequencing method that may or may not utilize a first enrichment step, such as bisulfite sequencing, reduced representation bisulfite sequencing, oxidative bisulfite sequencing, ACE-seq, enzymatic methyl-seq (EM-seq), nanopore sequencing, or equivalents thereof. In some embodiments, the results of the parallel mammalian 114 and non-mammalian 115 epigenetic analyses are combined and input into a meta-epigenomic computational workflow 112 to generate a meta-epigenomic machine learning feature set.
[0041] In some embodiments, the disclosure provided herein may include a method of generating a feature set for a disease of a subject through sequential isolation of mammalian and non-mammalian nucleic acids, as seen in FIG. 1C. In some cases, the method may include (a) providing one or more biological samples of one or more subjects, where the biological samples include mammalian and non-mammalian nucleic acid compositions 101, (b) isolating unfractionated mammalian and non-mammalian nucleic acid compositions 102, (c) concentrating the unfractionated mammalian and non-mammalian nucleic acid compositions to separate mammalian nucleic acid compositions and residual compositions 114, (d) concentrating the residual compositions of the non-mammalian nucleic acid compositions, and (e) converting the mammalian and non-mammalian nucleic acid compositions to a feature set for a disease 112. In some cases, converting the mammalian and non-mammalian nucleic acid compositions to a feature set for a disease may include inputting the mammalian and non-mammalian sequencing reads determined by 114 and 115 (FIG. 1C) into a meta-epigenomic computational workflow 112 at element 104 (FIG. 1A). In some embodiments, mammalian and non-mammalian epigenetic features may be enriched from the same nucleic acid sample 102 in a sequential manner 116, as shown in FIG. 1C, where mammalian epigenetic features 114 are enriched by affinity targeting of the epigenetic features with antibodies or non-antibody protein-based agents specific for the epigenetic features, thereby generating a sample depleted of mammalian nucleic acid molecules with the targeted epigenetic marks, which can then serve as input for enrichment of non-mammalian epigenetic features 115. In some embodiments, the order of enrichment is reversed, with targeted non-mammalian epigenetic enrichment 115 preceding mammalian epigenetic enrichment 114. The output of this sequential epigenetic analysis 116 may then be input to a meta-epigenomic computational workflow 112 to generate a meta-epigenomic machine learning feature set.
[0042] In some aspects, the disclosure provided herein may include methods for training predictive models incorporating a meta-epigenomic analysis module to enable meta-epigenomic-based discovery of healthy, non-cancer (unhealthy), and cancer-associated non-mammalian signatures (FIG. 1D). In some embodiments, the inventive systems and methods disclosed herein may include (a) determining meta-epigenomic features of a sample via sequencing, and (b) generating a predictive model. In some embodiments, the sequencing method may include next-generation sequencing or long-read sequencing (e.g., nanopore sequencing), or a combination thereof. In some embodiments, the predictive model 121 may include training a predictive model 120 on a meta-epigenomic machine learning feature set described elsewhere herein. In some embodiments, the predictive model may include a regularized machine learning model. In some embodiments, the predictive model may include a linear regression, a logistic regression, a decision tree, a support vector machine (SVM), a naive Bayes, a k-nearest neighbor (kNN), a k-means, a random forest algorithm model, or any combination thereof.
[0043] Aspects of the present disclosure may include a method of training a predictive model to determine a disease of a subject, as seen in FIG. 1D. In some cases, the method may include (a) providing one or more nucleic acid samples from a healthy subject 117, a cancerous subject 118, a non-cancerous and unhealthy subject, or any combination thereof; (b) isolating an unfractionated nucleic acid composition from the one or more nucleic acid samples 102; (c) enriching one or more non-mammalian and mammalian nucleic acid molecules of the unfractionated nucleic acid composition by affinity targeting 103; (d) converting the one or more non-mammalian and mammalian nucleic acid molecules into one or more feature sets corresponding to a disease of one or more subjects 112; and (e) training a predictive model 120 with the one or more feature sets and the corresponding disease, thereby generating a trained predictive model 121 configured to determine a disease of the subject. In some cases, the determined characterization of the subject may include health 122, cancerous disease 123, or non-cancerous disease 124. In some cases, the determined characterization of the subject may include health 122, cancerous disease 123, or non-cancerous disease 124. In some embodiments, the predictive model may be trained 120 with a meta-epigenomic feature set 112 derived from nucleic acids 102 from a plurality of known healthy subjects 117, a plurality of known cancer subjects 118, and a plurality of non-cancer, unhealthy subjects 119, enriched by affinity targeting 103 of epigenetic features shared between mammalian and non-mammalian nucleic acid molecules present in the sample, as shown in FIG. 1D. In some embodiments, training the predictive model 120 to generate a trained predictive model 121 generates machine learning identified meta-epigenomic signatures for healthy subjects 122, subjects with cancer 123, and unhealthy subjects without cancer 124.
[0044] Aspects of the disclosure provided herein may include methods of individual mammalian and non-mammalian nucleic acid analysis to train predictive models to determine disease in a subject, as seen in FIG. 1E. In some cases, the method may include (a) providing one or more nucleic acid samples from healthy subjects 117, cancerous subjects 118, non-cancerous and unhealthy subjects, or any combination thereof; (b) isolating an unfractionated nucleic acid composition from the one or more nucleic acid samples 102; (c) dividing the unfractionated nucleic acid composition 113 into two or more aliquots (114, 115); (d) enriching a first subset of the two or more aliquots of one or more mammalian nucleic acids 114 and a second subset of the two or more aliquots of one or more non-mammalian nucleic acid molecules 115; (e) transforming the one or more non-mammalian and mammalian nucleic acid molecules into one or more feature sets corresponding to a disease of one or more subjects 112; and (e) training a predictive 120 model with the one or more feature sets and the corresponding disease, thereby generating a trained predictive model 121 configured to determine the disease of the subject. In some cases, the determined disease of the subject may include health 122, cancerous disease 123, or non-cancerous disease 124. In some aspects, the disclosure provided herein may include a method of training a predictive model on a meta-epigenomic feature set to enable meta-epigenomic-based discovery of healthy, non-cancer (unhealthy), and cancer-related non-mammalian signatures, where the separate epigenetic analyses of FIG. 1B are combined to form a combined meta-epigenomic feature set for training a predictive model 120. In some embodiments, the meta-epigenomic feature set 112 configured to train the predictive model 120 may be derived from nucleic acids 102 from a plurality of known healthy subjects 117, a plurality of known cancer subjects 118, and a plurality of non-cancer, unhealthy subjects 119, physically partitioned to facilitate parallel analysis of mammalian and non-mammalian epigenetic features, as shown in FIG. 1E.
[0045] Aspects of the disclosure provided herein may include a method of continuous mammalian and non-mammalian nucleic acid analysis to train a predictive model to determine a disease in a subject. In some cases, the method may include (a) providing one or more nucleic acid samples from a healthy subject 117, a cancerous subject 118, a non-cancerous and unhealthy subject, or any combination thereof; (b) isolating an unfractionated nucleic acid composition from the one or more nucleic acid samples 102; (c) performing continuous epigenetic analysis using the isolated unfractionated nucleic acid composition, thereby generating one or more non-mammalian and mammalian nucleic acid molecules; (d) converting the one or more non-mammalian and mammalian nucleic acid molecules into one or more feature sets corresponding to a disease in one or more subjects 112; and (e) training a predictive model 120 with the one or more feature sets and the corresponding disease, thereby generating a trained predictive model 121 configured to determine a disease in a subject. In some cases, as shown in FIG. 1C, the sequential epigenetic analysis 116 includes enriching the unfractionated nucleic acid composition to separate a mammalian nucleic acid composition and a residual composition 114, and enriching the residual composition for a non-mammalian nucleic acid composition 115. In some cases, the determined characterization of the subject may include health 122, cancerous disease 123, or non-cancerous disease 124. In some aspects, the disclosure provided herein may include methods for training predictive models to enable meta-epigenomic based discovery of health, non-cancer (unhealthy), and cancer-associated non-mammalian signatures, combining the sequential epigenetic analysis of FIG. 1C to form a combined meta-epigenomic feature set for machine learning. In some embodiments, the meta-epigenomic feature set 112 for training the predictive model 120 may be derived from nucleic acids 102 from a plurality of known healthy subjects 117, a plurality of known cancer subjects 118, and a plurality of non-cancer, unhealthy subjects 119 that have undergone sequential analysis of mammalian and non-mammalian epigenetic features, as shown in FIG. 1F.
[0046] In some embodiments, specific mammalian epigenetic features targeted for enrichment or direct sequencing analysis may include 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), or N4-acetylcytosine (N4AcC), as shown in Figure 2 (Figures 2A-2E, respectively).
[0047] In some embodiments, specific non-mammalian epigenetic features targeted for enrichment or direct sequencing analysis may include 6-methyladenosine (6mA), 5-methylcytosine (5mC), 4-methylcytosine (4mC), N4-acetylcytosine (N4AcC), or 5-hydroxymethylcytosine (5hmC), as shown in Figure 3 (Figures 3A-E, respectively).
[0048] In some embodiments, the specific non-mammalian epigenetic feature targeted for enrichment may include phosphorothioate nucleotide linkages, as shown in FIG. 4.
[0049]
[0023] Embodiments disclosed herein may provide a method for creating a predictive model for diagnosing disease in a subject based on a combination of mammalian and non-mammalian epigenetic information contained in a nucleic acid sample (Figure 1A), the method including (a) enriching one or more mammalian and non-mammalian nucleic acid molecules by affinity targeting of epigenetic features present in one or more mammalian and non-mammalian nucleic acid molecules 103, (b) sequencing the enriched nucleic acids through targeting of the epigenetic features 104, and computationally analyzing both the mammalian and non-mammalian sequencing reads from the dataset to generate a meta-epigenomic machine learning feature set 111 that is used to train a predictive model to generate a trained diagnostic model (Figure 1D).
[0050] An embodiment disclosed herein provides a method of training a predictive model (FIG. 1D), the method including: (a) providing one or more sequenced metaepigenome abundances 112 of one or more subjects as a training dataset (i); (b) providing one or more sequenced metaepigenome abundances 112 of one or more subjects as a test set (i); (c) training a predictive model with a sample ratio of training samples to validation samples of 60:40, respectively; and (d) evaluating the predictive accuracy of the predictive model.
[0051] In some embodiments, predictions made by the trained predictive model may include a machine learning signature indicative of a healthy subject, or a machine learning derived signature indicative of a subject with cancer, or a machine learning derived signature indicative of a subject with a disease other than cancer. In some embodiments, the trained predictive model may identify and remove one or more non-mammalian or non-microbial nucleic acids classified as noise, while selectively retaining one or more other non-mammalian or non-microbial sequences, referred to as the signal.
[0052] While the above steps illustrate each of the methods or sequences of actions according to the embodiments, one of ordinary skill in the art will recognize many variations based on the teachings provided herein. Steps may be completed in different orders. Steps may be added or omitted. Some of the steps may include sub-steps. Many of these steps may be repeated as often as beneficial.
[0053] One or more of the steps of each of the methods or sequences of operations may be performed by one or more processors or logic circuitry, such as programmable array logic for a field programmable gate array. The circuitry may be programmed to provide one or more of the steps of each of the methods or sequences of operations, and the program may include, for example, program instructions stored in a computer readable memory, or programmed steps of logic circuitry, such as programmable array logic or a field programmable gate array.
[0054] Predictive Model The disclosed methods and systems may utilize or access external capabilities of artificial intelligence, predictive models, and / or machine learning techniques to determine whether one or more subjects have cancer from a biological sample of each of the one or more subjects. In some cases, the artificial intelligence techniques may identify features of one or more nucleic acid molecular sequencing reads (e.g., non-mammalian and / or mammalian) that may predict cancer in one or more subjects. In some cases, the features may be used to train one or more predictive models described elsewhere herein. These features may be used to accurately predict a disease or disorder, as described elsewhere herein. In some cases, the disease or disorder may include a cancer or non-cancerous disease, as described elsewhere herein. Using such predictive models, algorithms, and / or machine learning techniques, health care providers (e.g., doctors, nurses, medical technicians, etc.) may make informed and accurate risk-based decisions, thereby improving early stage disease diagnosis, disease progression, and monitoring, treatment and / or therapeutic suggestions to treat the disease in a subject, or any combination thereof.
[0055] The methods and systems of the present disclosure can analyze the presence and abundance of mammalian nucleic acid molecules and / or non-mammalian nucleic acid molecules to determine one or more mammalian features and / or one or more non-mammalian features that may predict disease in one or more subjects. In some cases, the methods and systems described elsewhere herein can train a predictive model with one or more mammalian features, one or more non-mammalian features, and the corresponding disease in one or more subjects. In some cases, the trained predictive model can then be used to generate a likelihood (e.g., a prediction) of disease (e.g., cancer or cancerous disease) in one or more other subjects that are different from the one or more subjects utilized to train the predictive model. The learned predictive model can include an artificial intelligence-based model, such as a machine learning-based classifier, configured to process one or more nucleic acid molecule sequencing reads to generate a likelihood that the subject has the disease. The model may be trained using the presence or abundance of one or more mammalian and / or non-mammalian nucleic acid sequence reads generated from one or more nucleic acid molecules of biological samples from one or more cohorts of patients (e.g., cancer patients, patients with a non-cancerous disease, disease-free and cancer-free patients, cancer patients undergoing treatment for cancer, patients undergoing treatment for a non-cancerous disease, or a combination thereof). In some cases, the predictive model may be trained to provide treatment predictions for treating cancer for one or more patients that are not part of the training dataset of the predictive model. Such a predictive model may output treatment recommendations for one or more patients that are not part of the training dataset, when provided with an input of the presence of the patient and the abundance of one or more nucleic acid molecule sequencing reads obtained from the biological sample.
[0056] The predictive model may include one or more predictive models. The predictive model may include one or more machine learning algorithms. Examples of machine learning algorithms may include support vector machines (SVMs), naive Bayes classification, random forests, neural networks, deep neural networks (DNNs), recurrent neural networks (RNNs), deep RNNs, long short-term memory (LSTM) recurrent neural networks (RNNs), gated recurrent units (GRUs), gradient boosting machines, linear regression, k-nearest neighbors, k-means, decision trees, logistic regression, other supervised learning algorithms, or unsupervised machine learning models, or any combination thereof. The predictive model may be used for classification or regression. The model may include estimation of an ensemble model composed of multiple predictive models, and may utilize techniques such as gradient boosting in the construction of gradient boosting decision trees. The model may be trained using one or more training datasets corresponding to patient and / or subject data (e.g., patient medical history, family medical history, blood pressure, pulse rate, temperature, oxygen saturation, or any combination thereof), in addition to one or more nucleic acid sequencing reads generated from one or more nucleic acid molecules of a subject's biological sample, as described elsewhere herein.
[0057] The training dataset may be generated, for example, from one or more cohorts of patients with a diagnosis of a common clinical disease or disorder. The training dataset may include a set of one or more non-mammalian features, one or more mammalian features, or a combination thereof, in the form of the presence and / or abundance of one or more mammalian nucleic acid molecules and / or one or more non-mammalian nucleic acid molecules in a biological sample of one or more subjects. In some cases, the one or more mammalian nucleic acid molecules and / or one or more non-mammalian nucleic acid molecules may include enriched nucleic acid molecules as described elsewhere herein. The features may include the corresponding cancer diagnosis of one or more subjects for the one or more mammalian features and / or one or more non-mammalian features. In some cases, the features may include patient information such as the patient's age, the patient's medical history, other medical conditions, current or past medications, clinical risk scores, and time since last observation. For example, a set of features collected from a given patient at a given time point collectively serves as a signature, which may be indicative of a disease or disease state of the patient and / or subject at a time point.
[0058] The training data labels can include, for example, clinical outcomes such as the presence, absence, diagnosis, or prognosis of a disease (e.g., cancer or non-cancer disease) or disorder of a subject and / or patient. Clinical outcomes can include treatment efficacy (e.g., whether a subject is a positive responder to a cancer-based treatment).
[0059] The input features may be structured by aggregating the data into bins, or alternatively by using one-hot encoding. The input may also include feature values or vectors derived from the aforementioned inputs, such as cross-correlations.
[0060] A training record can be constructed from characterization of the presence and / or abundance of one or more mammalian nucleic acid molecules and / or one or more non-mammalian nucleic acid molecules of a biological sample of one or more subjects.
[0061] The model may process the input features to generate output values including one or more classifications, one or more predictions, or a combination thereof. For example, such classifications or predictions may include a binary classification of whether a subject has cancer or not (e.g., absence of disease or disorder), a classification between groups of categorical labels (e.g., "no disease or disorder," "apparent disease or disorder," and "likely disease or disorder"), the likelihood (e.g., relative likelihood or likelihood) of developing a particular disease or disorder, a score indicating the presence of a disease or disorder, a "risk factor" for the likelihood of the patient's death, and a confidence interval for any numerical prediction. Various machine learning techniques may be cascaded such that the output of the machine learning techniques may be used as input features to subsequent layers or subsections of the predictive model.
[0062] To train the model (e.g., by determining model weights and correlations) to generate real-time classifications or predictions, the model may be trained using the datasets and / or features described elsewhere herein. Such datasets may be large enough to generate statistically significant classifications or predictions. For example, a dataset may include a database of data, where the data may include one or more nucleic acid molecule sequencing reads for one or more subjects and corresponding disease labels for the one or more subjects. The training dataset may be collected from training subjects (e.g., human and / or non-human mammals). Each subject's training dataset may have a diagnostic status indicating that the subject has been diagnosed with a disease (e.g., cancer or a non-cancerous disease) or has not been diagnosed with a biological condition.
[0063] The dataset may be divided into subsets (e.g., separate or overlapping), such as a training dataset, a development dataset, and a test dataset. For example, the dataset may be divided into a training dataset that includes 80% of the dataset, a development dataset that includes 10% of the dataset, and a test dataset that includes 10% of the dataset. The training dataset may include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The development dataset may include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The test dataset may include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, leave-one-out cross-validation may be used. The training set (e.g., training data set) may be selected by random sampling of a set of data corresponding to one or more patient cohorts to ensure independence of sampling. Alternatively, the training set (e.g., training data set) may be selected by proportional sampling of a set of data corresponding to one or more patient cohorts to ensure independence of sampling.
[0064] To improve the accuracy of model predictions and reduce overfitting of predictive models, the dataset may be expanded to increase the number of samples in the training set. For example, data expansion may include rearranging the order of observations in the training records. To accommodate datasets with missing observations, methods of imputing missing data such as forward filling, backfilling, linear interpolation, and multitask Gaussian processes may be used. The dataset may be filtered or batch corrected to remove or reduce confounding factors. For example, within the database, a subset of patients may be excluded.
[0065] The predictive model may include one or more neural networks, such as a neural network, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), or a deep RNN. The recurrent neural network may include units that may be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, the model may include an algorithmic architecture that includes a neural network with a set of input features (e.g., one or more nucleic acid molecular sequencing reads), vitals (described elsewhere herein), patient medical history, and / or patient demographics. Neural network techniques such as shedding or normalization may be used during training of the predictive model to prevent overfitting. The neural network may include multiple sub-networks, each of which is configured to generate classifications or predictions of different types of output information (e.g., which may be combined to form an overall output of the neural network). The machine learning model may alternatively utilize statistical or related algorithms, including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, and ensemble and gradient boosted variations thereof.
[0066] If the predictive model generates a classification or prediction of a disease or disorder, a notification (e.g., an alert or alarm) may be generated and sent to a health care provider, such as a doctor, nurse, or other member of the patient's treatment team in a hospital. The notification may be sent via an automated phone call, a Short Message Service (SMS) or Multimedia Message Service (MMS) message, an email, or an alert in a dashboard. The notification may include output information such as a prediction of the disease or disorder, a predicted likelihood of the disease or disorder, an expected time to onset of the disease or disorder, a confidence interval of the likelihood or time, or a recommended course of treatment for the disease or disorder.
[0067] To verify the performance of the predictive model, different performance metrics can be generated. For example, the area under the receiver operating characteristic curve (AUROC) can be used to determine the diagnostic ability of the predictive model. For example, the predictive model can use an adjustable classification threshold, such that the specificity and sensitivity are adjustable, and the receiver operating characteristic curve (ROC) can be used to identify different operating points corresponding to different values of specificity and sensitivity.
[0068] In some cases, such as when the dataset is not large enough, cross-validation may be performed to assess the robustness of the model across different training and testing datasets.
[0069] The following definitions may be used to calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), area under the precision recall curve (AUPR), AUROC, or the like. A "false positive" may refer to an outcome in which a positive outcome or result is generated erroneously or prematurely (e.g., before or without the actual onset of a disease or disorder). A "true positive" may refer to an outcome in which a positive outcome or result is correctly generated if the patient has a disease or disorder (e.g., the patient exhibits symptoms of a disease or disorder or the patient's records indicate the disease or disorder). A "false negative" may refer to an outcome in which a negative outcome or result is generated but the patient has a disease or disorder (e.g., the patient exhibits symptoms of a disease or disorder or the patient's records indicate the disease or disorder). A "true negative" may refer to an outcome in which a negative outcome or result is generated (e.g., before or without the actual onset of a disease or disorder).
[0070] A predictive model may be trained until certain predefined conditions for accuracy or performance are met, such as having a minimum desired value corresponding to a diagnostic accuracy measure. For example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of occurrence of a disease or disorder in a subject. As another example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of aggravation or recurrence of a disease or disorder for which a subject has previously been treated. Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, AUPR, and AUROC, which correspond to the diagnostic accuracy of detecting or predicting a disease or disorder.
[0071] For example, such a predetermined condition can be that the sensitivity of predicting a disease or disorder includes values, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0072] As another example, such a predetermined condition may be that the specificity of predicting a disease or disorder includes values, for example, of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0073] As another example, such a predetermined condition can be that the positive predictive value (PPV) of predicting the disease or disorder includes values, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0074] As another example, such a predetermined condition can be that the negative predictive value (NPV) of predicting the disease or disorder includes values, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0075] As another example, such a predetermined condition may be that the area under the curve (AUC) of a receiver operating characteristic (ROC) curve for predicting the disease or disorder (AUROC) comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0076] As another example, such a predetermined condition may be that the area under the precision recall curve (AUPR) predicting the disease or disorder includes a value of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0077] In some embodiments, the trained model may be trained or configured to predict a disease or disorder with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0078] In some embodiments, the model is a neural network or a convolutional neural network, see Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.
[0079] In some embodiments, the data is de-dimensionalized using independent component analysis (ICA), such as that described in Lee, T.-W. (1998): Independent component analysis: Theory and applications, Boston, Mass: Kluwer Academic Publishers, ISBN 0-7923-8261-7, and Hyvaerinen, A.; Karhunen, J.; Oja, E. (2001): Independent Component Analysis, New York: Wiley, ISBN 978-0-471-40540-5, which are incorporated herein by reference in their entireties.
[0080] In some embodiments, the data is de-dimensionalized using principal component analysis (PCA), such as that described in Jolliffe, IT (2002). Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag. doi:10.1007 / b98835. ISBN 978-0-387-95442-4, which are incorporated herein by reference in their entireties.
[0081] SVM is an SVM, Cristianini and Shawe-Taylor. Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, Duda, Pattern Classification, Second Edition, 2001, John Wiley&Sons, Inc., pp. 259, 262-265 and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al. al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a given set of binary labeled data using a hyperplane that is maximally far from the labeled data. When linear separation is not possible, SVMs can work in conjunction with the technique of "kernels" that automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.
[0082] Decision trees are generally described by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Decision tree based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that may be used is Classification and Regression Trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and Random Forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York. pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests-Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.
[0083] Clustering (e.g., unsupervised and supervised clustering model algorithms) are described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a data set. To identify natural groupings, two problems are addressed. First, determine how to measure the similarity (or dissimilarity) between two samples. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than samples in the other cluster. Second, determine a mechanism for splitting the data into clusters using the similarity measure. Similarity measures are discussed in section 6.7 of Duda 1973, which states that one way to begin a clustering study is to define a distance function and calculate a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster will be significantly smaller than the distance between reference entities in different clusters. However, as described on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x,x') may be used to compare two vectors x and x'. Traditionally, s(x,x') is a symmetric function whose value is large if x and x' are "similar" in some way. An example of a non-metric similarity function s(x,x') is given on page 218 of Duda 1973. Once a method for measuring "similarity" or "dissimilarity" between points in a data set has been selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. A partition of the data set that extremizes a criterion function is used to cluster the data. See Duda 1973, page 217. Criterion functions are discussed in Duda 1973, section 6.8.More recently, Duda et al., Pattern Classification, 2nd edition, John Wiley & Sons, Inc. New York has been published. Pages 537-563 describe clustering in detail. Details of clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is incorporated herein by reference. Certain exemplary clustering techniques that may be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum of squares algorithms), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering. In some embodiments, the clustering includes unsupervised clustering, where no preconceived notions are imposed of what clusters should form when a training set is clustered.
[0084] Regression models, such as those of the multi-category logit model, are described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, Inc., New York, Chapter 8, which is incorporated herein by reference in its entirety. In some embodiments, the model utilizes the regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, which is incorporated herein by reference in its entirety. In some embodiments, gradient boosting models are used, for example, for the classification algorithms described herein, and these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019). "Gradient Boosting". Hands-On Machine Learning with R. Chapman & Hall. pp. 221-245. ISBN 978-1-138-49568-5., which is incorporated herein by reference in its entirety. In some embodiments, ensemble modeling techniques are used, and these ensemble modeling techniques are described in the implementation of the classification models herein and in Zhou Zhihua (2012) Ensemble Methods: Foundations and Algorithms. Chapman and Hall / CRC. ISBN 978-1-439-83003-1, which is incorporated herein by reference in its entirety.
[0085] In some embodiments, the machine learning analysis is performed by a device executing one or more programs (e.g., one or more programs stored in non-persistent or persistent memory) that include instructions for performing the data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., a processing core) and a memory (e.g., one or more programs stored in non-persistent or persistent memory) that includes instructions for performing the data analysis.
[0086] system The present disclosure provides a computer system programmed to implement the method of the present disclosure. Figure 7 shows a computer system 201, which is programmed or otherwise configured to predict a disease (e.g., cancer or non-cancerous disease), train a predictive model, generate a recommended therapy, generate and / or predict a longitudinal course of treatment of a disease for one or more subjects, or any combination thereof, as described elsewhere herein. The computer system 201 may be a user's electronic device, or a computer system remotely located with respect to the electronic device. The electronic device may be a mobile electronic device.
[0087] The computer system 201 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 205, which may be a single-core or multiple-core processor, or multiple processors for parallel processing. The computer system 201 also includes memory or memory locations 204 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 206 (e.g., hard disk), a communication interface 208 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 207, such as cache, other memory, data storage devices, and / or electronic display adapters. The memory 204, the storage unit 206, the interface 208, and the peripheral devices 207 are in communication with the CPU 205 via a communication bus (solid lines), such as a motherboard. The storage unit 206 may be a data storage unit (or data repository) for storing data. The computer system 201 may be operatively coupled to a computer network ("network") 203 using the communication interface 208. Network 203 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, network 203 is a telecommunications and / or data network. Network 203 may include one or more computer servers that may enable distributed computing, such as cloud computing. Network 203 may implement a peer-to-peer network, in some cases with computer system 201, which may enable devices coupled to computer system 201 to operate as clients or servers.
[0088] CPU 205 may execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 204. The instructions may be directed to CPU 205, which may then be programmed or configured to implement the methods of the present disclosure, as described elsewhere herein. Examples of operations performed by CPU 205 may include fetch, decode, execute, and writeback.
[0089] The CPU 205 may be part of a circuit, such as an integrated circuit. One or more other components of the system 201 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0090] The storage unit 206 may store files such as drivers, libraries, and saved programs. The storage unit 206 may store user data (e.g., disease prediction and / or one or more mammalian and / or one or more non-mammalian features of a user and / or subject's nucleic acid sequencing reads, user preferences, user programs, or any combination thereof). The computer system 201 may include one or more additional data storage units that are external to the computer system 201, such as located on a remote server in communication with the computer system 201 via an intranet or the Internet, in some cases.
[0091] Computer system 201 may communicate with one or more remote computer systems via network 203. For example, computer system 201 may communicate with a remote computer system of a user. Examples of remote computer systems may include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user may access computer system 201 via network 203.
[0092] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored on electronic storage locations of computer system 201, such as, for example, on memory 204 or electronic storage unit 206. Machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by processor 205. In some cases, the code may be retrieved from storage unit 206 and stored on memory 204 for immediate access by processor 205. In some circumstances, electronic storage unit 206 may be excluded and machine executable instructions are stored in memory 204.
[0093] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that may be selected to enable the code to be executed in a pre-compiled or as-compiled manner.
[0094] Aspects of the systems and methods provided herein, such as the computer system 201, may be embodied in programming. Various aspects of the technology may be thought of as a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried or embodied in some type of machine-readable medium. The machine-executable code may be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or hard disk. A "storage" type medium may include any or all of the tangible memory of a computer, processor, etc., or their associated modules, such as various semiconductor memories, tape drives, disk drives, etc., and may provide non-transitory storage at any time for software programming. All or a portion of the software may be communicated, at times, over the Internet or various other communications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from a management server or host computer to a computer platform of an application server. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical terrestrial communications networks, and across various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0095] Thus, a machine-readable medium such as a computer executable code may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices of any computer(s) such as may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, a DVD or a DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium having a pattern of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, a cable or link transporting such a carrier wave, or any other medium from which a computer may read programming code or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0096] The computer system 201 may include or be in communication with an electronic display 202, which includes a user interface (UI) 209 to provide, for example, a display for visualization of prediction results or an interface for training a predictive model, as described elsewhere herein. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0097] The methods and systems of the present disclosure may be implemented by one or more algorithms and / or predictive models, as described elsewhere herein. The algorithms and / or predictive models may be implemented by software upon execution by the central processing unit 205. The algorithms and / or predictive models may, for example, predict cancer in a subject, determine tailored treatments and / or therapeutics for treating a disease in a subject or one or more subjects (e.g., cancers as described elsewhere herein), and predict the longitudinal course of a therapeutic for treating a disease in a subject or one or more subjects (e.g., cancers as described elsewhere herein).
[0098] Embodiment Numbered embodiment 1 includes a method of determining a disease in a subject, comprising providing a biological sample from the subject, enriching one or more nucleic acid molecules of the biological sample by affinity targeting of epigenetic features common to the one or more nucleic acid molecules, sequencing the enriched one or more nucleic acid molecules to generate one or more nucleic acid molecule sequencing reads, and determining the disease in the subject as an output of a predictive model when the enriched one or more nucleic acid molecules are provided as input to the predictive model. Numbered embodiment 2 includes the method of numbered embodiment 1, wherein the one or more nucleic acid molecules comprise one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. Numbered embodiment 3 includes the method of numbered embodiment 1 or 2, wherein the disease comprises cancer or a non-cancerous disease. Numbered embodiment 4 includes the method of any one of numbered embodiments 1-3, wherein the cancer comprises acute myeloid leukemia, adrenal cortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. Numbered embodiment 5 includes the method of any one of numbered embodiments 1-4, wherein the non-cancerous disease comprises a non-cancerous disease of lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof. Numbered embodiment 6 includes the method of any one of numbered embodiments 1-5, further comprising filtering the one or more nucleic acid molecule sequencing reads to identify one or more non-mammalian sequencing reads and one or more mammalian sequencing reads. Numbered embodiment 7 includes the method of any one of numbered embodiments 1-6, wherein the epigenetic signature comprises a nucleic acid epigenetic signature.Numbered embodiment 8 includes the method of any one of numbered embodiments 1-7, wherein the epigenetic feature comprises a mammalian nucleic acid epigenetic feature or a non-mammalian nucleic acid epigenetic feature. Numbered embodiment 9 includes the method of any one of numbered embodiments 1-8, wherein the non-mammalian nucleic acid epigenetic feature comprises phosphorothioate linked nucleotides. Numbered embodiment 10 includes the method of any one of numbered embodiments 1-9, wherein the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. Numbered embodiment 11 includes the method of any one of numbered embodiments 1-9, wherein the subject comprises a human or a non-human mammal. Numbered embodiment 12 includes the method of any one of numbered embodiments 1-11, wherein the one or more mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 13 includes the method of any one of numbered embodiments 1-12, wherein the one or more non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 14 includes the method of any one of numbered embodiments 1-13, wherein affinity targeting of the nucleic acid epigenetic feature comprises enriching the nucleic acid epigenetic feature. Numbered embodiment 15 includes the method of any one of numbered embodiments 1-14, wherein the nucleic acid epigenetic feature comprises a methylated CpG dinucleotide pair, an unmethylated CpG dinucleotide pair, or a combination thereof. Numbered embodiment 16 includes the method of any one of numbered embodiments 1-15, wherein the nucleic acid epigenetic feature comprises the nucleobase 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof. Numbered embodiment 17 includes the method of any one of numbered embodiments 1 to 16, wherein the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature.Numbered embodiment 18 includes the method of any one of numbered embodiments 1-17, wherein the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibody, aptamer, recombinant epigenetic protein, or any combination thereof. Numbered embodiment 19 includes the method of any one of numbered embodiments 1-18, wherein the recombinant epigenetic protein comprises an epigenetic reader, writer, eraser, or any combination thereof. Numbered embodiment 20 includes the method of any one of numbered embodiments 1-19, wherein the epigenetic reader comprises recombinant methyl-CpG binding protein Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or a recombinant methyl-binding domain derived therefrom. Numbered embodiment 21 includes the method of any one of numbered embodiments 1 to 20, wherein the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. Numbered embodiment 22 includes the method of any one of numbered embodiments 1 to 21, wherein the epigenetic reader comprises a microbial protein Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or a recombinant methyl-binding domain derived therefrom. Numbered embodiment 23 includes the method of any one of numbered embodiments 1 to 22, wherein the epigenetic writer and eraser are catalytically inactive. Numbered embodiment 24 includes the method of any one of numbered embodiments 1 to 23, wherein the epigenetic reader, writer, and eraser comprise an epitope tag.Numbered embodiment 25 includes the method of any one of numbered embodiments 1 to 24, wherein the epitope tag comprises an N-terminal or C-terminal 6x histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. Numbered embodiment 26 includes the method of any one of numbered embodiments 1 to 25, wherein the molecular recognition motif comprises a birA or a sortase motif. Numbered embodiment 27 includes the method of any one of numbered embodiments 1 to 26, further comprising concentrating the one or more mammalian and non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody against the epitope tag. Numbered embodiment 28 includes the method of any one of numbered embodiments 1 to 27, wherein the specific affinity reagent comprises a region for recognizing and binding to an epigenetic feature. Numbered embodiment 29 includes the method of any one of numbered embodiments 1-28, wherein the affinity targeting comprises incubating the biological sample with a solid support comprising a plurality of immobilized affinity agents. Numbered embodiment 30 includes the method of any one of numbered embodiments 1-29, wherein the plurality of immobilized affinity agents comprises a region that will bind to the epigenetic feature. Numbered embodiment 31 includes the method of any one of numbered embodiments 1-30, wherein the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymer, or any combination thereof. Numbered embodiment 32 includes the method of any one of numbered embodiments 1-31, wherein the filtering comprises filtering the one or more mammalian sequencing reads and the one or more non-mammalian sequencing reads against a genome database. Numbered embodiment 33 includes the method of any one of numbered embodiments 1-32, wherein the genome database is a human genome database. Numbered embodiment 34 includes a method according to any one of numbered embodiments 1 to 33, wherein the predictive model is trained with one or more mammalian, one or more non-mammalian, or combinations thereof features determined from one or more nucleic acid molecules of a biological sample of one or more subjects and a corresponding disease of one or more subjects.Numbered embodiment 35 includes the method of any one of numbered embodiments 1-34, wherein the one or more mammalian features include mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. Numbered embodiment 36 includes the method of any one of numbered embodiments 1-35, wherein the one or more mammalian features include mammalian functional gene and biochemical pathway abundances. Numbered embodiment 37 includes the method of any one of numbered embodiments 1-36, wherein the one or more non-mammalian features include microbial taxonomic assignments and a number of sequencing reads associated therewith. Numbered embodiment 38 includes the method of any one of numbered embodiments 1-37, wherein the one or more non-mammalian features include microbial functional gene and biochemical pathway abundances. Numbered embodiment 39 includes the method of any one of numbered embodiments 1-38, wherein the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. Numbered embodiment 40 includes the method of any one of numbered embodiments 1-39, wherein the accuracy of the predictive model for determining a disease is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% when the biological sample is enriched for one or more nucleic acid molecules compared to when the biological sample is not enriched for one or more nucleic acid molecules. Numbered embodiment 41 includes the method of any one of numbered embodiments 1-40, wherein the predictive model for determining a disease of the subject includes an area under the curve of at least about 0.70, at least about 0.80, at least about 0.85, at least about 0.90, or at least about 0.95. Numbered embodiment 42 includes the method of any one of numbered embodiments 1-41, wherein the output of the trained predictive model includes an analysis of a combination of one or more mammalian feature abundances and one or more non-mammalian feature abundances.Numbered embodiment 43 includes the method of any one of numbered embodiments 1-42, wherein the input of the trained predictive model includes epigenome abundance information from one or more of the following biological kingdoms: mammals, bacteria, archaea, fungi, and / or viruses. Numbered embodiment 44 includes the method of any one of numbered embodiments 1-43, wherein the predictive model is further trained on tissue-specific locations of disease. Numbered embodiment 45 includes the method of any one of numbered embodiments 1-44, wherein the predictive model is further trained on cancer type, subtype, stage, prognosis, or any combination thereof. Numbered embodiment 46 includes the method of any one of numbered embodiments 1-45, wherein the predictive model outputs a cancer type, subtype, stage, prognosis, or any combination thereof when provided with nucleic acid sequencing reads of the subject's biological sample. Numbered embodiment 47 includes the method of any one of numbered embodiments 1-46, wherein the predictive model outputs a cancer therapy response of the subject. Numbered embodiment 48 includes the method of any one of numbered embodiments 1-47, where the trained predictive model outputs a therapy for the subject that results in at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% reduction in cancer area in the subject. Numbered embodiment 49 includes the method of any one of numbered embodiments 1-48, where the trained predictive model outputs a longitudinal model of the subject's cancer in response to a therapy, an adjustment to a therapy to treat the subject's cancer, or a combination thereof. Numbered embodiment 50 includes the method of any one of numbered embodiments 1-49, where the predictive model removes contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. Numbered embodiment 51 includes a method of any one of numbered embodiments 1 to 50, wherein the concentrating reduces the total of the one or more nucleic acid molecules by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99%.
[0099] Numbered embodiment 51 includes a method of training a predictive model, comprising providing a biological sample of one or more subjects having a disease, enriching the biological sample of the one or more subjects by affinity targeting of epigenetic features common to one or more nucleic acid molecules of the biological sample, sequencing the enriched one or more nucleic acid molecules to generate one or more nucleic acid molecule sequencing reads, and training a predictive model with one or more features of the one or more nucleic acid molecule sequencing reads and the disease of the one or more subjects. Numbered embodiment 52 includes the method of numbered embodiment 51, wherein the epigenetic features include mammalian epigenetic features or non-mammalian epigenetic features. Embodiment 53 includes the method of numbered embodiment 51 or 52, wherein the one or more features include one or more disease features. Numbered embodiment 54 includes the method of any one of numbered embodiments 51-53, wherein the trained predictive model determines the disease of one or more other subjects different from the one or more subjects when the trained predictive model is provided with nucleic acid sequencing reads of a biological sample of another one or more subjects. Numbered embodiment 55 includes the method of any one of numbered embodiments 51-54, wherein the one or more nucleic acid molecules comprise one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. Numbered embodiment 56 includes the method of any one of numbered embodiments 51-55, further comprising filtering the one or more nucleic acid sequencing reads to identify one or more non-mammalian sequencing reads, one or more mammalian sequencing reads, or a combination thereof. Numbered embodiment 57 includes the method of any one of numbered embodiments 51-56, wherein the epigenetic signature comprises a nucleic acid epigenetic signature. Numbered embodiment 58 includes the method of any one of numbered embodiments 51-57, wherein the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. Numbered embodiment 58 includes the method of any one of numbered embodiments 51-57, wherein the one or more subjects are human or non-human mammals.Numbered embodiment 59 includes the method of any one of numbered embodiments 51 to 58, wherein the one or more mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 60 includes the method of any one of numbered embodiments 51 to 59, wherein the one or more non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 61 includes the method of any one of numbered embodiments 51 to 60, wherein affinity targeting of the nucleic acid epigenetic feature comprises enriching the nucleic acid epigenetic feature. Numbered embodiment 62 includes the method of any one of numbered embodiments 51 to 61, wherein the nucleic acid epigenetic feature comprises a methylated CpG dinucleotide pair, an unmethylated CpG dinucleotide pair, or a combination thereof. Numbered embodiment 63 includes the method of any one of numbered embodiments 51 to 62, wherein the nucleic acid epigenetic feature comprises the nucleobase 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof. Numbered embodiment 64 includes the method of any one of numbered embodiments 51 to 63, wherein the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature. Numbered embodiment 65 includes the method of any one of numbered embodiments 51 to 64, wherein the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibody, aptamer, recombinant epigenetic protein, or any combination thereof. Numbered embodiment 66 includes the method of any one of numbered embodiments 51 to 65, wherein the recombinant epigenetic protein comprises an epigenetic reader, writer, eraser, or any combination thereof.Numbered embodiment 67 includes the method of any one of numbered embodiments 51 to 66, wherein the epigenetic reader comprises a recombinant methyl-CpG binding protein Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or a recombinant methyl-binding domain derived therefrom.Numbered embodiment 68 includes the method of any one of numbered embodiments 51 to 67, wherein the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. Numbered embodiment 69 includes the method of any one of numbered embodiments 51 to 68, wherein the epigenetic writer and eraser are catalytically inactive. Numbered embodiment 70 includes the method of any one of numbered embodiments 51 to 69, wherein the epigenetic reader, writer, and eraser comprise an epitope tag. Numbered embodiment 71 includes the method of any one of numbered embodiments 51 to 70, wherein the epitope tag comprises an N-terminal or C-terminal 6× Histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. Numbered embodiment 72 includes the method of any one of numbered embodiments 51 to 71, wherein the molecular recognition motif comprises a birA or a sortase motif. Numbered embodiment 73 includes the method of any one of numbered embodiments 51 to 72, further comprising concentrating the one or more mammalian nucleic acid molecules and the one or more non-mammalian nucleic acid molecules with a solid support, wherein the solid support comprises an immobilized complementary antibody to the epitope tag. Numbered embodiment 74 includes the method of any one of numbered embodiments 51 to 73, wherein the specific affinity reagent comprises a region for recognizing and binding to an epigenetic feature.Numbered embodiment 75 includes the method of any one of numbered embodiments 51-74, wherein the affinity targeting comprises incubating the biological sample with a solid support comprising a plurality of immobilized affinity agents. Numbered embodiment 76 includes the method of any one of numbered embodiments 51-75, wherein the plurality of immobilized affinity agents comprises a region that will bind to the epigenetic feature. Numbered embodiment 77 includes the method of any one of numbered embodiments 51-76, wherein the solid support comprises magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, pH-sensitive polymer, or any combination thereof. Numbered embodiment 78 includes the method of any one of numbered embodiments 51-77, wherein the filtering comprises filtering one or more of the mammalian and non-mammalian sequencing reads against a genome database. Numbered embodiment 79 includes the method of any one of numbered embodiments 51-78, wherein the genome database is a human genome database. Numbered embodiment 80 includes the method of any one of numbered embodiments 51-79, wherein the one or more features include one or more mammalian features, one or more non-mammalian features, or a combination thereof. Numbered embodiment 81 includes the method of any one of numbered embodiments 51-80, wherein the one or more mammalian features include mammalian genomic coordinates or annotated genomic loci and a number of sequencing reads associated therewith. Numbered embodiment 82 includes the method of any one of numbered embodiments 51-81, wherein the one or more mammalian features include mammalian functional genes and biochemical pathway abundances. Numbered embodiment 83 includes the method of any one of numbered embodiments 51-82, wherein the one or more non-mammalian features include microbial taxonomic assignments and a number of sequencing reads associated therewith. Numbered embodiment 84 includes the method of any one of numbered embodiments 51-83, wherein the one or more non-mammalian features include microbial functional genes and biochemical pathway abundances. Numbered embodiment 85 includes a method of any one of numbered embodiments 51 to 84, wherein the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.Numbered embodiment 86 includes the method of any one of numbered embodiments 51-85, wherein the disease comprises cancer or a non-cancer disease. Numbered embodiment 87 includes the method of any one of numbered embodiments 51-86, wherein the accuracy of the predictive model for determining the disease is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% when the biological sample is enriched for one or more nucleic acid molecules, compared to when the biological sample is not enriched for one or more nucleic acid molecules. Numbered embodiment 88 includes the method of any one of numbered embodiments 51-87, wherein the predictive model for determining the disease in another subject comprises an area under the curve of at least about 0.70, at least about 0.80, at least about 0.85, at least about 0.90, or at least about 0.95. Numbered embodiment 89 includes the method of any one of numbered embodiments 51-88, wherein the non-mammalian nucleic acid epigenetic feature comprises phosphorothioate linked nucleotides. Numbered embodiment 90 includes the method of any one of numbered embodiments 51-89, wherein the epigenetic reader comprises microbial proteins Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or recombinant methyl-binding domains derived therefrom. Numbered embodiment 91 includes the method of any one of numbered embodiments 51-90, wherein the output of the trained predictive model comprises an analysis of a combination of one or more mammalian feature abundances and one or more non-mammalian feature abundances. Numbered embodiment 92 includes the method of any one of numbered embodiments 51-91, wherein the input of the trained predictive model comprises epigenome abundance information from one or more of the following kingdoms of life: mammals, bacteria, archaea, fungi, and / or viruses. Numbered embodiment 93 includes a method according to any one of numbered embodiments 51 to 92, wherein the predictive model is further trained on tissue-specific locations of the disease.Numbered embodiment 94 includes the method of any one of numbered embodiments 51-93, where the predictive model is further trained on a cancer type, subtype, stage, prognosis, or any combination thereof. Numbered embodiment 95 includes the method of any one of numbered embodiments 51-94, where the predictive model outputs a cancer type, subtype, stage, prognosis, or any combination thereof when provided with nucleic acid sequencing reads of a biological sample of one or more additional subjects. Numbered embodiment 96 includes the method of any one of numbered embodiments 51-95, where the trained predictive model outputs a cancer therapy response of one or more additional subjects. Numbered embodiment 97 is a numbered embodiment in which the trained predictive model outputs a therapy for one or more other subjects that results in at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% reduction in the cancerous area of the one or more other subjects. Numbered embodiment 98 includes the method of any one of numbered embodiments 51 to 97, wherein the trained predictive model outputs a longitudinal model of the cancer in one or more additional subjects in response to a therapy, an adjustment to a therapy to treat the cancer in the subject, or a combination thereof. Numbered embodiment 99 includes the method of any one of numbered embodiments 51-98, wherein the cancer includes acute myeloid leukemia, adrenal cortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. Numbered embodiment 100 includes the method of any one of numbered embodiments 51-99, wherein the predictive model is configured to remove contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. Numbered embodiment 101 includes the method of any one of numbered embodiments 51-100, wherein the non-cancerous disease comprises a non-cancerous disease of lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof. Numbered embodiment 102 includes the method of any one of numbered embodiments 51-101, wherein enriching reduces the total of the one or more nucleic acid molecules by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99%.
[0100] Numbered embodiment 103 includes a computer system for determining a disease in a subject, the computer system comprising one or more processors and a non-transitory computer readable storage medium comprising software, the software comprising executable instructions that, as a result of execution, cause the one or more processors of the computer system to: (i) receive one or more nucleic acid molecule sequencing reads of one or more nucleic acid molecules of a biological sample of the subject, the one or more nucleic acid molecules being enriched by affinity targeting of epigenetic features common to the one or more nucleic acid molecules; and (ii) determine the disease in the subject as an output of a predictive model when the one or more nucleic acid molecule sequencing reads are provided to a predictive model. Numbered embodiment 104 includes the system of numbered embodiment 103, wherein the one or more nucleic acid molecules comprise one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. Numbered embodiment 105 includes the system of numbered embodiment 103 or 104, wherein the disease comprises a cancer or a non-cancerous disease. Numbered embodiment 106 includes the system of any one of numbered embodiments 103 to 105, wherein the cancer includes acute myeloid leukemia, adrenal cortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid cancer, uterine carcinoma sarcoma, uterine endometrial cancer, uveal melanoma, or any combination thereof. Numbered embodiment 107 includes a system described in any one of numbered embodiments 103 to 106, wherein the non-cancerous disease includes a non-cancerous disease of lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof.Numbered embodiment 108 includes the system of any one of numbered embodiments 103-107, wherein the executable instructions further comprise filtering the one or more nucleic acid molecule sequencing reads to identify one or more non-mammalian sequencing reads and one or more mammalian sequencing reads. Numbered embodiment 109 includes the system of any one of numbered embodiments 103-108, wherein the epigenetic feature comprises a nucleic acid epigenetic feature. Numbered embodiment 110 includes the system of any one of numbered embodiments 103-109, wherein the epigenetic feature comprises a mammalian nucleic acid epigenetic feature or a non-mammalian nucleic acid epigenetic feature. Numbered embodiment 111 includes the system of any one of numbered embodiments 103-110, wherein the non-mammalian nucleic acid epigenetic feature comprises phosphorothioate linked nucleotides. Numbered embodiment 112 includes the system of any one of numbered embodiments 103-111, wherein the biological sample comprises a tissue, a liquid biopsy sample, or a combination thereof. Numbered embodiment 113 includes the system of any one of numbered embodiments 103-112, wherein the subject is a human or non-human mammal. Numbered embodiment 114 includes the system of any one of numbered embodiments 103-113, wherein the one or more mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 115 includes the system of any one of numbered embodiments 103-114, wherein the one or more non-mammalian nucleic acid molecules comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. Numbered embodiment 116 includes the system of any one of numbered embodiments 103-115, wherein affinity targeting of nucleic acid epigenetic features comprises enriching for nucleic acid epigenetic features. Numbered embodiment 117 includes a system described in any one of numbered embodiments 103 to 116, wherein the nucleic acid epigenetic feature includes a methylated CpG dinucleotide pair, an unmethylated CpG dinucleotide pair, or a combination thereof.Numbered embodiment 118 includes the system of any one of numbered embodiments 103 to 117, wherein the nucleic acid epigenetic feature comprises the nucleobase 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof. Numbered embodiment 119 includes the system of any one of numbered embodiments 103 to 118, wherein the affinity targeting utilizes a specific affinity reagent to bind to the epigenetic feature. Numbered embodiment 120 includes the system of any one of numbered embodiments 103 to 119, wherein the specific affinity reagent comprises streptavidin, NeutrAvidin, polyclonal, monoclonal, recombinant antibody, aptamer, recombinant epigenetic protein, or any combination thereof. Numbered embodiment 121 includes the system of any one of numbered embodiments 103 to 120, wherein the recombinant epigenetic protein comprises an epigenetic reader, writer, eraser, or any combination thereof. Numbered embodiment 122 includes the system of any one of numbered embodiments 103 to 121, wherein the epigenetic reader comprises a recombinant methyl-CpG binding protein Mecp2, Mbd1-6, SETDB1, SETDB2, TIP5 / BAZ2A, Zbtb38, Kaiso, Zbtb4, Np95, Np97, or a recombinant methyl-binding domain derived therefrom.Numbered embodiment 123 includes the system of any one of numbered embodiments 103 to 122, wherein the epigenetic reader comprises a recombinant zinc finger CXXC domain-containing protein KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, or a recombinant CXXC domain derived therefrom. Numbered embodiment 124 includes a system described in any one of numbered embodiments 103 to 123, wherein the epigenetic reader comprises a microbial protein Dam, CcrM, ModA13, SpnD39III, Dcm, JHP1050, M2.Hpy.AII, or a recombinant methyl-binding domain derived therefrom.Numbered embodiment 125 includes the system of any one of numbered embodiments 103-124, wherein the epigenetic writer and eraser are catalytically inactive. Numbered embodiment 126 includes the system of any one of numbered embodiments 103-125, wherein the epigenetic reader, writer, and eraser comprise an epitope tag. Numbered embodiment 127 includes the system of any one of numbered embodiments 103-126, wherein the epitope tag comprises an N-terminal or C-terminal 6× histidine tag, green fluorescent protein (MA), myc, hemagglutinin (HA), Fc fusion, a molecular recognition motif, or any combination thereof. Numbered embodiment 128 includes the system of any one of numbered embodiments 103-127, wherein the molecular recognition motif comprises a birA or a sortase motif. Numbered embodiment 129 includes the system of any one of numbered embodiments 103-128, further comprising concentrating one or more mammalian and non-mammalian nucleic acid molecules with a solid support, the solid support comprising an immobilized complementary antibody against the epitope tag. Numbered embodiment 130 includes the system of any one of numbered embodiments 103-129, wherein the specific affinity reagent comprises a region for recognizing and binding to the epigenetic feature. Numbered embodiment 131 includes the system of any one of numbered embodiments 103-130, wherein the affinity targeting comprises incubating the biological sample with a solid support comprising a plurality of immobilized affinity agents. Numbered embodiment 132 includes the system of any one of numbered embodiments 103-131, wherein the plurality of immobilized affinity agents comprises a region that will bind to the epigenetic feature. Numbered embodiment 133 includes a system described in any one of numbered embodiments 103 to 132, wherein the solid support includes magnetic beads, agarose beads, non-magnetic latex, functionalized sepharose, a pH-sensitive polymer, or any combination thereof.Numbered embodiment 134 includes the system of any one of numbered embodiments 103-133, where the filtering includes filtering the one or more mammalian sequencing reads and the one or more non-mammalian sequencing reads against a genome database. Numbered embodiment 135 includes the system of any one of numbered embodiments 103-134, where the genome database is a human genome database. Numbered embodiment 136 includes the system of any one of numbered embodiments 103-135, where the predictive model is trained with one or more mammalian, one or more non-mammalian, or combinations thereof features determined from one or more nucleic acid molecules of a biological sample of one or more subjects and a corresponding disease of one or more subjects. Numbered embodiment 137 includes the system of any one of numbered embodiments 103-136, where the one or more mammalian features include mammalian genome coordinates or annotated genomic loci and a number of sequencing reads associated therewith. Numbered embodiment 138 includes the system of any one of numbered embodiments 103-137, wherein the one or more mammalian features include mammalian functional gene and biochemical pathway abundance. Numbered embodiment 139 includes the system of any one of numbered embodiments 103-138, wherein the one or more non-mammalian features include a microbial taxonomic assignment and a number of sequencing reads associated therewith. Numbered embodiment 140 includes the system of any one of numbered embodiments 103-139, wherein the one or more non-mammalian features include a microbial functional gene and biochemical pathway abundance. Numbered embodiment 141 includes the system of any one of numbered embodiments 103-140, wherein the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof.Numbered embodiment 142 includes the system of any one of numbered embodiments 103-141, wherein the accuracy of the predictive model for determining a disease is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% when the biological sample is enriched for one or more nucleic acid molecules compared to when the biological sample is not enriched for one or more nucleic acid molecules. Numbered embodiment 143 includes the system of any one of numbered embodiments 103-142, wherein the predictive model for determining a disease of the subject includes an area under the curve of at least about 0.70, at least about 0.80, at least about 0.85, at least about 0.90, or at least about 0.95. Numbered embodiment 144 includes the system of any one of numbered embodiments 103-143, wherein the output of the trained predictive model includes an analysis of a combination of one or more mammalian feature abundances and one or more non-mammalian feature abundances. Numbered embodiment 145 includes the system of any one of numbered embodiments 103-143, wherein the input of the trained predictive model includes one or more of the following biological kingdoms: mammals, bacteria, archaea, fungi, and / or viruses. Numbered embodiment 146 includes the system of any one of numbered embodiments 103-145, where the predictive model is further trained on tissue-specific locations of disease. Numbered embodiment 147 includes the system of any one of numbered embodiments 103-146, where the predictive model is further trained on cancer type, subtype, stage, prognosis, or any combination thereof. Numbered embodiment 148 includes the system of any one of numbered embodiments 103-147, where the predictive model outputs a cancer type, subtype, stage, prognosis, or any combination thereof when provided with nucleic acid sequencing reads of the subject's biological sample. Numbered embodiment 149 includes the system of any one of numbered embodiments 103-148, where the predictive model outputs a cancer therapy response of the subject. Numbered embodiment 150 includes the system of any one of numbered embodiments 103-149, where the trained predictive model outputs a therapy for the subject that results in at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% reduction in the cancer area of the subject. Numbered embodiment 151 includes the system of any one of numbered embodiments 103-150, where the trained predictive model outputs a longitudinal model of the cancer of the subject in response to a therapy, an adjustment to a therapy to treat the cancer of the subject, or a combination thereof. Numbered embodiment 152 includes the system of any one of numbered embodiments 103-151, where the predictive model removes contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features.Numbered embodiment 153 includes a system described in any one of numbered embodiments 103 to 152, wherein the concentrated nucleic acid comprises a reduction of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99% of one or more nucleic acid molecules prior to concentration. EXAMPLES
[0101] Example 1: Discovery of 5-hydroxymethylcytosine microbial epigenetic biomarkers and evaluation in cancer diagnostic models Figures 5A-5D show the experimental parameters and resulting classification accuracy of a study of 5-hydroxymethylcytosine (5hmC) microbial epigenetic biomarker discovery and cancer diagnostic model evaluation. Figure 5A shows the cell-free DNA study from which 5-hydroxymethylcytosine enriched sequencing data was obtained, and the sample types present in the sequencing data. The resulting non-human sequencing data was then aligned to a reference database of microbial genomes ("rep206"). Figure 5B shows the dataset of alignment of non-human reads. Figure 5C shows the clinical details of the pancreatic cancer samples present in the aligned dataset. A machine learning model was then trained on 5hmC enriched microbial nucleic acids from pancreatic cancer patients and healthy individuals shown in Figure 5D ("5hmC sample" ROC curve, left; "input sample" ROC curve was generated from non-enriched nucleic acids). Due to the small number of samples, leave-one-out (LOO) cross-validation was performed instead of the more traditional 70 / 30 train-test split of the sample feature set. Figure 5E shows the clinical details of the lung cancer samples present in the dataset. Figure 5F shows the performance of machine learning models trained on 5hmC-enriched microbial nucleic acids from lung cancer patients and healthy individuals. Similar to Figure 5D, LOO was utilized to develop a lung cancer classifier.
[0102] FIG. 6A shows the cell-free DNA study from which 5-hydroxymethylcytosine enriched sequencing data was obtained, and the sample types present therein. FIG. 6B shows the performance of a random forest machine learning classifier trained on 5hmC-enriched microbial nucleic acids from various cancer types and healthy individuals. The ROC curves for each cancer type versus healthy are given with the cancer type specified on the respective ROC curve. FIG. 6C shows the performance of a random forest machine learning classifier trained on 5hmC-enriched microbial nucleic acids from colon and gastric cancer, and benign tumors from colon and gastric cancer. FIG. 6D shows the performance of a random forest machine learning classifier trained on the same samples from FIG. 6C. However, in this example, the microbial 5hmC feature set was limited to certain microbial kingdoms (i.e., bacteria, fungi, and viruses), thereby demonstrating that all three kingdoms of life contain features with 5hmC that have discriminatory power for cancer versus benign.
[0103] Example 2: Identification of 5hmC-positive microbial genomic regions by hMeDIP-seq Enrichment of 5hmC is performed using Active Motif's hMeDIP kit (#55010) according to the manufacturer's protocol. Briefly, 3-5 μg of human brain DNA (Zyagen #HG0201), Pseudomonas aeruginosa strain PAO1-LAC DNA (ATCC #47085D-5), Escherichia coli strain EDL933 DNA (ATCC #700927D-5), and Bacillus subtilis strain 168 DNA (ATCC #23857D-5) are fragmented using enzymatic digestion (Roche's KAPA frag kit for enzymatic fragmentation, #07962517001) according to the manufacturer's protocol. Samples are incubated at 37°C for 8 min and then purified using AMPure XP beads (Beckman Coulter #A63881). Fragmented DNA is quantified using Qubit1x dsDNA HS Assay Kit (ThermoFisher #Q33231) and fragmentation profiles are visualized using TapeStation Genomic (Agilent #5067-5365) and D1000 (Agilent #5067-5582) tapes. 100ng of fragmented human brain gDNA and 500ng of DNA from Pseudomonas aeruginosa, Escherichia coli, and Bacillus subtilis are incubated with 4μg of either rabbit anti-5hmC antibody or control IgG overnight at 4°C with rotation. 10% of the material (10ng and 50ng, respectively) is retained as input and stored at -80°C until downstream purification and analysis. Protein-antibody complexes are captured by adding 25 μL of Pierce Protein A / G Plus Agarose Beads (ThermoFisher #20423) and rotating the samples at room temperature for 2 hours, followed by washing as indicated in the manufacturer's protocol. The captured antibody-protein complexes are eluted from the beads using SDS-mediated elution. Similarly, an equal volume of elution buffer is also added to the inputs. The eluted immunoprecipitation (IP) material and their respective inputs are purified using Qiagen MinElute columns.They are then subjected to a qPCR-based QC analysis to assess the IP efficiency prior to library preparation.
[0104] Libraries are prepared using the 2S™ Plus DNA Library Kit (IDT #10009878) and 2S™ MID Adapter Set A+B (IDT #10009902) according to the manufacturer's protocol. Briefly, 9 and 14 PCR cycles were used to amplify the input and IP, respectively. The final libraries are eluted in a volume of 25 μL. The final libraries are quantified using the KAPA Library Quantification Kit (Roche #07960140001) and Qubit 1× dsDNA HS Assay Kit and visualized using TapeStation D1000 tapes. They are paired-end (150×150 8×0) sequenced on a NextSeq2000 using P3 chemistry (Illumina #20040561). Genome-wide 5hmC enrichment is computationally identified via the MeDIPS package (Lienhard, M., Grimm, C., Morkel, M., Herwig, R., & Chavez, L. (2014). MEDIPS: genome-wide differential coverage analysis of sequencing data derived from DNA enrichment experiments. Bioinformatics (Oxford, England), 30(2), 284-286. https: / / doi.org / 10.1093 / bioinformatics / btt650), where a statistically significant increase in sequencing reads at genomic loci of interest over the number of reads found in the non-immunoprecipitated input control is calculated and tabulated.
[0105] definition Unless otherwise defined, all technical terms, notations, and other technical and scientific or terminology used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms having commonly understood meanings are defined herein for clarity and / or ease of reference, and the inclusion of such definitions herein should not necessarily be construed as representing a substantial difference from what is commonly understood in the art.
[0106] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Thus, the description of a range should be considered to specifically disclose all possible subranges as well as individual values within that range. For example, the description of a range such as 1-6 should be considered to specifically disclose subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc., as well as individual numbers within that range, such as 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0107] As used in this specification and claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a sample" includes multiple samples, including mixtures thereof.
[0108] The terms "determining," "measuring," "evaluating," "assessing," "assaying," and "analyzing" are often used interchangeably herein to refer to forms of measurement. These terms include determining whether an element is present or not (e.g., detecting). These terms can include quantitative, qualitative, or quantitative and qualitative determinations. Evaluating can be relative or absolute. "Detecting the presence of" can include determining the amount of something present in addition to determining whether it is present or absent, depending on the context.
[0109] The terms "subject," "individual," or "patient" are used interchangeably herein. A "subject" may be a biological entity that contains expressed genetic material. The biological entity may be, for example, a plant, an animal, or a microorganism, including bacteria, viruses, fungi, and protozoa. A subject may be tissues, cells, and their progeny of a biological entity obtained in vivo or cultured in vitro. A subject may be a mammal. A mammal may be a human. A subject may be diagnosed or suspected of being at high risk for a disease. In some cases, a subject is not necessarily diagnosed or suspected of being at high risk for a disease.
[0110] The term "epigenetic signature" is used to describe heritable and reversible chemical modifications to nucleic acids that are incorporated or removed by the biochemical machinery (enzymes) of the cell, as opposed to nucleic acid modifications introduced by chemicals or environmental factors. It also applies to chemical modifications to viral nucleic acids produced via viral recruitment of the host cell's enzymatic machinery and / or viral enzymes during the infection process.
[0111] The terms "metaepigenetic" and "metaepigenomic" are used to describe the combined analysis of epigenetic data, such as nucleic acid sequencing data, derived from the analysis of nucleic acids from multiple kingdoms of life. In these examples, the sequencing data is derived from enrichment of nucleic acids with one or more epigenetic features to enrich for nucleic acids with targeted epigenetic features.
[0112] The term "epigenetic writer" is used to describe an enzyme that performs the biochemical reaction(s) required to incorporate a specific nucleotide modification. For example, mammalian DNA methyltransferases are "epigenetic writers" that incorporate methyl groups at selected cytosine nucleotides in the genome.
[0113] The term "epigenetic reader" is used to describe a protein that is able to recognize an epigenetic mark and promote / modulate cellular or transcriptional events that depend on the recognition of the epigenetic mark in question.
[0114] The term "epigenetic eraser" is used to describe an enzyme that carries out the biochemical reaction(s) required to remove a specific nucleotide modification.
[0115] The term "taxonomic abundance" is used to describe the number of sequencing reads that can be assigned to a specified microbial taxon in each sample.
[0116] The term "cross-kingdom" is used to describe analyses that combine biological or molecular data or features from two or more taxonomic kingdoms (here: mammals, bacteria, archaea, fungi, and viruses).
[0117] The term "in vivo" describes events that take place inside a subject's body.
[0118] The term "ex vivo" describes an event that occurs outside of a subject's body. An ex vivo assay is not performed on a subject. Rather, it is performed on a sample separate from the subject. An example of an ex vivo assay performed on a sample is an "in vitro" assay.
[0119] The term "in vitro" is used to describe events that occur within a container for holding laboratory reagents such that the material is separated from the biological source from which it is obtained. In vitro assays can include cell-based assays in which live or dead cells are used. In vitro assays can also include cell-free assays in which no intact cells are used.
[0120] As used herein, the term "about" a number refers to that number plus or minus 10% of that number. The term "about" a range refers to that range minus 10% of its minimum value and plus 10% of its maximum value.
[0121] The use of absolute or sequential terms, such as "will," "will not," "shall," "shall not," "must," "must not," "first," "initially," "next," "consequently," "before," "after," "lastly," and "finally" are not intended to limit the scope of the embodiments disclosed herein, but are exemplary.
[0122] Any systems, methods, software, compositions, and platforms described herein are modular and not limited to sequential steps, and thus, terms such as "first" and "second" do not necessarily imply a priority, order of importance, or order of actions.
[0123] As used herein, the term "treatment" or "treating" is used in reference to a pharmaceutical or other intervention regimen to obtain a beneficial or desired result in a recipient. Beneficial or desired results include, but are not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit may refer to the eradication or amelioration of the condition or underlying disorder being treated. Therapeutic benefit may also be achieved with the eradication or amelioration of one or more of the physiological symptoms associated with the underlying disorder such that an improvement is observed in the subject, although the subject may still be afflicted with the underlying disorder. Prophylactic benefits include delaying, preventing, or eliminating the appearance of the disease or condition, delaying or eliminating the onset of symptoms of the disease or condition, slowing, halting, or reversing the progression of the disease or condition, or any combination thereof. For prophylactic benefit, subjects at risk of developing a particular disease or who report one or more of the physiological symptoms of the disease may be treated, even though they may not have been diagnosed with the disease.
[0124] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
Claims
1. 1. A method for determining a disease in a subject, comprising: (a) enriching one or more nucleic acid molecules from a biological sample provided by a subject by affinity targeting an epigenetic feature common to said one or more nucleic acid molecules; (b) sequencing the enriched one or more nucleic acid molecules to generate one or more nucleic acid molecule sequencing reads; (c) determining the disease in the subject as an output of a predictive model when the enriched one or more nucleic acid molecules are provided as input to the predictive model.
2. The method described in claim 1, wherein the disease includes cancer or a non-cancerous disease.
3. The method of claim 2, wherein the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, invasive breast carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid cancer, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof.
4. The method described in claim 2, wherein the non-cancerous disease comprises lupus erythematosus, type 2 diabetes, chronic obstructive pulmonary disease (COPD), sarcoidosis, or any combination thereof.
5. The method of claim 1, further comprising filtering the one or more nucleic acid molecule sequencing reads to identify one or more non-mammalian sequencing reads and one or more mammalian sequencing reads.
6. The method described in claim 1, wherein the epigenetic feature comprises a mammalian nucleic acid epigenetic feature or a non-mammalian nucleic acid epigenetic feature.
7. The method described in claim 2, wherein the predictive model is trained with features of one or more mammals, one or more non-mammals, or a combination thereof determined from one or more nucleic acid molecules of a biological sample of one or more subjects and a corresponding disease of the one or more subjects.
8. The method described in claim 2, wherein the predictive model outputs the type, subtype, stage, prognosis, or any combination thereof of the cancer when provided with nucleic acid sequencing reads of the subject's biological sample.
9. A method for training a predictive model, comprising: (a) enriching biological samples from one or more subjects with a disease by affinity targeting epigenetic features common to one or more nucleic acid molecules in the biological samples; (b) sequencing the enriched one or more nucleic acid molecules to generate one or more nucleic acid molecule sequencing reads; (c) training the predictive model with one or more features of the one or more nucleic acid molecule sequencing reads and the disease of the one or more individual subjects.
10. The method described in claim 9, wherein the epigenetic features include mammalian epigenetic features or non-mammalian epigenetic features.
11. The method of claim 9, wherein the one or more features include one or more disease features.
12. The method described in claim 9, wherein the epigenetic feature is a nucleic acid epigenetic feature comprising a methylated CpG dinucleotide pair, an unmethylated CpG dinucleotide pair, or a combination thereof.
13. The method described in claim 9, wherein the epigenetic feature is a nucleic acid epigenetic feature comprising the nucleic acid bases 5-methylcytosine, 5-hydroxymethylcytosine, N4-acetylcytosine, N6-methyladenine, or any combination thereof.
14. The method described in claim 10, wherein the non-mammalian nucleic acid epigenetic feature comprises phosphorothioate-linked nucleotides.
15. The method of claim 9, wherein the disease includes cancer and the predictive model outputs the type, subtype, stage, prognosis, or any combination thereof of the cancer when provided with nucleic acid sequencing reads of the biological sample from one or more other subjects.
16. The method described in claim 9, wherein the predictive model is configured to remove contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features.
17. A computer system for determining a disease in a subject, comprising: (a) one or more processors; (b) a non-transitory computer-readable storage medium containing software, the software causing the one or more processors of the computer system to: (i) receiving one or more nucleic acid molecule sequencing reads for one or more nucleic acid molecules of a biological sample of a subject, wherein the one or more nucleic acid molecules are enriched by affinity targeting of an epigenetic feature common to the one or more nucleic acid molecules; (ii) determining a disease in the subject as an output of a predictive model when the predictive model is provided with one or more nucleic acid molecule sequences.