Disease classifiers from targeted microbial amplicon sequencing
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2026-03-18
AI Technical Summary
The prior art is difficult to effectively distinguish and diagnose cancerous and non-cancer health conditions, especially in the use of microbial genomic characteristics.
The microbial genome characteristics of non-mammalians are analyzed through target amplification and sequencing technology, including microbial phylogenetic marker genes or fragments, and amplification and sequencing using polymerase chain reaction (PCR) and other technologies to generate feature sets to distinguish between cancer and non-cancer health conditions.
The ability to distinguish between cancer and non-cancer health conditions through microbial genomic characteristics is achieved, and the accuracy and efficiency of diagnosis is improved.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 318,479, filed March 10, 2022, which is incorporated herein by reference for all purposes. Summary of the Invention
[0002] Aspects of the present disclosure provide a method for generating a feature set for distinguishing between cancer and non-cancer health conditions of one or more subjects. In some embodiments, the method is based on targeted amplicon sequencing of one or more microbial genomic features. In some embodiments, the method includes: (a) providing one or more nucleic acids and corresponding health conditions of one or more subjects; (b) amplifying one or more genomic features of one or more non-mammalian nucleic acids of the one or more nucleic acids, thereby generating one or more amplified genomic features; (c) sequencing the amplified one or more genomic features to generate one or more non-mammalian sequencing reads; and (d) generating a feature set configured to distinguish between cancer and non-cancer health conditions by combining the abundance of the one or more genomic features of the one or more non-mammalian sequencing reads and the health conditions of the one or more subjects. In some embodiments, the genomic features include a microbial phylogenetic marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene includes a bacterial marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene includes a fungal marker gene or a marker gene fragment thereof. In some embodiments, the bacterial marker genes include ribosomal RNA gene 5S, ribosomal RNA gene 16S, ribosomal RNA gene 23S, bacterial housekeeping genes dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf, or any combination thereof. In some embodiments, the fungal marker genes include ribosomal RNA gene 18S, ribosomal RNA gene 5.8S, ribosomal RNA gene 28S, internal transcribed spacer regions 1 and 2, or any combination thereof. In some embodiments, the microbial phylogenetic marker genes include bacterial, fungal, or any combination thereof marker genes.In some embodiments, the amplifying comprises performing a polymerase chain reaction or a derivative thereof. In some embodiments, the derivative of the polymerase chain reaction comprises inverse PCR, anchored PCR, primer-directed rolling circle amplification, or any combination thereof. In some embodiments, the polymerase chain reaction comprises a blocking primer, a marker gene primer, or any combination thereof configured to prevent amplification of the one or more genomic features. In some embodiments, the one or more genomic features comprise a mitochondrial DNA genomic feature. In some embodiments, the blocking primer inhibits amplification of the mitochondrial DNA genomic feature. In some embodiments, the method further comprises enriching the one or more nucleic acids. In some embodiments, the one or more nucleic acids comprise mammalian, non-mammalian, or any combination thereof nucleic acids. In some embodiments, the enrichment of nucleic acids comprises: (a) combining one or more mammalian and non-mammalian nucleic acids with a hybridization probe, where the hybridization probe comprises a nucleic acid sequence complementarity to a non-mammalian genomic feature; (b) incubating the hybridization probe and one or more mammalian and non-mammalian nucleic acids under conditions that promote nucleic acid base pairing between the target nucleic acid feature and the hybridization probe; (c) separating unbound hybridization probe and hybridized probe bound to the non-mammalian nucleic acid; and (d) washing the hybridized probe bound to the non-mammalian nucleic acid, thereby producing one or more enriched non-mammalian nucleic acids. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and other reaction components. In some embodiments, the enrichment of one or more nucleic acids comprises non-mammalian DNA enrichment.In some embodiments, non-mammalian DNA enrichment comprises: (a) combining one or more mammalian and non-mammalian nucleic acids with one or more recombinant CXXC domain proteins to form a protein-DNA binding reaction; (b) incubating the protein-DNA binding reaction under conditions that promote interaction between the recombinant CXXC domain protein and the unmethylated CpG motifs of one or more mammalian or non-mammalian nucleic acids; (c) separating unbound recombinant CXXC domain protein and recombinant CXXC domain protein bound to unmethylated CpG nucleic acid fragments from the remainder of the protein-DNA binding reaction; and (d) washing the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid fragments, thereby generating one or more enriched nucleic acids for amplification. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and the remainder of the protein-DNA binding reaction components. In some embodiments, the one or more nucleic acids are derived from one or more biological samples of the one or more subjects. In some embodiments, the one or more biological samples include a biopsy sample of tissue, liquid, or any combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the one or more subjects comprise a human, a non-human mammal, or any combination thereof. In some embodiments, the mammalian and non-mammalian nucleic acids comprise DNA, RNA, microbial cell-free DNA, microbial cell-free RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the method comprises filtering one or more non-mammalian sequencing reads. In some embodiments, the filtering comprises filtering one or more non-mammalian sequencing reads to generate one or more mitochondrial DNA-depleted non-mammalian sequencing reads.In some embodiments, the filtering comprises mapping one or more mitochondrial DNA-depleted non-mammalian sequencing reads against one or more microbial reference databases to determine a microbial taxonomic identity of the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the method comprises decontaminating the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the decontamination comprises in silico purification. In some embodiments, the purification is configured to remove non-endogenous microbial sequencing reads, thereby generating a purified microbial taxonomic assignment and associated quantity of the sequencing reads. In some embodiments, the non-mammalian sequencing read mapping is performed in QIIME2 or other supported versions. In some embodiments, the one or more microbial reference databases comprise Greengenes, a database of bacterial 16S rRNA, SILVA, a database of bacterial, fungal, and archaeal rRNA, UNITE, a database of eukaryotic nuclear ribosomal ITS regions, a custom database derived from publicly available complete microbial genome sequences, or any combination thereof. In some embodiments, the abundance of one or more genomic features of the one or more non-mammalian sequencing reads comprises abundance of microbial functional genes, biochemical pathways, or any combination thereof. In some embodiments, the method comprises predicting metagenomic functional content of the purified microbial taxonomic assignments, thereby generating one or more functional abundances. In some embodiments, the prediction of metagenomic functional content is performed by PICRUSt2. In some embodiments, the cancer comprises lung, breast, ovarian, gastrointestinal, head and neck, liver, pancreatic, prostate, skin, or any combination thereof. In some embodiments, the lung cancer comprises non-small cell lung cancer. In some embodiments, the cancer comprises stage I, II, or III cancer. In some embodiments, the non-cancerous condition comprises health, disease, or any combination thereof.In some embodiments, the disease state comprises a pulmonary disease, the pulmonary disease comprises carcinoid, hamartoma, granuloma, interstitial fibrosis, emphysema, bronchitis, chronic obstructive pulmonary disease, pneumonia, sarcoidosis, or any combination thereof. In some embodiments, the method comprises generating a trained predictive model, where the trained predictive model is trained on one or more subject feature sets and health states. In some embodiments, the trained predictive model comprises a machine learning model, one or more machine learning models, an ensemble of machine learning models, or any combination thereof. In some embodiments, the trained predictive model comprises a regularized machine learning model. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a gradient boosting machine, a neural network, a support vector machine, k-means, a classification tree, a random forest, regression, or any combination thereof.
[0003] Aspects disclosed herein provide methods of using the output of the trained predictive model to diagnose a cancer or non-cancer health condition of one or more subjects. In some embodiments, the method includes: (a) providing one or more nucleic acids of one or more subjects; (b) amplifying one or more genomic features of the one or more non-mammalian nucleic acids, thereby generating one or more amplified genomic features; (c) sequencing the amplified one or more genomic features to generate one or more non-mammalian sequencing reads; and (d) outputting a diagnosis of a cancer or non-cancer health condition of the one or more subjects as a result of providing at least the one or more genomic features as input to the trained predictive model. In some embodiments, the non-mammalian nucleic acid includes a microbial nucleic acid. In some embodiments, the one or more nucleic acids are derived from one or more biological samples of the one or more subjects. In some embodiments, the one or more biological samples include a biopsy sample of tissue, liquid, or any combination thereof. In some cases, the liquid biopsy sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the one or more subjects include a human, a non-human mammal, or any combination thereof. In some embodiments, the one or more nucleic acid molecules include a total population of DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, cell-free microbial DNA, cell-free microbial RNA, or any combination thereof. In some embodiments, the one or more genomic features include a microbial phylogenetic marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene may include a bacterial marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene includes a fungal marker gene or a marker gene fragment thereof.In some embodiments, the bacterial marker genes include ribosomal RNA gene 5S, ribosomal RNA gene 16S, ribosomal RNA gene 23S, bacterial housekeeping genes dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf, or any combination thereof. In some embodiments, the fungal marker genes include ribosomal RNA gene 18S, ribosomal RNA gene 5.8S, ribosomal RNA gene 28S, internal transcribed spacer regions 1 and 2, or any combination thereof. In some embodiments, the microbial phylogenetic marker genes include bacterial, fungal, or any combination thereof marker genes. In some embodiments, the amplification comprises performing a polymerase chain reaction or a derivative thereof. In some embodiments, the derivative of the polymerase chain reaction comprises inverse PCR, anchored PCR, primer-directed rolling circle amplification, or any combination thereof. In some embodiments, the polymerase chain reaction comprises a blocking primer, a marker gene primer, or any combination thereof configured to prevent amplification of one or more genomic features. In some embodiments, the one or more genomic features comprise a mitochondrial DNA genomic feature. In some embodiments, the blocking primer inhibits amplification of a mitochondrial DNA genomic feature. In some embodiments, the method comprises enriching one or more nucleic acids. In some embodiments, the one or more nucleic acids comprise mammalian, non-mammalian, or any combination thereof. In some embodiments, the mammalian and non-mammalian nucleic acids comprise DNA, RNA, microbial cell-free DNA, microbial cell-free RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof.In some embodiments, the enrichment of nucleic acids comprises: (a) combining the one or more mammalian and non-mammalian nucleic acids with a hybridization probe, the hybridization probe comprising a nucleic acid sequence complementarity to a non-mammalian genomic feature; (b) incubating the hybridization probe and the one or more mammalian and non-mammalian nucleic acids under conditions that promote nucleic acid base pairing between the target nucleic acid feature and the hybridization probe; (c) separating unbound hybridization probe and hybridized probe bound to the non-mammalian nucleic acid; and (d) washing the hybridized probe bound to the non-mammalian nucleic acid, thereby producing one or more enriched non-mammalian nucleic acids. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and other reaction components. In some embodiments, the enrichment of the one or more nucleic acids comprises non-mammalian DNA enrichment. In some embodiments, non-mammalian DNA enrichment comprises (a) combining one or more mammalian and non-mammalian nucleic acids with one or more recombinant CXXC domain proteins to form a protein-DNA binding reaction, (b) incubating the protein-DNA binding reaction under conditions that promote interaction between the recombinant CXXC domain protein and one or more mammalian or non-mammalian nucleic acid unmethylated CpG motifs, (c) separating unbound recombinant CXXC domain protein and recombinant CXXC domain protein bound to unmethylated CpG nucleic acid fragments from the remainder of the protein-DNA binding reaction, and (d) washing the recombinant CXXC domain protein bound to unmethylated CpG nucleic acid fragments, thereby generating one or more enriched nucleic acids for amplification. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and such remainder of the protein-DNA binding reaction components.In some embodiments, the recombinant CXXC domain protein comprises recombinant zinc finger CXXC domain containing proteins KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, recombinant CXXC domains derived therefrom, or any combination thereof. In some embodiments, the method comprises filtering one or more non-mammalian sequencing reads. In some embodiments, the filtering comprises filtering one or more non-mammalian sequencing reads to generate one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the filtering comprises mapping one or more mitochondrial DNA-depleted non-mammalian sequencing reads against one or more microbial reference databases to determine a microbial taxonomic identity of the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the method comprises purifying one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the purification comprises in silico purification. In some embodiments, the cleanup is configured to remove non-endogenous microbial sequencing reads, thereby generating a cleaned microbial taxonomic assignment and associated quantity of sequencing reads. In some embodiments, the non-mammalian sequencing read mapping is performed in QIIME2 or other supported versions. In some embodiments, the one or more microbial reference databases include Greengenes, a database of bacterial 16S rRNA, SILVA, a database of bacterial, fungal, and archaeal rRNA, UNITE, a database of eukaryotic nuclear ribosomal ITS regions, a custom database derived from publicly available complete microbial genome sequences, or any combination thereof. In some embodiments, the one or more genomic features include abundances of microbial functional genes, biochemical pathways, or any combination thereof, of the one or more non-mammalian sequencing reads. In some embodiments, the method includes predicting metagenomic functional content of the cleaned microbial taxonomic assignment, thereby generating one or more functional abundances.In some embodiments, the metagenomic functional content is performed by PICRUSt2. In some embodiments, the cancer health condition comprises lung, breast, ovarian, gastrointestinal, head and neck, liver, pancreatic, prostate, skin, or any combination thereof. In some embodiments, the lung cancer comprises non-small cell lung cancer. In some embodiments, the cancer comprises stage I, II, or III cancer. In some embodiments, the non-cancerous condition comprises a non-cancerous condition that is healthy, diseased, or any combination thereof. In some embodiments, the disease condition comprises a lung disease, and the lung disease comprises carcinoid, hamartoma, granuloma, interstitial fibrosis, emphysema, bronchitis, chronic obstructive pulmonary disease, pneumonia, sarcoidosis, or any combination thereof. In some embodiments, the trained predictive model is trained on one or more subject feature sets and health conditions. In some embodiments, the trained predictive model comprises a machine learning model, one or more machine learning models, an ensemble of machine learning models, or any combination thereof. In some embodiments, the trained predictive model comprises a regularized machine learning model. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a gradient boosting machine, a neural network, a support vector machine, k-means, a classification tree, a random forest, regression, or any combination thereof.
[0004] Aspects disclosed herein provide a system for diagnosing a cancerous or non-cancerous health condition of one or more subjects. In some embodiments, the system includes (a) a processor; and (b) a non-transitory computer-readable storage medium including software configured to cause the processor to (i) receive one or more nucleic acid sequencing reads of one or more subjects of a biological sample of the one or more subjects, the one or more nucleic acid sequencing reads including one or more amplified genomic features of one or more non-mammalian nucleic acids, and (ii) provide at least one or more genomic features of the one or more non-mammalian nucleic acid sequencing reads as inputs to a trained predictive model, thereby outputting a diagnosis of a cancerous or non-cancerous health condition of the one or more subjects. In some embodiments, the non-mammalian nucleic acid may include a microbial nucleic acid. In some embodiments, the one or more biological samples include a biopsy sample of tissue, fluid, or any combination thereof. In some embodiments, the one or more subjects may include a human, a non-human mammal, or any combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the one or more nucleic acids comprise DNA, RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, cell-free microbial DNA, cell-free microbial RNA, or any combination thereof. In some embodiments, the genomic features may comprise a microbial phylogenetic marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene may comprise a bacterial marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene may comprise a fungal marker gene or a marker gene fragment thereof. In some embodiments, the bacterial marker gene comprises a ribosomal RNA gene. In some embodiments, the ribosomal RNA gene comprises a 5S, 16S, 23S, or any combination thereof ribosomal RNA gene.In some embodiments, the bacterial marker genes include ribosomal RNA gene 5S, ribosomal RNA gene 16S, ribosomal RNA gene 23S, bacterial housekeeping genes dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf, or any combination thereof. In some embodiments, the fungal marker genes include ribosomal RNA gene 18S, ribosomal RNA gene 5.8S, ribosomal RNA gene 28S, internal transcribed spacer regions 1 and 2, or any combination thereof. In some embodiments, the microbial phylogenetic marker genes may include bacterial, fungal, or any combination thereof marker genes. In some embodiments, the amplified one or more genomic features of the one or more non-mammalian nucleic acids are amplified by polymerase chain reaction or a derivative thereof. In some embodiments, the derivative of polymerase chain reaction includes inverse PCR, anchored PCR, primer-directed rolling circle amplification, or any combination thereof. In some embodiments, the polymerase chain reaction includes a blocking primer, a marker gene primer, or any combination thereof configured to prevent amplification of the one or more genomic features. In some embodiments, the one or more genomic features include a mitochondrial DNA genomic feature. In some embodiments, the blocking primer inhibits amplification of the mitochondrial DNA genomic feature. In some embodiments, the one or more nucleic acid sequencing reads include sequencing reads of one or more enriched nucleic acids. In some embodiments, the one or more nucleic acids may include mammalian, non-mammalian, or any combination thereof nucleic acids.In some embodiments, the one or more enriched nucleic acids are generated by (a) combining one or more mammalian and non-mammalian nucleic acids with a hybridization probe, where the hybridization probe comprises a nucleic acid sequence complementarity to a non-mammalian genomic feature, (b) incubating the hybridization probe and the one or more mammalian and non-mammalian nucleic acids under conditions that promote nucleic acid base pairing between the target nucleic acid feature and the hybridization probe, (c) separating unbound hybridization probe and hybridized probe bound to the non-mammalian nucleic acid, and (d) washing the hybridized probe bound to the non-mammalian nucleic acid, thereby generating one or more enriched non-mammalian nucleic acids. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and other reaction components. In some embodiments, the one or more enriched nucleic acids are generated by non-mammalian DNA enrichment. In some embodiments, non-mammalian enrichment comprises (a) combining one or more mammalian and non-mammalian nucleic acids with one or more recombinant CXXC domain proteins to form a protein-DNA binding reaction, (b) incubating the protein-DNA binding reaction under conditions that promote interaction between the recombinant CXXC domain protein and the unmethylated CpG motifs of one or more mammalian or non-mammalian nucleic acids, (c) separating unbound recombinant CXXC domain protein and recombinant CXXC domain protein bound to unmethylated CpG nucleic acid fragments from the remainder of the protein-DNA binding reaction, and (d) washing the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid fragments, thereby generating one or more enriched nucleic acids for amplification. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and remaining protein-DNA binding reaction components.In some embodiments, the recombinant CXXC domain protein comprises recombinant zinc finger CXXC domain containing proteins KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, recombinant CXXC domains derived therefrom, or any combination thereof. In some embodiments, the software configures the processor to filter one or more nucleic acid sequencing reads. In some embodiments, the filtering comprises filtering one or more sequencing reads to generate one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the filtering comprises mapping one or more mitochondrial DNA-depleted non-mammalian sequencing reads against one or more microbial reference databases to determine a microbial taxonomic identity of the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the software configures the processor to clean up the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the cleaning comprises in silico cleaning. In some embodiments, the cleanup is configured to remove non-endogenous microbial sequencing reads, thereby generating cleaned microbial taxonomic assignments and associated quantities of sequencing reads. In some embodiments, the mapping is performed in QIIME2 or other supported versions. In some embodiments, the one or more microbial reference databases include Greengenes, a database of bacterial 16S rRNA, SILVA, a database of bacterial, fungal, and archaeal rRNA, UNITE, a database of eukaryotic nuclear ribosomal ITS regions, custom databases derived from publicly available complete microbial genome sequences, or any combination thereof. In some embodiments, the one or more genomic features amplified include abundances of microbial functional genes, biochemical pathways, or any combination thereof, of one or more non-mammalian sequencing reads.In some embodiments, the metagenomic functional content prediction is performed on the cleaned microbial taxonomic assignments, thereby generating one or more functional abundances. In some embodiments, the software configures a processor to predict the metagenomic functional content of the cleaned microbial taxonomic assignments, thereby generating one or more functional abundances. In some embodiments, the metagenomic functional content prediction is performed by PICRUSt2. In some embodiments, the cancerous health condition comprises lung, breast, ovarian, gastrointestinal, head and neck, liver, pancreatic, prostate, skin, or any combination thereof. In some embodiments, the lung cancer comprises non-small cell lung cancer. In some embodiments, the cancerous condition comprises stage I, II, or III cancer. In some embodiments, the non-cancerous health condition comprises a non-cancerous condition, health, disease, or any combination thereof. In some embodiments, the disease condition may comprise a lung disease, the lung disease comprises carcinoid, hamartoma, granuloma, interstitial fibrosis, emphysema, bronchitis, chronic obstructive pulmonary disease, pneumonia, sarcoidosis, or any combination thereof. In some embodiments, the trained predictive model is trained on one or more genomic feature sets of one or more subjects and the health state of interest. In some embodiments, the trained predictive model comprises a machine learning model, one or more machine learning models, an ensemble of machine learning models, or any combination thereof. In some embodiments, the trained predictive model comprises a regularized machine learning model. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a gradient boosting machine, a neural network, a support vector machine, k-means, a classification tree, a random forest, a regression, or any combination thereof. In some embodiments, the cancerous health state comprises one or more types of cancer, one or more subtypes of cancer, a cancer stage, a cancer prognosis, or any combination thereof. In some embodiments, the cancerous or non-cancerous health state comprises a cancer or disease category, a tissue specific location, or any combination thereof.In some embodiments, the trained predictive model is used to predict the cancer therapy response of one or more subjects.In some embodiments, the trained predictive model is utilized to select the optimal therapy for one or more subjects.In some embodiments, the trained predictive model is utilized to longitudinally model the course of one or more cancers' response to therapy of one or more subjects, and then adjust the treatment regimen. In some embodiments, the cancerous condition may include acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, bile duct carcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the trained predictive model removes contaminating non-mammalian features while selectively retaining other non-contaminating non-mammalian features. The device may be configured to:
[0005] Aspects disclosed herein provide a method of generating a feature set for distinguishing types of cancer in one or more subjects, the method comprising: (a) providing one or more nucleic acids and corresponding health states of one or more subjects; (b) amplifying one or more genomic features of one or more non-mammalian nucleic acids of the one or more nucleic acids, thereby generating one or more amplified genomic features; (c) sequencing the amplified one or more genomic features to generate one or more non-mammalian sequencing reads; and (d) generating a feature set configured to distinguish types of cancer by combining an abundance of the one or more genomic features of the one or more non-mammalian sequencing reads with the health state of the one or more subjects. In some embodiments, the genomic features comprise a microbial phylogenetic marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene comprises a bacterial marker gene or a marker gene fragment thereof. In some embodiments, the microbial phylogenetic marker gene comprises a fungal marker gene or a marker gene fragment thereof. In some embodiments, the bacterial marker genes include ribosomal RNA gene 5S, ribosomal RNA gene 16S, ribosomal RNA gene 23S, bacterial housekeeping genes dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf, or any combination thereof. In some embodiments, the fungal marker genes include ribosomal RNA gene 18S, ribosomal RNA gene 5.8S, ribosomal RNA gene 28S, internal transcribed spacer regions 1 and 2, or any combination thereof. In some embodiments, the microbial phylogenetic marker genes include bacterial, fungal, or any combination thereof marker genes. In some embodiments, amplifying comprises performing a polymerase chain reaction or a derivative thereof.In some embodiments, the derivative of the polymerase chain reaction comprises inverse PCR, anchored PCR, primer-directed rolling circle amplification, or any combination thereof. In some embodiments, the polymerase chain reaction comprises a blocking primer, a marker gene primer, or any combination thereof configured to prevent amplification of the one or more genomic features. In some embodiments, the one or more genomic features comprise a mitochondrial DNA genomic feature. In some embodiments, the blocking primer inhibits amplification of the mitochondrial DNA genomic feature. In some embodiments, the method further comprises enriching the one or more nucleic acids. In some embodiments, the one or more nucleic acids comprise mammalian, non-mammalian, or any combination thereof nucleic acids. In some embodiments, the enrichment of nucleic acids comprises: (a) combining one or more mammalian and non-mammalian nucleic acids with a hybridization probe, where the hybridization probe comprises a nucleic acid sequence complementarity to a non-mammalian genomic feature; (b) incubating the hybridization probe and one or more mammalian and non-mammalian nucleic acids under conditions that promote nucleic acid base pairing between the target nucleic acid feature and the hybridization probe; (c) separating unbound hybridization probe and hybridized probe bound to the non-mammalian nucleic acid; and (d) washing the hybridized probe bound to the non-mammalian nucleic acid, thereby producing one or more enriched non-mammalian nucleic acids. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and other reaction components. In some embodiments, the enrichment of one or more nucleic acids comprises non-mammalian DNA enrichment.In some embodiments, non-mammalian DNA enrichment comprises: (a) combining one or more mammalian and non-mammalian nucleic acids with one or more recombinant CXXC domain proteins to form a protein-DNA binding reaction; (b) incubating the protein-DNA binding reaction under conditions that promote interaction between the recombinant CXXC domain protein and the unmethylated CpG motifs of one or more mammalian or non-mammalian nucleic acids; (c) separating unbound recombinant CXXC domain protein and recombinant CXXC domain protein bound to unmethylated CpG nucleic acid fragments from the remainder of the protein-DNA binding reaction; and (d) washing the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid fragments, thereby generating one or more enriched nucleic acids for amplification. In some embodiments, the washing is configured to remove non-specifically associated nucleic acids and the remainder of the protein-DNA binding reaction components. In some embodiments, the one or more nucleic acids are derived from one or more biological samples of the one or more subjects. In some embodiments, the one or more biological samples include a biopsy sample of tissue, liquid, or any combination thereof. In some embodiments, the liquid biopsy sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the one or more subjects comprise a human, a non-human mammal, or any combination thereof. In some embodiments, the mammalian and non-mammalian nucleic acids comprise DNA, RNA, microbial cell-free DNA, microbial cell-free RNA, cell-free DNA, cell-free RNA, exosomal DNA, exosomal RNA, or any combination thereof. In some embodiments, the method comprises filtering one or more non-mammalian sequencing reads. In some embodiments, the filtering comprises filtering one or more non-mammalian sequencing reads to generate one or more mitochondrial DNA-depleted non-mammalian sequencing reads.In some embodiments, the filtering comprises mapping one or more mitochondrial DNA-depleted non-mammalian sequencing reads against one or more microbial reference databases to determine a microbial taxonomic identity of the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the method comprises cleaning the one or more mitochondrial DNA-depleted non-mammalian sequencing reads. In some embodiments, the cleaning comprises in silico cleaning. In some embodiments, the cleaning is configured to remove non-endogenous microbial sequencing reads, thereby generating a cleaned microbial taxonomic assignment and associated quantity of the sequencing reads. In some embodiments, the non-mammalian sequencing read mapping is performed in QIIME2 or other supported versions. In some embodiments, the one or more microbial reference databases comprise Greengenes, a database of bacterial 16S rRNA, SILVA, a database of bacterial, fungal, and archaeal rRNA, UNITE, a database of eukaryotic nuclear ribosomal ITS regions, a custom database derived from publicly available complete microbial genome sequences, or any combination thereof. In some embodiments, the abundance of one or more genomic features of the one or more non-mammalian sequencing reads comprises the abundance of microbial functional genes, biochemical pathways, or any combination thereof. In some embodiments, the method includes predicting metagenomic functional content of the purified microbial taxonomic assignments, thereby generating one or more functional abundances. In some embodiments, the prediction of metagenomic functional content is performed by PICRUSt2. In some embodiments, the cancer comprises lung, breast, ovarian, gastrointestinal, head and neck, liver, pancreatic, prostate, skin, or any combination thereof. In some embodiments, the lung cancer comprises non-small cell lung cancer. In some embodiments, the cancer comprises stage I, II, or III cancer. In some embodiments, the method includes generating a trained predictive model, where the trained predictive model is trained on one or more subject feature sets and health states.In some embodiments, the trained predictive model comprises a machine learning model, one or more machine learning models, an ensemble of machine learning models, or any combination thereof. In some embodiments, the trained predictive model comprises a regularized machine learning model. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a gradient boosting machine, a neural network, a support vector machine, k-means, a classification tree, a random forest, regression, or any combination thereof.
[0006] Aspects of the disclosure provided herein describe a method of determining disease in a subject, comprising receiving a biological sample, electronic medical record information, and one or more radiological images of the subject, sequencing one or more nucleic acid molecules isolated from the biological sample, thereby generating one or more nucleic acid molecule sequencing reads, and determining disease in the subject as an output of the predictive model when the predictive model is provided with data derived from the one or more nucleic acid molecule sequencing reads, electronic medical record information, and one or more radiological images of the subject as input. In some embodiments, the method further comprises identifying one or more protein biomarkers from the biological sample of the subject. In some embodiments, the predictive model is provided with one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the biological sample comprises a liquid biopsy, a tissue biopsy, or any combination thereof. In some embodiments, the one or more radiological images include images of X-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of one or more nucleic acid molecules.In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD, lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine endometrial carcinoma, uveal melanoma, or and any combination thereof. In some embodiments, the method further comprises calculating one or more features of the one or more radiological images, wherein the one or more features of the one or more radiological images are provided as input to the predictive model. In some embodiments, the one or more features comprise a block cancer probability score, a diameter of the lesion, a spiculation of the lesion, a solidity of the lesion, or any combination thereof. In some embodiments, the method further comprises mapping or aligning the one or more nucleic acid sequencing reads to a genomic database to determine one or more human, non-human, or combinations thereof features of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the predictive model comprises a machine learning model.In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a distiller vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or any combination thereof. In some embodiments, the predictive model is trained with leave-one-out validation. In some embodiments, the predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises cleaning the one or more nucleic acid molecule sequencing reads to generate one or more cleaned sequencing reads. In some embodiments, the cleaning comprises in silico cleaning, experimental control cleaning, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun metagenomic sequencing, next generation sequencing, long read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject.In some embodiments, the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.
[0007] Another aspect of the disclosure provided herein describes a method comprising receiving one or more subject biological samples, electronic medical record information, data derived from one or more radiological images, and a corresponding disease; sequencing one or more nucleic acid molecules isolated from the biological samples, thereby generating one or more nucleic acid molecule sequencing reads; and identifying one or more features of the one or more nucleic acid molecule sequencing reads, electronic medical record information, and data derived from the one or more radiological images that correspond to the one or more subject diseases. In some embodiments, the identifying comprises aligning the one or more sequencing reads to a genomic database. In some embodiments, the method further comprises training a predictive model with the one or more subject nucleic acid molecule sequencing reads, electronic medical record information, and data derived from the one or more radiological images, and one or more features of the corresponding disease. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the method further comprises identifying one or more features of one or more protein biomarkers of the subject biological sample. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the biological sample comprises a liquid biopsy, a tissue biopsy, or any combination thereof. In some embodiments, the one or more radiological images comprise images of x-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing.In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules include mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney renal clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor. In some embodiments, the one or more radiographic features include tumors, thymoma, thyroid cancer, uterine carcinosarcoma, uterine endometrial cancer, uveal melanoma, or any combination thereof. In some embodiments, the one or more radiographic features include a block cancer probability score, a diameter of the lesion, a spiculation of the lesion, a solidity of the lesion, or any combination thereof. In some embodiments, the method further includes mapping or aligning the one or more nucleic acid sequencing reads to a genomic database to determine one or more human, non-human, or combinations thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database includes a human genomic database. In some embodiments, the predictive model includes a machine learning model.In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a distiller vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or any combination thereof. In some embodiments, the predictive model is trained with leave-one-out validation. In some embodiments, the predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, or stage III, or stage IV. In some embodiments, the method further comprises cleaning the one or more nucleic acid molecule sequencing reads to generate one or more cleaned sequencing reads. In some embodiments, the cleaning comprises in silico cleaning, experimental control cleaning, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules comprise features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancerous disease in the subject. In some embodiments, the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.
[0008] Another aspect of the disclosure provided herein describes a computer system configured to determine a disease in a subject, the computer system comprising: (a) one or more processors; and (b) a non-transitory computer readable storage medium comprising software, the non-transitory computer readable storage medium comprising executable instructions that, upon execution, cause the one or more processors of the computer system to: (i) receive one or more sequencing reads of a biological sample of the subject, electronic medical record information, and one or more images; and (ii) determine a disease in the subject as an output of the predictive model when the predictive model is provided with data from one or more nucleic acid molecule sequencing reads, electronic medical record information, and one or more radiological images of the subject as inputs. In some embodiments, the disease comprises a cancer or a non-cancerous disease. In some embodiments, the biological sample comprises a tissue biopsy, a liquid biopsy, or a combination thereof. In some embodiments, the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject. In some embodiments, the predictive model is provided with one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, or a combination thereof. In some embodiments, the predictive model is trained with data derived from one or more subject nucleic acid molecular sequencing reads, electronic medical record information, and one or more radiological images, and one or more features of the corresponding disease. In some embodiments, the executable instructions include identifying one or more features of one or more protein biomarkers in a biological sample of the subject. In some embodiments, the one or more protein biomarkers comprise carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof.In some embodiments, the one or more radiological images include images of x-ray, computed tomography (CT), low-dose computed tomography, magnetic resonance imaging (MRI), ultrasound, positron emission tomography, fluoroscopy, angiography, or any combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the one or more nucleic acid molecule sequencing reads include one or more amplicon-based 16S rRNA sequencing reads. In some embodiments, the amplicon-based 16S rRNA sequencing reads include sequencing reads of a V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecule sequencing reads include sequencing reads of mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD, lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, ribosome-binding markers ... The cancers include diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, cutaneous melanoma of the skin, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the one or more radiographic features include a block cancer probability score, a diameter of the lesion, a spiculation of the lesion, a solidity of the lesion, or any combination thereof.In some embodiments, the executable instructions further comprise mapping or aligning the one or more nucleic acid sequencing reads to a genome database to determine one or more human, non-human, or combinations thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genome database comprises a human genome database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a distiller vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or any combination thereof. In some embodiments, the predictive model is trained with leave-one-out validation. In some embodiments, the predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the executable instructions further comprise purifying the one or more nucleic acid molecular sequencing reads to generate one or more purified sequencing reads. In some embodiments, the purification comprises in silico purification, experimental control purification, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the one or more sequencing reads are generated by shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof. In some embodiments, the executable instructions further comprise determining one or more features of the one or more nucleic acid molecule sequencing reads.In some embodiments, the one or more features of the one or more nucleic acid molecules include non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancer disease of the subject. In some embodiments, the mapping or alignment is completed with Deblur, Bowtie2, Kraken, or any combination thereof.
[0009] Another aspect of the disclosure provided herein describes a method of determining a disease in a subject, comprising receiving a biological sample from a subject, sequencing one or more nucleic acid molecules of the biological sample, thereby generating one or more nucleic acid molecule sequencing reads, and determining a disease in the subject as an output of a predictive model when the one or more nucleic acid molecule sequencing reads of the subject are provided as input to the predictive model, wherein the predictive model is trained with the one or more nucleic acid molecule sequencing reads of the one or more liquid biological samples and the one or more tissue biological samples of the one or more subjects and the corresponding disease. In some embodiments, the disease comprises cancer, non-cancerous disease, or a combination thereof. In some embodiments, the method further comprises identifying one or more protein biomarkers from the biological sample of the subject. In some embodiments, the predictive model is provided with the one or more protein biomarkers from the biological sample of the subject. In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof.In some embodiments, the liquid biopsy comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD, lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, liver hepatocellular carcinoma, lymphoid tumors. In some embodiments, the method includes mapping or aligning one or more nucleic acid sequencing reads to a genomic database to identify one or more nucleic acid sequences that are provided as input to the predictive model. The method further comprises determining one or more human, non-human, or combinations thereof characteristics of the determined lead. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a diminution vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or any combination thereof. In some embodiments, the predictive model is trained with leave-one-out validation. In some embodiments, the predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV.In some embodiments, the method further comprises cleaning one or more nucleic acid molecule sequencing reads to generate one or more cleaned nucleic acid molecule sequencing reads, and the one or more cleaned nucleic acid molecules are provided as input to the predictive model. In some embodiments, the cleaning comprises in silico cleaning, experimental control cleaning, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancer disease of the subject. In some embodiments, the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
[0010] Another aspect of the disclosure provided herein describes a method of identifying one or more non-human genomic features, the method comprising receiving one or more liquid biological samples, one or more tissue biological samples, and corresponding diseases of one or more subjects, sequencing one or more nucleic acid molecules of the one or more liquid biological samples and the one or more tissue biological samples, thereby generating one or more sequencing reads, and identifying one or more non-human genomic features corresponding to the disease of the one or more subjects from the one or more sequencing reads. In some embodiments, the identifying comprises aligning or mapping the one or more sequencing reads to a genome database to determine one or more human, non-human, or combination thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the method further comprises training a predictive model with the one or more non-human genomic features of the one or more subjects and the corresponding disease. In some embodiments, the disease comprises cancer or a non-cancerous disease. In some embodiments, the method further comprises identifying one or more characteristics of one or more protein biomarkers of the one or more liquid biological samples, the one or more tissue biological samples, or combinations thereof. In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or a combination thereof. In some embodiments, the cancer comprises a tumor mass less than 3 centimeters in diameter. In some embodiments, the sequencing comprises amplicon-based 16S rRNA sequencing. In some embodiments, the amplicon-based 16S rRNA sequencing sequences the V6 region of the one or more nucleic acid molecules.In some embodiments, the one or more nucleic acid molecules comprise mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biological sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD, lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney renal clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinosarcoma, uterine endometrial carcinoma, uveal melanoma, or any combination thereof. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a suprasegment vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier. In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or any combination thereof. In some embodiments, the predictive model is trained with leave-one-out validation. In some embodiments, the predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof.In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the method further comprises purifying the one or more nucleic acid molecule sequencing reads to generate one or more purified sequencing reads. In some embodiments, the purification comprises in silico purification, experimental control purification, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the sequencing comprises shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof. In some embodiments, the method further comprises determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancer disease of the subject. In some embodiments, the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
[0011] Another aspect of the disclosure provided herein describes a computer system configured to determine a disease in a subject, the computer system comprising: (a) one or more processors; and (b) a non-transitory computer readable storage medium comprising software, the software comprising executable instructions, as a result of execution, that cause the one or more processors of the computer system to: (i) receive one or more sequencing reads of a biological sample of the subject; and (ii) determine a disease in the subject as an output of the predictive model when the one or more nucleic acid molecule sequencing reads of the subject are provided as an input to the predictive model, the predictive model being trained with the one or more nucleic acid molecule sequencing reads of one or more liquid biological samples and one or more tissue biological samples of the one or more subjects and the corresponding disease. In some embodiments, the disease comprises a cancer or a non-cancerous disease. In some embodiments, the executable instructions comprise receiving one or more protein biomarkers from the biological sample of the subject. In some embodiments, the predictive model is provided with one or more protein biomarkers from the biological sample of the subject. In some embodiments, the executable instructions include identifying one or more characteristics of one or more protein biomarkers of the subject's biological sample. In some embodiments, the one or more protein biomarkers include carcinoembryonic antigen, osteopontin, cancer antigen 15-3, cancer antigen 19-9, cancer antigen 125, interleukin-8, prolactin, cytokeratin 19 fragment (CYFRA21-1), MMP-9, sTNFRII, MMP-7, resistin, MPO, MCP-1, GRO, sVEGFR2, sKDR, sFlk-1, VEGF-A, VEGF-C, VEGF-D, HGF, CRp, MIF, PDGF, AB / bb, RANTES, SAA, TNFRII, or combinations thereof. In some embodiments, the cancer includes a tumor mass less than 3 centimeters in diameter. In some embodiments, the one or more nucleic acid molecule sequencing reads include 16S rRNA sequencing reads based on one or more amplicons.In some embodiments, the amplicon-based 16S rRNA sequencing reads include sequencing reads of the V6 region of one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecule sequencing reads include sequencing reads of mammalian RNA, mammalian DNA, mammalian cell-free DNA, mammalian cell-free RNA, mammalian exosomal DNA, mammalian exosomal RNA, non-human RNA, non-human DNA, non-human cell-free DNA, non-human cell-free RNA, non-human exosomal DNA, non-human exosomal RNA, circulating tumor DNA, circulating tumor RNA, or any combination thereof. In some embodiments, the liquid biological sample includes plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some embodiments, the cancer comprises lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), small cell lung cancer (SCLC), or any combination thereof. In some embodiments, the cancer comprises acute myeloid leukemia, adrenal cortical carcinoma, bladder urothelial carcinoma, brain low-grade glioma, breast invasive carcinoma, cervical squamous cell carcinoma and adenocarcinoma, cholangiocarcinoma, colon adenocarcinoma, esophageal carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, kidney chromophobe cell, kidney renal clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lymphoma diffuse large B-cell lymphoma, mesothelioma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, pheochromocytoma and paraganglioma, prostate adenocarcinoma, rectal adenocarcinoma, sarcoma, skin cutaneous melanoma, gastric adenocarcinoma, testicular germ cell tumor, thymoma, thyroid carcinoma, uterine carcinoma sarcoma, uterine corpus intrauterine In some embodiments, the executable instructions further comprise: mapping or aligning the one or more nucleic acid sequencing reads to a genomic database to determine one or more human, non-human, or combinations thereof characteristics of the one or more nucleic acid sequencing reads. In some embodiments, the genomic database comprises a human genomic database. In some embodiments, the predictive model comprises a machine learning model. In some embodiments, the predictive model comprises a neural network, a convolutional neural network, a logistic regression, a random forest, a dinner vector machine, or any combination thereof. In some embodiments, the machine learning model comprises a machine learning classifier.In some embodiments, the machine learning model comprises a stacked machine learning model, one or more machine learning models, an ensemble machine learning model, or any combination thereof. In some embodiments, the predictive model is trained with leave-one-out validation. In some embodiments, the predictive model is configured to determine a stage of the cancer, an anatomical origin of the cancer, or a combination thereof. In some embodiments, the stage of the cancer is stage I, stage II, stage III, or stage IV. In some embodiments, the executable instructions further comprise purifying the one or more nucleic acid molecule sequencing reads to generate one or more purified sequencing reads. In some embodiments, the purification comprises in silico purification, experimental control purification, or a combination thereof. In some embodiments, the predictive model determines the disease with an accuracy of at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the one or more sequencing reads are generated by shotgun sequencing, next generation sequencing, long read sequencing, or any combination thereof. In some embodiments, the executable instructions further comprise determining one or more features of the one or more nucleic acid molecule sequencing reads. In some embodiments, the one or more features of the one or more nucleic acid molecules include features of non-microbial taxonomic abundance, mammalian genome coordinates, annotated genomic loci, mammalian functional gene and / or biochemical pathway abundance, or any combination thereof, and a number of sequencing reads associated with the one or more features. In some embodiments, the predictive model is configured to distinguish between cancer and non-cancerous diseases of the subject. In some embodiments, the mapping or alignment is completed with Deblur, PICRUSt2, Bowtie2, Kraken, or any combination thereof.
[0012] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0013] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief description of the drawings]
[0014] [Figure 1] 1 shows a flow diagram of a microbial nucleic acid amplification and / or enrichment method described in some embodiments herein. [Diagram 2] 1 shows a schematic flow diagram of a microbial taxonomy calculation method described in some embodiments herein. [Diagram 3] 1 shows a flow diagram of the microbial function annotated computational method described in some embodiments herein. [Figure 4] 1 shows a flow diagram for a method for generating predictive model classifiers based on one or more microbial taxa from nucleic acid samples of healthy, cancer, and / or non-cancerous non-healthy subjects. [Diagram 5] 1 shows a flow diagram for a method for generating one or more microbial functional annotation predictive model classifiers from nucleic acid samples of healthy, cancer, and / or non-cancerous non-healthy subjects. [Figure 6] 1 illustrates a system configured to perform the methods of the present disclosure provided herein. [Figure 7A] 1 shows the 16S ribosomal RNA hypervariable region and corresponding 16S primers used to amplify the 16S region of phylogenetically diverse bacteria described in several embodiments herein. [Figure 7B] 1 shows the 16S ribosomal RNA hypervariable region and corresponding 16S primers used to amplify the 16S region of phylogenetically diverse bacteria described in several embodiments herein. [Figure 8] FIG. 1 shows a schematic diagram of a fungal ribosomal RNA gene cluster with an internal transcribed (ITS) region as described in some embodiments herein. [Figure 9] 1 shows experimental data for 16S ribosomal DNA amplification using the V6 primer pair described in some embodiments herein and a microbial DNA standard composition amplified with the V6 primer pair. [Figure 10] 1 shows experimental data for microbial 16S ribosomal DNA amplification using the V6 primer pair in the presence and absence of human genomic DNA. [Figure 11] 1 shows experimental data on the specificity of 16S ribosomal DNA amplification using the V6 primer pair described in some embodiments herein. [Figure 12A] 1 shows a flow diagram for 16S sequencing library preparation as described in some embodiments herein. [Figure 12B] 1 shows a flow diagram for Western validation as described in some embodiments herein. [Figure 13] 1 shows a flow diagram for the 16S sequencing process described in some embodiments herein. [Figure 14A] 1 shows experimental data sequencing read counts at various points throughout the 16S sequencing process described in some embodiments herein. [Figure 14B] 1 shows experimental data sequencing read counts at various points throughout the 16S sequencing process described in some embodiments herein. [Figure 15A] 4 shows receiver operating characteristic curves for predictive models trained in distinguishing between non-small cell lung cancer samples and non-cancer nucleic acid samples from one or more subjects described elsewhere herein. [Figure 15B] 4 shows receiver operating characteristic curves for predictive models trained in distinguishing between non-small cell lung cancer samples and non-cancer nucleic acid samples from one or more subjects described elsewhere herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] The disclosure provided herein describes methods and systems for determining, identifying, classifying, and / or generating one or more nucleic acid molecular features of one or more subjects that may differentiate, classify, and / or diagnose the health status of one or more subjects and / or a group of one or more subjects. In some cases, the one or more nucleic acid molecular features may be derived from, obtained, received, and / or determined from one or more nucleic acid molecules of one or more biological samples of a subject and / or a plurality of subjects. In some cases, the one or more nucleic acid molecules may include one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or a combination thereof. In some examples, the one or more non-mammalian nucleic acid molecules may include one or more nucleic acid molecules from bacteria, fungi, or a combination thereof. In some cases, the one or more subject health statuses described elsewhere herein may include a cancerous health status, a non-cancerous disease health status, a healthy health status, or a combination thereof. In some examples, the cancerous health status may include an individual having cancer. In some cases, the cancer may include cancer of the lung, breast, ovary, gastrointestinal, head and neck, liver, pancreas, prostate, skin, or any combination thereof. In some cases, the lung cancer may include non-small cell lung cancer. In some cases, the cancerous health condition may include a diagnosis of a stage of the cancer (e.g., stage I, stage II, stage II, etc.). In some cases, the health condition may include a spatial location (i.e., anatomical location) of the cancer and / or disease within a subject or subjects. In some cases, the biological sample may include a liquid biological sample, a tissue biological sample, or a combination thereof. In some cases, the non-cancerous disease health condition may include a lung disease. In some examples, the lung disease may include carcinoid, hamartoma, granuloma, interstitial fibrosis, emphysema, bronchitis, chronic obstructive pulmonary disease, pneumonia, sarcoidosis, or any combination thereof. In some cases, the liquid biological sample may include a liquid biopsy. In some cases, the liquid biopsy may include plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled breath condensate, or any combination thereof. In some cases, the tissue biosample may include a tissue biopsy of one or more regions, organs, and / or anatomical locations of a subject (e.g., lung, skin, liver, pancreas, brain, etc.).
[0016] In some cases, the amplification, enrichment, filtering, and / or purification of one or more nucleic acid molecules and / or one or more sequencing reads of one or more nucleic acid molecules may provide better-than-expected results if the corresponding one or more enriched, filtered, and / or purified nucleic acid molecule features determine, classify, identify, and / or diagnose a health condition of one or more subjects with an accuracy of at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 92%, at least about 94%, at least about 96%, at least about 98%, or at least about 99%.
[0017] In some cases, as seen in Figure 1, one or more subject biological samples may contain one or more microbial nucleic acid molecule compositions, one or more mammalian nucleic acid molecule compositions, or a combination thereof, 201. In some cases, one or more microbial nucleic acid molecules, one or more mammalian nucleic acid molecule compositions, or a combination thereof may be enriched via a microbial nucleic acid enrichment and amplification workflow 202. In some cases, the biological sample containing one or more microbial nucleic acid molecules and one or more mammalian nucleic acid molecules may be enriched by hybridization probe enrichment 203 and / or protein-based microbial DNA enrichment 204. In some cases, hybridization-based enrichment may include combining one or more mammalian nucleic acid molecules and one or more non-mammalian nucleic acid molecules with a hybridization probe, where the hybridization probe may comprise a nucleic acid sequence complementary to a non-mammalian genomic feature; incubating the hybridization probe, the one or more mammalian nucleic acid molecules, and the one or more mammalian nucleic acid molecules under conditions that promote nucleic acid molecule base pairing between the target nucleic acid feature of the one or more non-mammalian nucleic acid molecules and the hybridization probe; separating unbound hybridization probes and hybridized probes bound to the one or more non-mammalian nucleic acid molecules; and washing the hybridized probes bound to the one or more non-mammalian nucleic acid molecules, thereby generating one or more enriched non-mammalian nucleic acid molecules.
[0018] method In some examples, the disclosure provided herein describes a method of enriching one or more non-mammalian nucleic acid molecules (e.g., non-mammalian DNA). In some cases, enriching one or more non-mammalian nucleic acid molecules can be by protein-based non-mammalian (e.g., microbial) nucleic acid molecule enrichment 204. In some cases, non-mammalian DNA enrichment can include combining one or more mammalian nucleic acid molecules and one or more non-mammalian nucleic acid molecules with one or more recombinant CXXC domain proteins to form a protein-DNA binding reaction, incubating the protein-DNA binding reaction under conditions that promote interaction between the recombinant CXXC domain protein and the unmethylated CpG motifs of one or more mammalian nucleic acid molecules or one or more non-mammalian nucleic acid molecules, separating the unbound recombinant CXXC domain protein and the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid fragments from the remainder of the protein-DNA binding reaction, and washing the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid fragments, thereby generating one or more enriched nucleic acid molecules for amplification.
[0019] In some cases, the amplification may include marker gene amplification 205. In some cases, the marker gene may include a microbial phylogenetic marker gene or a marker gene fragment thereof. In some examples, the microbial phylogenetic marker gene may include a bacterial marker gene or a marker gene fragment thereof. In some cases, the microbial phylogenetic marker gene may include a bacterial marker gene or a marker gene fragment thereof. In some cases, the microbial phylogenetic marker gene may include a fungal marker gene or a marker gene fragment thereof. In some cases, the bacterial marker genes may include ribosomal RNA gene 5S, ribosomal RNA gene 16S, ribosomal RNA gene 23S, bacterial housekeeping genes dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf, or any combination thereof. In some examples, bacterial ribosomal RNA genes as shown in Figure 7A may include hypervariable regions (V1-V9) that can be utilized to distinguish and / or classify microbial taxa. In some cases, as shown in FIG. 7B, one or more forward and / or reverse primers can be used to amplify the 16S region of the bacterial ribosomal RNA gene to distinguish a phylogenetically diverse set of bacteria that can be used as a signature to distinguish, determine, and / or diagnose the health status of a subject and / or group of subjects.
[0020] In some examples, the fungal marker genes may include ribosomal RNA gene 18S, ribosomal RNA gene 5.8S, ribosomal RNA gene 28S, internal transcribed spacer regions 1 and 2, or any combination thereof. In some cases, as shown in FIG. 8, internal transcribed spacer regions 1 and 2 (ITS1 and ITS2, respectively) are located between the small ribosomal RNA (rRNA) and the large ribosomal RNA (rRNA), and represent 18S rRNA, 5.8S rRNA, and 28S RNA. In some cases, as described elsewhere herein, amplification and sequencing of the ITS1 and / or ITS2 regions provide genomic features and / or markers for detecting and / or determining the presence of one or more fungi in a biological sample of a subject and / or group of subjects. In some examples, the one or more fungi may provide taxonomic features that may differentiate, classify, and / or diagnose the health status of a subject and / or a group of subjects. In some cases, the ITS1 and / or ITS2 regions. In some cases, the amplification may be accomplished and / or completed by performing a polymerase chain reaction (PCR) or a derivative thereof. In some cases, the derivative of the polymerase chain reaction may include reverse primer PCR, inverse PCR, anchored PCR, primer-directed rolling circle amplification, or any combination thereof. In some examples, the polymerase chain reaction amplification may include blocking primers, marker gene primers, or a combination thereof configured to prevent amplification of one or more genomic features. In some cases, the one or more genomic features may include mitochondrial DNA genomic features.
[0021] In some cases, the enriched and / or amplified one or more nucleic acid molecules may be prepared for sequencing via sequence library preparation 300, as shown in FIG. 12A. In some cases, the method may include providing 302 the amplified and / or enriched one or more nucleic acids, coupling 304 a barcoded index sequence, and coupling 306 one or more adapter sequences to the barcoded index sequence. In some cases, the library of nucleic acid molecules may include a length of about 256 bp base pairs. In some cases, the amplified and / or enriched one or more nucleic acid molecules may include a length of about 90 bp. In some examples, the amplified one or more nucleic acid molecules may include cell-free DNA of a plasma biological sample amplified with V6 primers. As shown in the gel electrophoresis results of the exemplary experiment shown in FIG. 12B, the various lengths of the enriched and / or amplified nucleic acid molecules increase from before 308 library preparation to after 310 library preparation.
[0022] The resulting prepared library of one or more enriched and / or amplified nucleic acid molecule composition(s) of mammalian and / or non-mammalian nucleic acid molecules may then be sequenced by a targeted amplicon sequencing 206 method, for example, targeted microbial amplicon sequencing may be used in a microbial classification characteristic method 213 and / or a microbial function characteristic method 216, as shown in Figures 2 and 3, respectively. In some cases, targeted microbial amplicon sequencing may include microbial 16S amplicon sequencing. In some cases, one or more sequencing reads generated by sequencing one or more enriched and / or amplified nucleic acid molecule compositions may be pre-processed, as shown in Figure 13. In some examples, preprocessing may include processing one or more sequencing reads of the enriched and / or amplified nucleic acid composition via fastp to remove adapter sequences and perform quality control to generate one or more processed sequencing reads 312, generating sub-operational taxonomic units from the processed one or more sequencing reads and performing quality control 314, and querying the sub-operational taxonomic units against a genome database to assign one or more sub-operational unit classifications 316. In some cases, the quality control may include an average read quality of about 30 or at least about 30. In some cases, the genome database may include 16S GreenGenes13.8. In some cases, the Qiime2 sklearn classifier may be used to assign the sub-operational unit classifications. In some cases, the sub-operational taxonomic units may be generated using Deblur, a noise removal tool that models the error profile of a sequence based on quality scores, expected error rates, observed frequencies of each unique sequence, or a combination thereof. In some cases, targeted microbial amplicon sequencing 206 may include shotgun sequencing, next generation sequencing 207, sequencing by synthesis, or a combination thereof. In some cases, the microbial classification features 213 and / or microbial function features may be part of a set of one or more nucleic acid molecular features, as described elsewhere herein.In some cases, the microbial taxonomic characterization method 213 may determine one or more microbial taxonomic assignments and associated microbial abundances of the enriched and / or amplified nucleic acid molecules. In some examples, the microbial functional characterization method may determine one or more microbial functional pathways of the enriched and / or amplified nucleic acid molecules.
[0023] In some cases, a microbial functional signature method 216 (FIG. 3) may include sequencing the enriched and / or amplified nucleic acid molecule library, e.g., using next generation sequencing, to generate a set of sequencing reads 207, filtering one or more nucleic acid molecule sequences (e.g., mitochondrial DNA) from the set of sequencing reads 208, thereby generating one or more mitochondrial DNA-depleted sequencing reads 209, identifying one or more microbial taxonomic assignments of the one or more mitochondrial DNA-depleted sequencing reads 210, purifying the one or more microbial taxonomic assignments 211, annotating and / or identifying one or more microbial functional features of the one or more purified microbial taxonomic sequencing reads 214, and outputting a feature set of one or more identified and / or annotated microbial functional features 215. In some cases, the one or more microbial functional features 215 may be used in combination with a known health state of the subject (217, 218, 219) to train a predictive model (e.g., a machine learning classifier), as shown in FIG. 5 described elsewhere herein. In some cases, the microbial functional features may include metagenomic functional features. In some cases, PICRUSt2 may determine and / or identify one or more metagenomic functional features of the one or more purified microbial taxonomic sequencing reads. In some cases, the microbial classification workflow may include mapping to determine microbial taxonomic assignments from the mitochondrial DNA-depleted sequencing reads 210. In some cases, the cleanup may include in silico cleanup. In some examples, the cleanup may remove one or more non-endogenous microbial sequencing reads, thereby generating one or more purified microbial taxonomic assignments and associated amounts of sequencing reads from one or more microbial taxonomic identities of the mitochondrial DNA-depleted non-mammalian sequencing reads.
[0024] In some cases, the microbial function method 213 may include, for example, sequencing the enriched and / or amplified nucleic acid molecule library using next generation sequencing to generate a set of sequencing reads 207, filtering one or more nucleic acid molecule sequences (e.g., mitochondrial DNA) from the set of sequencing reads 208, thereby generating one or more mitochondrial DNA-depleted sequencing reads 209, identifying one or more microbial taxonomic assignments of the one or more mitochondrial DNA-depleted sequencing reads 210, purifying the one or more microbial taxonomic assignments 211, and outputting one or more purified microbial taxonomic features 212 of the enriched and / or amplified nucleic acid molecule library. In some cases, the one or more microbial taxonomic features may be used in combination with the subject's known health status (217, 218, 219) to train a predictive model (e.g., a machine learning classifier), as shown in FIG. 5 described elsewhere herein. In some cases, the microbial classification workflow may include mapping to determine microbial taxonomic assignments from the mitochondrial DNA-depleted sequencing reads 210. In some cases, the mapping may include mapping one or more mitochondrial DNA-depleted nucleic acid molecule sequencing reads against one or more microbial reference databases to determine the microbial taxonomic identity of the mitochondrial DNA-depleted non-mammalian sequencing reads. In some cases, the mapping may be performed by QIME2 or other supported versions. In some cases, the one or more microbial reference databases may include Greengenes, a database of bacterial 16S rRNA, SILVA, a database of bacterial, fungal, and archaeal rRNA, UNITE, a database of eukaryotic nuclear ribosomal ITS regions, a custom database derived from publicly available complete microbial genome sequences, or any combination thereof. In some cases, the cleanup may include in silico cleanup.In some examples, the cleanup can remove one or more non-endogenous microbial sequencing reads, thereby generating one or more cleaned microbial taxonomic assignments and associated quantities of sequencing reads from one or more microbial taxonomic identities of the mitochondrial DNA-depleted non-mammalian sequencing reads.
[0025] In some cases, the one or more nucleic acid molecular features may include one or more nucleic acid molecular genomic features. In some cases, the genomic features may include a microbial phylogenetic marker gene or a marker gene fragment thereof. In some examples, the microbial phylogenetic marker gene may include a bacterial marker gene or a marker gene fragment thereof. In some examples, the microbial phylogenetic marker gene may include a fungal marker gene or a marker gene fragment thereof. In some cases, the one or more nucleic acid molecular features may include a feature, set of features, and / or feature cluster of one or more non-mammalian nucleic acid molecules (e.g., microbial nucleic acid molecules) described elsewhere herein.
[0026] In some cases, the one or more nucleic acid molecular features may be used to train 220 one or more predictive models 221 (e.g., machine learning classifiers) as described elsewhere herein and shown in Figures 4 and 5. In some cases, the one or more nucleic acid molecular features (213, 216) may include a microbial, bacterial, fungal, or combination thereof classification and / or functional classification and / or characterization of one or more nucleic acid molecules of a biological sample of a subject or subjects, as described elsewhere herein.
[0027] In some cases, a predictive model (e.g., a machine learning classifier) may be trained with one or more nucleic acid molecule features of one or more nucleic acid molecules of a biological sample of a subject with a known health condition of healthy 217, non-cancerous disease 219, or cancerous 218. In some cases, a predictive model may be trained 220 with one or more microbial taxonomic features 213 and one or more associated health conditions of the subject, as shown in FIG. 4. In some cases, a predictive model may be trained with one or more microbial functional features 216 and one or more associated health conditions of the subject, as shown in FIG. 5. In some cases, a trained predictive model 221 may include one or more classifiers (222, 223, 224) that may differentiate, classify, and / or diagnose one or more subject health conditions that were not included in the training of the predictive model. In some cases, the one or more classifiers may include a healthy vs. cancer health condition classifier 222, a cancerous vs. non-cancerous disease health classifier 223, a non-cancerous disease vs. healthy classifier, or any combination thereof.
[0028] Predictive Model The disclosed methods and systems may utilize or access external capabilities of artificial intelligence, predictive models, and / or machine learning trained on one or more nucleic acid molecular features that may classify, diagnose, and / or characterize the health status of a subject, a plurality of subjects, and / or a group of one or more subjects. In some cases, one or more nucleic acid molecular features (e.g., microbial functional features, microbial taxonomic features, etc.) described elsewhere herein may predict, classify, and / or identify cancer and / or non-cancerous disease of one or more subjects. In some cases, one or more nucleic acid molecular features may be used to train one or more predictive models described elsewhere herein. These features may accurately predict, classify, and / or characterize the health status, e.g., cancer, non-cancerous disease, disorder, or any combination thereof, of a subject, a plurality of subjects, and / or a group of one or more subjects. Using such predictive capabilities, health care providers (e.g., physicians) may make informed and accurate risk-based decisions, thereby improving the quality of care and monitoring provided to subjects with cancer, non-cancerous disease, disorder, or any combination thereof.
[0029] The methods and systems of the present disclosure may analyze the presence and / or abundance of microorganisms (e.g., abundance of a particular genus of microorganisms, taxonomy, microbial functional pathways). The presence and / or abundance of microorganisms may then be used to determine one or more nucleic acid molecular features, e.g., non-mammalian nucleic acid molecular features, that may predict cancer and / or non-cancerous disease in one or more subjects. In some cases, the methods and systems described elsewhere herein may train a predictive model using one or more nucleic acid molecular features indicative of a subject's health status, e.g., cancer and / or non-cancerous disease. In some cases, the trained predictive model may then be used to generate a likelihood (e.g., prediction) of cancer and / or non-cancerous disease for one or more subjects different from the one or more subjects utilized to train the predictive model. The trained predictive model may include an artificial intelligence-based model, such as a machine learning-based classifier, configured to process one or more nucleic acid molecular features from one or more nucleic acid molecules and / or enriched, filtered, and / or amplified one or more nucleic acid molecules to generate a likelihood that the subject(s) have cancer, a non-cancerous disease, or a disorder. The model may be trained using microbial taxonomic features or abundances of microbial functional pathways from one or more cohorts of subjects, such as cancer subjects, subjects with a non-cancerous disease, disease-free and non-cancer subjects, cancer subjects undergoing treatment for cancer, subjects undergoing treatment for a non-cancerous disease, or any combination thereof. In some cases, the predictive model may be trained to provide a treatment prediction for treating cancer in one or more subjects that are not part of the training dataset of the predictive model. Such a predictive model may output a treatment recommendation for one or more subjects that are not part of the training dataset when provided with an input of the presence and abundance of one or more microorganisms in a hybridization-enriched biological sample of a patient.
[0030] The predictive model may include one or more predictive models. The model may include one or more machine learning algorithms. Examples of machine learning algorithms may include support vector machines (SVM), naive Bayes classification, random forests, neural networks (such as deep neural networks (DNN)), recurrent neural networks (RNN), deep RNN, long short-term memory (LSTM) recurrent neural networks (RNN), gated recurrent units (GRU), gradient boosting machines, random forests, or other supervised learning algorithms or unsupervised machine learning, statistics, linear regression, k-nearest neighbors, k-means, decision trees, logistic regression, or any combination thereof. The model may be used for classification or regression. The model may also involve the estimation of an ensemble model consisting of multiple predictive models, for example, utilizing techniques such as gradient boosting in the construction of gradient boosting decision trees. The model may be trained using one or more training data sets including one or more nucleic acid molecular features, subject data, such as the subject's medical history, the subject's family's medical history, subject's vitals (e.g., blood pressure, pulse, temperature, oxygen saturation), or any combination thereof.
[0031] The predictive model may include any number of machine learning algorithms. In some embodiments, the random forest machine learning algorithm may be an ensemble of bagged decision trees. The ensemble may be at least about 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 250, 500, 1000 or more bagged decision trees. The ensemble may be up to about 1000, 500, 250, 200, 180, 160, 140, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2 or fewer bagged decision trees. The ensemble can be about 1-1000, 1-500, 1-200, 1-100, or 1-10 bagged decision trees.
[0032] In some embodiments, the machine learning algorithm may have various parameters, which may be, for example, a learning rate, a mini-batch size, a number of epochs to train, momentum, learning weight decay, or neural network layers.
[0033] In some embodiments, the learning rate may be between about 0.00001 and 0.1.
[0034] In some embodiments, the mini-batch size can be about 16-128.
[0035] In some embodiments, the neural network may include neural network layers. The neural network may have at least about 2 to 1000 or more neural network layers.
[0036] In some embodiments, the number of epochs for training may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 500, 1000, 10000, or more.
[0037] In some embodiments, the momentum can be at least about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or more. In some embodiments, the momentum can be up to about 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, or less.
[0038] In some embodiments, the learning weight decay may be at least about 0.00001, 0.0001, 0.001, 0.002, 0.003, 0.004, 0.005, 0.006, 0.007, 0.008, 0.009, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, or more. In some embodiments, the learning weight decay may be up to about 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, 0.001, 0.0001, 0.00001, or less.
[0039] In some embodiments, the machine learning algorithm may use a loss function, which may be, for example, a regression loss, a mean absolute error, a mean biased error, a hinge loss, an Adam optimizer, and / or a cross entropy.
[0040] In some embodiments, the parameters of the machine learning algorithm may be tuned with the aid of a human and / or a computer system.
[0041] In some embodiments, the machine learning algorithm may prioritize certain features. The machine learning algorithm may prioritize features that may be more relevant to detecting cancer, non-cancerous diseases, disorders, or any combination thereof. If a feature is classified more frequently than another feature in determining cancer, non-cancerous diseases, and / or disorders, the feature may be more relevant to detecting cancer, non-cancerous diseases, and / or disorders. In some cases, features may be prioritized using a weighting system. In some cases, features may be prioritized on a probability statistic based on the frequency and / or amount of occurrence of the feature. The machine learning algorithm may prioritize features with the help of a human and / or a computer system.
[0042] In some cases, machine learning algorithms may prioritize certain features to reduce computational cost, conserve processing power, conserve processing time, improve reliability, or reduce random access memory usage, etc.
[0043] The training dataset may be generated, for example, from one or more cohorts of subjects with a diagnosis of a common cancer, non-cancerous disease, or disorder. The training dataset may include one or more nucleic acid molecular features in the form of abundance taxonomic assignment features of microorganisms present in one or more subjects' biological samples and / or microbial functional pathway features of microorganisms present in the biological samples. The features may include one or more subjects' cancer diagnoses corresponding to the microbial features. In some cases, the features may include patient information such as the patient's age, the patient's medical history, other medical conditions, current or past medications, clinical risk scores, and time since last observation. For example, a set of features collected from a given patient at a given time point may collectively function as a signature that may indicate the patient's health state or status at that given time point.
[0044] The label may include, for example, a clinical outcome, such as the presence, absence, diagnosis, and / or prognosis of a cancer, a non-cancer disease, disorder, or a combination thereof, in a subject (e.g., a patient). Clinical outcomes may include treatment efficacy (e.g., whether the subject is a positive or negative responder to a cancer and / or disease-based treatment).
[0045] The input features may be structured by aggregating the data into bins, or alternatively by using one-hot encoding. The input may also include feature values or vectors derived from the aforementioned inputs, such as cross-correlations.
[0046] The training dataset can be constructed from the presence and / or abundance of one or more nucleic acid molar features, for example, one or more microbial taxonomic features, one or more microbial functional pathways, or combinations thereof, identified and / or classified from enriched and / or amplified nucleic acid molecules of a biological sample that are indicative of cancer, a non-cancerous disease, disorder, or any combination thereof.
[0047] The model may process the input features to generate output values including one or more classifications, one or more predictions, or a combination thereof. For example, such classifications or predictions may include a binary classification of the presence or absence of cancer, the presence of a non-cancerous disease, the presence of a disorder, or any combination of those classifications of the subject. In some cases, one or more predictive models and / or machine learning algorithms may classify the subject between a group of categorical labels (e.g., "no cancer, non-cancer disease and / or disorder", "definite cancer, non-cancer disease and / or disorder", and "probable cancer, non-cancer disease and / or disorder") and a score indicating the likelihood (e.g., relative likelihood or probability) of developing a particular cancer, non-cancer disease, and / or disorder, a "risk factor" for the presence of cancer, non-cancer disease and / or disorder, the likelihood of the patient's death, and a confidence interval for any numerical prediction. Various machine learning techniques may be cascaded such that the output of the machine learning techniques may also be used as input features to subsequent layers or subsections of the model.
[0048] To train a model (e.g., by determining model weights and correlations) for generating real-time classifications or predictions, the model can be trained using a training dataset and / or one or more training features described elsewhere herein. Such datasets and / or features can be large enough to generate statistically significant classifications or predictions. For example, a dataset can include one or more nucleic acid molecular features derived from sequencing data from the presence and / or abundance of fungal, viral, archaeal, bacterial, or any combination thereof microorganisms in one or more subject biological samples.
[0049] The dataset may be divided into subsets (e.g., separate or overlapping), such as a training dataset, a development dataset, and a test dataset. For example, the dataset may be divided into a training dataset that includes 80% of the dataset, a development dataset that includes 10% of the dataset, and a test dataset that includes 10% of the dataset. The training dataset may include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The development dataset may include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The test dataset may include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, leave-one-out cross-validation may be used. The training set (e.g., training data set) may be selected by random sampling of a set of data corresponding to one or more patient cohorts to ensure independence of sampling. Alternatively, the training set (e.g., training data set) may be selected by proportional sampling of a set of data corresponding to one or more patient cohorts to ensure independence of sampling.
[0050] To improve the accuracy of model predictions and reduce overfitting of the model, the data set may be expanded to increase the number of samples in the training set. For example, data expansion may include rearranging the order of observations in the training records. To accommodate a data set with missing observations, methods of imputing missing data such as forward filling, backfilling, linear interpolation, and multitask Gaussian processes may be used. The data set may be filtered or batch corrected to remove or reduce confounding factors. For example, a subset of subjects may be excluded within the database.
[0051] The model may include one or more neural networks, such as a neural network, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), or a deep RNN. The recurrent neural network may include units that may be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, the model may include an algorithmic architecture including a neural network with an input feature set, such as one or more nucleic acid molecular features, vital measurements, a subject's medical history, a subject's demographics, or any combination thereof, as described elsewhere herein. Neural network techniques such as dropout or regularization may be used during training of the model to prevent overfitting. The neural network may include multiple sub-networks, each of which is configured to generate classifications or predictions of different types of output information, which may be combined to form an overall output of the neural network. The machine learning model may alternatively utilize statistical or related algorithms, including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, and ensembles and gradient boosted variations thereof.
[0052] If the model generates a classification or prediction of cancer, a non-cancerous disease, a disorder, or a combination thereof, a notification (e.g., an alert or alarm) may be generated and transmitted to a healthcare provider, such as a doctor, a nurse, or other member of the subject's treatment team in a hospital. The notification may be transmitted via an automated phone call, a short message service (SMS), a multimedia message service (MMS) message, an email, and / or an alert in a dashboard. The notification may include output information such as a prediction of cancer, a non-cancerous disease, and / or a disorder; a predicted probability of the cancer, a non-cancerous disease, and / or a disorder; an expected time to onset of the cancer, a non-cancerous disease, and / or a disorder; a confidence interval of the probability or time, a recommended course of treatment for the cancer, a non-cancerous disease, and / or a disorder, or any combination of that information.
[0053] To validate the performance of the model, different performance metrics may be generated. For example, the area under the receiver operating characteristic curve (AUROC) may be used to determine the diagnostic, prognostic, screening, or any combination of their capabilities of the model. For example, the model may use adjustable classification thresholds so that the specificity and sensitivity are adjustable, and the receiver operating characteristic curve (ROC) may be used to identify different operating points that correspond to different values of specificity and sensitivity.
[0054] In some cases, such as when the dataset is not large enough, cross-validation may be performed to assess the robustness of the model across different training and testing datasets.
[0055] The following definitions may be used to calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), area under the precision recall curve (AUPR), AUROC, or the like. A "false positive" may refer to an outcome in which a positive outcome or result is generated incorrectly or prematurely (e.g., before or without the actual onset of a cancer, non-cancer disease and / or disorder). A "true positive" may refer to an outcome in which a positive outcome or result is correctly generated when a patient has a cancer, non-cancer disease and / or disorder (e.g., the patient exhibits symptoms of a cancer, non-cancer disease and / or disorder or the patient's records indicate a cancer, non-cancer disease and / or disorder). A "false negative" may refer to an outcome in which a negative outcome or result is generated but the patient has a cancer, non-cancer disease and / or disorder (e.g., the patient exhibits symptoms of a cancer, non-cancer disease and / or disorder or the patient's records indicate a cancer, non-cancer disease and / or disorder). A "true negative" can refer to an outcome in which a negative outcome or result is produced (e.g., prior to or without the actual onset of cancer, a non-cancerous disease and / or disorder).
[0056] The model may be trained until certain predefined conditions for accuracy or performance are met, such as having a minimum desired value corresponding to a diagnostic accuracy measure. For example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of occurrence of cancer, non-cancerous disease and / or disorder in a subject. As another example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of worsening or recurrence of a cancer, non-cancerous disease and / or disorder for which a subject has previously been treated. Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, AUPR, and AUROC, which correspond to diagnostic accuracy for detecting or predicting cancer, non-cancerous disease and / or disorder.
[0057] For example, such a predetermined condition can be that the sensitivity for predicting cancer, a non-cancerous disease and / or disorder includes values, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0058] As another example, such a predetermined condition can be a specificity for predicting cancer, a non-cancerous disease and / or disorder comprising a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0059] As another example, such a predetermined condition can be a positive predictive value (PPV) for predicting cancer, a non-cancerous disease and / or disorder including, for example, a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0060] As another example, such a predetermined condition can be a negative predictive value (NPV) for predicting cancer, a non-cancerous disease and / or disorder that includes a value of, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0061] As another example, such a predetermined condition may be that the area under the curve (AUC) (AUROC) of a receiver operating characteristic curve (ROC) predicting cancer, non-cancerous disease and / or disorder comprises a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0062] As another example, such a predetermined condition may be that the area under the precision-recall curve (AUPR) predicting cancer, non-cancer diseases and / or disorders includes a value of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0063] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0064] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0065] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with a positive predictive value (PPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0066] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases and / or disorders with a negative predictive value (NPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%.
[0067] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancerous diseases and / or disorders with an area under the curve (AUC) (AUROC) of a receiver operating characteristic curve (ROC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0068] In some embodiments, the trained model may be trained or configured to predict cancer, non-cancer diseases and / or disorders with an area under the precision-recall curve (AUPR) of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.
[0069] The training dataset may be collected from training subjects (e.g., humans), each having a diagnosis status indicating that they have been diagnosed with a biological condition or have not been diagnosed with cancer, a non-cancerous disease and / or disorder.
[0070] In some embodiments, the model is a neural network or a convolutional neural network, see Vincent et al., 2010, "Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion," J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, "Exploring strategies for training deep neural networks," J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.
[0071] In some embodiments, the data is de-dimensionalized using independent component analysis (ICA), such as that described in Lee, T.-W. (1998): Independent component analysis: Theory and applications, Boston, Mass: Kluwer Academic Publishers, ISBN 0-7923-8261-7, and Hyvaerinen, A.; Karhunen, J.; Oja, E. (2001): Independent Component Analysis, New York: Wiley, ISBN 978-0-471-40540-5, which are incorporated herein by reference in their entireties.
[0072] In some embodiments, the data is de-dimensionalized using principal component analysis (PCA), such as that described in Jolliffe, IT (2002). Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag. doi:10.1007 / b98835. ISBN 978-0-387-95442-4, which are incorporated herein by reference in their entireties.
[0073] SVM is an SVM, Cristianini and Shawe-Taylor. Theory, Wiley, New York, Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, and Furey. et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a given set of binary labeled data using a hyperplane that is maximally far from the labeled data. When linear separation is not possible, SVMs can work in conjunction with the technique of "kernels" that automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.
[0074] Decision trees are reviewed by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Decision tree based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that may be used is Classification and Regression Trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and Random Forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York. pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests-Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.
[0075] Clustering (e.g., unsupervised and supervised clustering model algorithms) are described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a data set. To identify natural groupings, two problems are addressed. First, determine how to measure the similarity (or dissimilarity) between two samples. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than samples in the other cluster. Second, determine a mechanism for splitting the data into clusters using the similarity measure. Similarity measures are discussed in section 6.7 of Duda 1973, which states that one way to begin a clustering study is to define a distance function and compute a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster will be significantly smaller than the distance between reference entities in different clusters. However, as described on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x,x') may be used to compare two vectors x and x'. Traditionally, s(x,x') is a symmetric function whose value is large if x and x' are "similar" in some way. An example of a non-metric similarity function s(x,x') is given on page 218 of Duda 1973. Once a method for measuring "similarity" or "dissimilarity" between points in a data set has been selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. A partition of the data set that extremizes a criterion function is used to cluster the data. See Duda 1973, page 217. Criterion functions are discussed in Duda 1973, section 6.8.More recently, Duda et al., Pattern Classification, 2nd edition, John Wiley & Sons, Inc. New York has been published. Pages 537-563 describe clustering in detail. Details of clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is incorporated herein by reference. Certain exemplary clustering techniques that may be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum of squares algorithms), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering. In some embodiments, the clustering comprises unsupervised clustering, where no preconceived notion of what clusters should form when a training set is clustered is imposed.
[0076] Regression models, such as those of the multi-category logit model, are described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, Inc., New York, Chapter 8, which is incorporated herein by reference in its entirety. In some embodiments, the model utilizes the regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, which is incorporated herein by reference in its entirety. In some embodiments, gradient boosting models are used, for example, for the classification algorithms described herein, and these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019). "Gradient Boosting". Hands-On Machine Learning with R. Chapman & Hall. pp. 221-245. ISBN 978-1-138-49568-5., which is incorporated herein by reference in its entirety. In some embodiments, ensemble modeling techniques are used, and these ensemble modeling techniques are described in the implementation of the classification models herein and in Zhou Zhihua (2012) Ensemble Methods: Foundations and Algorithms. Chapman and Hall / CRC. ISBN 978-1-439-83003-1, which is incorporated herein by reference in its entirety.
[0077] In some embodiments, the machine learning analysis is performed by a device executing one or more programs (e.g., one or more programs stored in non-persistent or persistent memory) that include instructions for performing the data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., a processing core) and a memory (e.g., one or more programs stored in non-persistent or persistent memory) that includes instructions for performing the data analysis.
[0078] Computer Systems The present disclosure provides a computer system programmed to implement the method of the present disclosure. Figure 6 shows a computer system 600 programmed or otherwise configured to predict the health status of one or more subjects of cancer, non-cancerous disease, or any combination thereof, train a predictive model as described elsewhere herein, generate a recommended treatment, or perform any combination thereof as described elsewhere herein. The computer system 600 may be a user's electronic device, or a computer system located remotely with respect to the electronic device. The electronic device may be a mobile electronic device.
[0079] The computer system 600 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 606, which may be a single-core or multiple-core processor, or multiple processors for parallel processing. The computer system 600 also includes memory or memory locations 604 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 602 (e.g., hard disk), a communication interface 608 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 610, such as cache, other memory, data storage devices, and / or electronic display adapters. The memory 604, the storage unit 602, the interface 608, and the peripheral devices 610 are in communication with the CPU 606 via a communication bus (solid lines), such as a motherboard. The storage unit 602 may be a data storage unit (or data repository) for storing data. The computer system 600 may be operatively coupled to a computer network ("network") 612 using the communication interface 608. Network 612 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, network 612 is a telecommunications network and / or a data network. Network 612 may include one or more computer servers that may enable distributed computing, such as cloud computing. Network 612 may implement a peer-to-peer network, in some cases using computer system 600, which may enable devices coupled to computer system 600 to operate as clients or servers.
[0080] CPU 606 may execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 604. The instructions may be directed to CPU 606, which may then be programmed or otherwise configured to implement the methods of the present disclosure, as described elsewhere herein. Examples of operations performed by CPU 606 may include fetch, decode, execute, and writeback.
[0081] The CPU 606 may be part of a circuit, such as an integrated circuit. One or more other components of the system 600 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0082] The storage unit 602 may store files such as drivers, libraries, and saved programs. The storage unit 602 can store user data, such as user preferences and user programs. In some cases, the computer system 600 may include one or more additional data storage units that are external to the computer system 600, such as located on a remote server that communicates with the computer system 600 through an intranet or the Internet.
[0083] Computer system 600 may communicate with one or more remote computer systems via network 612. For example, computer system 600 may communicate with a remote computer system of a user. Examples of remote computer systems may include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user may access computer system 600 via network 612.
[0084] The methods described herein may be implemented by machine (e.g., a computer processor) executable code stored on electronic storage locations of computer system 600, such as, for example, on memory 604 or electronic storage unit 602. Machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by processor 606. In some cases, the code may be retrieved from storage unit 602 and stored on memory 604 for immediate access by processor 606. In some circumstances, electronic storage unit 602 may be excluded and machine executable instructions are stored in memory 604.
[0085] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that may be selected to enable the code to be executed in a pre-compiled or as-compiled manner.
[0086] In some embodiments, the systems described elsewhere herein may include a system for diagnosing a cancerous or non-cancerous health condition of one or more subjects. In some cases, the system may include (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software configured to cause the one or more processors to: (i) receive one or more nucleic acid molecule sequencing reads of one or more subjects of a biological sample of the one or more subjects, the one or more nucleic acid molecule sequencing reads including sequences of one or more amplified genomic features of one or more non-mammalian nucleic acid molecules; and (ii) provide at least the one or more genomic features of the one or more non-mammalian nucleic acid sequencing reads as inputs to a trained predictive model, thereby outputting a diagnosis of a cancerous or non-cancerous health condition of the one or more subjects.
[0087] Aspects of the systems and methods provided herein, such as computer system 600, may be embodied in programming. Various aspects of the technology may be thought of as a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried or embodied in some type of machine-readable medium. The machine-executable code may be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or hard disk. A "storage" type medium may include any or all of the tangible memory of a computer, processor, etc., or their associated modules, such as various semiconductor memories, tape drives, disk drives, etc., and may provide non-transitory storage at any time for software programming. All or a portion of the software may be communicated, at times, over the Internet or various other communications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from a management server or host computer to a computer platform of an application server. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical terrestrial communications networks, and across various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0088] Thus, a machine-readable medium such as a computer executable code may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices of any computer(s) such as may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, a DVD or a DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium having a pattern of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, a cable or link transporting such a carrier wave, or any other medium from which a computer may read programming code or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0089] The computer system 600 may include or be in communication with an electronic display 616 that includes a user interface (UI) 614, for example, for providing a display for visualizing prediction results or an interface for training a predictive model. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0090] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be used in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0091] Although the method steps described and claimed herein represent each of the methods or sequences of actions according to the embodiments, one of ordinary skill in the art will recognize many variations based on the teachings described herein. Steps may be completed in different orders. Steps may be added or omitted. Some of the steps may include sub-steps. Many of these steps may be repeated as often as is beneficial.
[0092] One or more of the steps of each of the methods or sequences of operations may be implemented in one or more of the circuits described herein, e.g., a processor or logic circuitry, such as programmable array logic for a field programmable gate array. The circuitry may be programmed to provide one or more of the steps of each of the methods or sequences of operations, and the program may include, for example, program instructions stored in a computer readable memory, or programmed steps of a logic circuit, such as a programmable array logic or a field programmable gate array.
[0093] Example 1: 16S rDNA V6 primer amplification efficiency and specificity As can be seen in FIG. 9, the efficiency of 16S rDNA V6 amplification primers, e.g., 967F and 1064R, was evaluated and compared between human genomic and microbial standard dilution series. FIG. 9 shows polymerase chain reaction (PCR) cycles plotted against observed signal intensity of PCR reaction product production of various compositions: human genomic DNA 318, microbial standard dilution series 322, and negative control 320 where no DNA is present in the reaction. The microbial standard dilution series consisted of five standard dilutions with 16S copy numbers of 790, 7,896, 78,955, 789,554, and 7,895,540. From the PCR cycle plot, it can be seen that when only human genomic DNA is present 318, the V6 primers amplify human gDNA less efficiently compared to the microbial standard dilution 322. These results suggest that the V6 primers selectively amplify microbial nucleic acid molecules, thereby reducing noise or contaminant amplicons generated through amplification of host mammalian mitochondrial DNA.
[0094] FIG. 10 shows an experiment performed to evaluate the efficiency of 16S rDNA V6 amplification primers (e.g., 967F, 1064R) amplified in the presence of human gDNA. Four groups were prepared for PCR amplification with V6 primers: microbial DNA standards spiked into human genomic DNA from white blood cells (V6 / wbDNA) (324, 326, 328, 330, 332), microbial DNA standards (V6 / mbDNA) (338, 340, 342, 344, 346), human gDNA only (334), and a no-template control (i.e., negative control) (336). The human gDNA content used in the reaction was fixed at 5 ng, corresponding to approximately 1,396 genome equivalents. On a mass basis, human gDNA was present in each experimental group at more than 1000 times the amount of microbial DNA standards. The five microbial standards include 16S copy numbers of 790 (324, 338), 7,896 (326, 340), 78,955 (328, 342), 789,554 (330, 344), and 7,895,540 (332, 346). From the results shown in FIG. 10, it can be seen that although amplification was observed when only human gDNA was present in the PCR reaction, the human gDNA primer hybridization event does not prevent the specific amplification of microbial DNA when both DNA sources are mixed.
[0095] FIG. 11 shows an experiment that was performed to evaluate the specificity of 16S rDNA V6 primers (e.g., 967F, 1064R) for the amplification of microbial DNA in the presence of human DNA. Six experimental groups with varying amounts of human genomic DNA were prepared: 3 ng (410), 0.3 ng (402), 0.03 ng (408), 0.003 ng (414), 0.0003 ng (406), and 0 ng (400) and amplified in the presence of 5 pg of microbial DNA standard (7,895 genome equivalents). In addition, a no template control group (412) was also utilized as a negative control. From the PCR cycle plots shown in FIG. 37, it can be seen that at all levels of spiked human gDNA, microbial gDNA was preferentially amplified. In the absence of microbial DNA, the human DNA only control "hDNA 3 ng" curve amplifies with decreasing efficiency, resulting in a Cp of 31. These results indicated that the V6 primers were highly specific and not hindered by the prominent foreground of human DNA, as the PCR amplification curves were superimposed.
[0096] Example 2: Amplicon sequencing blank filtering Amplicon sequencing read counts at various steps of the workflow and / or methods described herein are shown in Figures 14A-14B. Figure 14B shows a zoomed-in view of the "OTU" and "OTU_filtered" sequencing reads / sample shown in Figure 14A. Experimental groups included plasma (500), blank negative control (502), and an industrial Zymo commercial sample of DNA microorganisms mixed at defined concentrations (504). A plot of the number of reads per sample shows the number of reads / sample at various points in the nucleic acid molecule sequencing reads described elsewhere herein: "raw_reads" are the total reads / sample before quality filters or taxonomic assignment, "qf_reads" are the sequencing reads remaining after a quality filtering step to remove PCR duplicates, "OTUs" correspond to the number of reads per sample that correspond to sub-operational taxonomic units identified via Deblur processing, and "OTUs_filtered" corresponds to the sOTUs remaining after subtracting the sOTUs present in the DNA extraction blank control (i.e., "blank"). Features with an abundance of at least 10 within the entire dataset were retained for further downstream processing.
[0097] From the results shown in Figures 14A-14B, after taking into account the contamination present in the DNA extraction process itself from subtracting the sequencing reads from the blank control, there are thousands of purified reads for further processing to determine one or more nucleic acid features for classifying and / or identifying one or more subject health conditions, as described elsewhere herein. 75.75% of the total abundance was removed from the plasma sample after blank filtering, which means that 12,509 microorganisms were removed from a total of 30,689 microorganisms (40.76% of the microorganisms were removed). Thus, what is detected and analyzed by the methods and systems described elsewhere herein is not DNA contamination, but rather, nucleic acid molecular signals present in the subject's biological sample.
[0098] Example 3: Performance of machine learning classifiers The machine learning classifiers were trained on 16S amplicon sequences (e.g., V6 hypervariable region) of one or more subjects with known health status markers, i.e., specific non-cancerous disease markers and / or stages of non-small cell lung cancer. The distribution of the number of subjects and / or samples with various labeled health statuses is shown in Figure 15A. Plasma from all subjects in each group was obtained and amplified using V6 16S primers, followed by next-generation sequencing as described elsewhere herein. The sequencing reads were then cleaned to identify one or more microbial taxonomic features. The microbial taxonomic features included the abundance of the identified microbial taxa down to the genus level. The associated read counts of the microbial taxonomic features were then used to train three random forest machine learning models using 5-fold cross-validation. The three random forest machine learning models included classifiers for classifying and / or characterizing cancer health status of stages I, II, and III. The performance receiver operating characteristic curves of the classifiers and associated areas under the curve (AUC), namely 0.891 for stage I, 0.71 for stage II, and 0.88 for stage III cancer, are shown in FIG. 15B.
[0099] Three separate machine learning classifiers for stage I (506), stage II (510), and stage III (508) cancer health status classification and / or identification were generated using shotgun sequencing of sequencing reads and Web of Life (WoL) genus relative abundance. The associated AUCs of the machine learning classifiers trained on shotgun sequencing and WoL abundance were 0.837 for stage I cancer, 0.66 for stage II cancer, and 0.871 for stage III cancer.
[0100] These results show that the V6 16S amplification primers providing one or more enriched and / or amplified nucleic acid molecules can be used to develop one or more microbial taxonomic signatures that provide high accuracy in distinguishing between stage I, stage II, and stage III cancers.
Claims
1. A method for diagnosing one or more target cancers, wherein the method is (a) To provide one or more nucleic acid molecules of one or more targets, (b) Amplifying one or more genomic features of one or more non-mammalian nucleic acid molecules of one or more nucleic acid molecules, thereby generating one or more amplified genomic features, (c) Sequence one or more of the amplified genomic features to generate one or more non-mammalian sequencing reads, (d) A method comprising providing at least one of the genomic features as input to a trained predictive model, thereby outputting a diagnosis of one or more cancerous or non-cancerous health conditions for the target.
2. The method according to claim 1, wherein the one or more non-mammalian nucleic acid molecules include microbial nucleic acids.
3. The method according to claim 1, wherein the one or more subjects include humans, non-human mammals, or any combination thereof.
4. The method according to claim 1, wherein the one or more nucleic acid molecules include DNA, RNA, cell-free DNA, cell-free RNA, exosome DNA, exosome RNA, cell-free microbial DNA, cell-free microbial RNA, or any combination thereof.
5. The method according to claim 1, wherein one or more genomic features include a microbial phylogenetic marker gene or a fragment of the marker gene.
6. The method according to claim 5, wherein the microbial phylogenetic marker gene comprises a bacterial marker gene or a fragment of the marker gene.
7. The method according to claim 5, wherein the microbial phylogenetic marker gene comprises a fungal marker gene or a fragment of the marker gene.
8. The method according to claim 6, wherein the bacterial marker gene includes ribosomal RNA gene 5S, ribosomal RNA gene 16S, ribosomal RNA gene 23S, bacterial housekeeping gene dnaG, frr, infC, nusA, pgk, pyrG, rplA, rplB, rplC, rplD, rplE, rplF, rplK, rplL, rplM, rplN, rplP, rplS, rplT, rpmA, rpoB, rpsB, rpsC, rpsE, rpsI, rpsJ, rpsK, rpsM, rpsS, smpB, tsf, or any combination thereof.
9. The method according to claim 7, wherein the fungal marker gene may include ribosomal RNA gene 18S, ribosomal RNA gene 5.8S, ribosomal RNA gene 28S, internally transcribed spacer regions 1 and 2, or any combination thereof.
10. The method according to claim 5, wherein the microbial phylogenetic marker gene comprises a marker gene of bacteria, fungi, or any combination thereof.
11. The method according to claim 1, wherein amplification comprises carrying out a polymerase chain reaction or a derivative thereof.
12. The method according to claim 11, wherein the derivative comprises inverse PCR, anchored PCR, primer-directed rolling circle amplification, or any combination thereof.
13. The method according to claim 11, comprising a blocking primer, a marker gene primer, or any combination thereof, configured to prevent the polymerase chain reaction from amplifying one or more genomic features.
14. The method according to claim 13, wherein one or more genomic features include mitochondrial DNA genomic features.
15. The method according to claim 1, comprising enriching one or more nucleic acid molecules.
16. The method according to claim 15, wherein the one or more nucleic acid molecules include one or more mammalian nucleic acid molecules, one or more non-mammalian nucleic acid molecules, or any combination thereof.
17. Enrichment is (a) Combining one or more mammalian nucleic acid molecules and one or more non-mammalian nucleic acid molecules with a hybridization probe, wherein the hybridization probe includes nucleic acid sequence complementarity to non-mammalian genome features, (b) Incubating the hybridization probe and the one or more mammalian nucleic acid molecules and the one or more non-mammalian nucleic acid molecules under conditions that promote nucleic acid base pairing between the target nucleic acid features of the one or more non-mammalian nucleic acid molecules and the hybridization probe, (c) Separating the unbound hybridization probe from the hybridized probe bound to one or more non-mammalian nucleic acid molecules, (d) Washing the hybridized probe bound to one or more non-mammalian nucleic acid molecules to generate one or more enriched non-mammalian nucleic acid molecules, the method according to claim 16.
18. The method according to claim 17, wherein the washing is configured to remove nonspecifically associated nucleic acid molecules and other reaction components.
19. The method according to claim 16, wherein the enrichment of one or more nucleic acid molecules includes enrichment of non-mammalian DNA.
20. The non-mammalian DNA enrichment is (a) Combining one or more mammalian nucleic acid molecules and one or more non-mammalian nucleic acid molecules with one or more recombinant CXXC domain proteins to form a protein-DNA binding reaction, (b) Incubating the protein-DNA binding reaction under conditions that promote interaction between the recombinant CXXC domain protein and the one or more mammalian nucleic acid molecules or the one or more non-mammalian nucleic acid unmethylated CpG motifs, (c) Separating the unbound recombinant CXXC domain protein and the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid molecule fragment from the remainder of the protein-DNA binding reaction. (d) The method according to claim 19, comprising washing the recombinant CXXC domain protein bound to the unmethylated CpG nucleic acid molecular fragment to generate one or more enriched nucleic acids for amplification.
21. The method according to claim 20, wherein the recombinant CXXC domain protein comprises recombinant zinc finger CXXC domain-containing proteins KDM2A, KDM2A, KDM2B, FBXL19, CFP1, DNMT1, MLL1, MLL2, MDB1, TET1, TET3, IDAX, CXXC5, CGBP, recombinant CXXC domains derived therefrom, or any combination thereof.
22. The method according to claim 20, wherein the washing is configured to remove nonspecifically associated nucleic acid molecules and the residue of the protein-DNA binding reaction.
23. The method according to claim 1, wherein the one or more nucleic acid molecules are derived from one or more biological samples of the one or more subjects.
24. The method according to claim 23, wherein the one or more biological samples include a tissue biopsy sample, a liquid biopsy sample, or any combination thereof.
25. The method according to claim 16, wherein the one or more mammalian nucleic acid molecules include DNA, RNA, cell-free DNA, cell-free RNA, exosome DNA, exosome RNA, or any combination thereof.
26. The method according to claim 16, wherein the one or more non-mammalian nucleic acid molecules include microbial DNA, microbial RNA, microbial cell-free DNA, microbial cell-free RNA, or any combination thereof.
27. The method according to claim 16, comprising filtering one or more non-mammalian sequencing reads.
28. The method according to claim 27, wherein filtering comprises filtering one or more non-mammalian sequencing reads to produce one or more mitochondrial DNA-depleted non-mammalian sequencing reads.
29. The method of claim 28, wherein filtering includes mapping the one or more mitochondrial DNA-depleted non-mammalian sequencing reads to one or more microbial reference databases to determine the microbial taxonomic identity of the one or more mitochondrial DNA-depleted non-mammalian sequencing reads.
30. The method according to claim 29, comprising purifying one or more mitochondrial DNA-depleted non-mammalian sequencing reads.
31. The method according to claim 30, wherein the purification includes in silico purification.
32. The method according to claim 30, wherein the purification is configured to remove non-endogenous microbial sequencing reads, thereby producing a relevant amount of purified microbial taxonomic assignment and sequencing reads.
33. The method according to claim 29, wherein the mapping is performed using QIIME2 or another supported version.
34. The method according to claim 29, wherein the one or more microbial reference databases include the bacterial 16S rRNA database Greengenes, the bacterial, fungal, and archaeal rRNA database SILVA, the eukaryotic ribosome ITS region database UNITE, a custom database derived from publicly available and complete microbial genome sequences, or any combination thereof.
35. The method according to claim 1, wherein the one or more genomic features include the abundance of a microbial functional gene, a biochemical pathway, or any combination thereof of the one or more non-mammalian sequencing reads.
36. The method according to claim 32, comprising predicting the metagenomic functional content of the purified microbial taxonomic allocation and thereby generating one or more functional abundances.
37. The method according to claim 36, wherein the prediction of the metagenomic functional content is performed by PICRUSt2.
38. The method according to claim 24, wherein the liquid biopsy sample comprises plasma, serum, whole blood, urine, cerebrospinal fluid, saliva, sweat, tears, exhaled condensate, or any combination thereof.
39. The method according to claim 1, wherein the cancerous health condition includes lung cancer, breast cancer, ovarian cancer, gastrointestinal cancer, head and neck cancer, liver cancer, pancreatic cancer, prostate cancer, skin cancer, or any combination thereof.
40. The method according to claim 39, wherein the lung cancer includes non-small cell lung cancer.
41. The method according to claim 1, wherein the non-cancerous state includes a non-cancerous state of health, disease, or any combination thereof.
42. The method according to claim 41, wherein the disease includes a lung disease, and the lung disease includes carcinoid, hamartoma, granuloma, interstitial fibrosis, emphysema, bronchitis, chronic obstructive pulmonary disease, pneumonia, or any combination thereof.
43. The method according to claim 1, wherein the health status of the cancer includes stage I, II, or III cancer.
44. The method according to claim 1, wherein the trained predictive model is trained on one or more feature sets and health conditions of a subject.
45. The method according to claim 1, wherein the trained predictive model includes a machine learning model, one or more machine learning models, an ensemble of machine learning models, or any combination thereof.
46. The method according to claim 1, wherein the trained predictive model includes a regularized machine learning model.
47. The method according to claim 45, wherein the machine learning model includes a machine learning classifier.
48. The method according to claim 45, wherein the machine learning model includes a machine learning model such as a gradient boosting machine, a neural network, a support vector machine, a k-means, a classification tree, a random forest, regression, or any combination thereof.