System and method for detection of disease
A database-independent microbial profiling platform using long-amplicon sequencing and machine learning algorithms addresses the limitations of conventional microbiome analysis, enabling accurate and early detection of systemic diseases and disease risk through unique genetic features.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-03-26
Smart Images

Figure IMGF000019_0001 
Figure 00000056_0000 
Figure 00000058_0000
Abstract
Description
[0001] Attorney Docket No. 24831-100810
[0002] SYSTEM AND METHOD FOR DETECTION OF DISEASE
[0003] CROSS-REFERENCE TO RELATED APPLICATIONS
[0004] This application is a PCT International Patent Application claiming priority to U.S. Nonprovisional Patent Application Ser. No. 18 / 445,677, filed September 23, 2024, and this application is designated as a continuation-in-part of said Application Ser. No. 18 / 445,677 in the U.S. This application further claims priority to U.S. Provisional Patent Application Ser. No. 63 / 798,263, filed May 1, 2025. Each of these applications are incorporated by reference herein in their entirety.
[0005] TECHNICAL FIELD
[0006] The present invention provides systems and methods for the detection of systemic disease from bacterial content samples, such as microbiome samples, obtained from a subject. The systems and methods are based on microbiome-based disease diagnostics using DNA sequencing and a machine learning algorithm. Also provided are methods for training a machine learning (“ML”) algorithm to correlate or associate patient microbiome sequence data with a disease state or the risk of developing a disease state. These systems and methods utilize a high-resolution, database-independent, high-throughput microbial profiling platform to diagnose systemic disease in patients or to identify those patients at risk of developing systemic disease. Also provided are kits for carrying out the methods.
[0007] BACKGROUND
[0008] There is an ongoing need to develop safe, reliable, and noninvasive systems and methods for detecting disease in a subject or for determining risk factors for or the propensity to develop the disease.
[0009] A prime example of this diagnostic need is in the medical area of colorectal cancer (“CRC”). As reported by the Centers for Disease Control (CDC), in 2019, 142,462 cases of colon and rectum cancer were reported, corresponding to an incidence rate of 36 per 100,000 standard population. Especially concerning is the trend of the increasing incidence of CRC amongst those under 50 years of age. Attorney Docket No. 24831-100810
[0010] Similarly, there is a need for earlier and more accurate diagnostic methods for a wide range of disease states and indications. The systems and methods disclosed herein can be applied to numerous conditions. Non-limiting examples of such diseases and indications include, but are not limited to: neurodegenerative diseases, Alzheimer’s Disease, Parkinson’s Disease, Amyotrophic Lateral Sclerosis (ALS), Multiple Sclerosis (MS), Lewy Body Dementia, Frontotemporal Dementia, Spinocerebellar Ataxia, autoimmune diseases, Celiac Disease, Crohn’s Disease, Ulcerative Colitis, Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis, Type 1 Diabetes, Hashimoto’s Thyroiditis, Graves' Disease, Psoriasis, Sjogren's Syndrome, Systemic Lupus Erythematosus (SLE), Myasthenia Gravis, Vasculitis, Pemphigus Vulgaris, Dermatomyositis, Guillain-Barre Syndrome, digestive disorders, Diverticulitis, Pancreatitis, Irritable Bowel Syndrome (IBS), Gastroesophageal Reflux Disease (GERD), Peptic Ulcer Disease, Non-Alcoholic Fatty Liver Disease (NAFLD), metabolic disorders, Type 2 Diabetes, Obesity, Hyperthyroidism, Hypothyroidism, cardiovascular diseases, Coronary Artery Disease, Hypertension (High Blood Pressure), Congestive Heart Failure, Stroke, Atherosclerosis, Renal (Kidney) disease, Chronic Kidney Disease (CKD), Polycystic Kidney Disease, Nephrotic Syndrome, Cancer, Lung Cancer, Breast Cancer, Prostate Cancer, Colon Cancer, Colorectal cancer, Early Onset Colorectal Cancer, Leukemia, Lymphoma, Pancreatic Cancer, Ovarian Cancer, Melanoma, Bladder Cancer, Liver Cancer, Kidney (renal cell and renal pelvis) Cancer, mental health disorders, Depression, Anxiety Disorders, Bipolar Disorder, Schizophrenia, Obsessive-Compulsive Disorder (OCD), Post- Traumatic Stress Disorder (PTSD), substance use disorders, Alcohol Use Disorder, Opioid Use Disorder, Nicotine Dependence, Chronic Obstructive Pulmonary Disease (COPD), Asthma, Fibromyalgia, Gout, Osteoarthritis, and Osteoporosis.
[0011] In the past, identification and quantification of microbes has been used as a direct measure of their presence and amount to identify the cause and severity of diseases, such as bacterial, fungal, or viral infections. Traditional microbiological diagnostics, originating with microscopy and cell culture, focused on identifying a single, causative pathogen responsible for an infectious disease. This paradigm has shifted toward understanding the role of the entire microbial community (the “microbiome”) on an individual’s overall health. However, the modem challenge is to identify complex patterns within the microbiome that can serve as biomarkers for a wide range of diseases, particularly non-infectious and chronic conditions. Attorney Docket No. 24831-100810
[0012] There is a need to improve the sensitivity of non-invasive tests to detect diseases or risk of diseases, particularly where early detection is important, such as, for example with colorectal and breast cancer.
[0013] Metagenomics and shotgun sequencing have been used in the past to identify and determine the relative prevalence of pathogenic bacteria from biological samples, such as gut microbiome samples from human patients. However, metagenomic and shotgun sequencing methods will not work efficiently and have limitations such as requiring a very large and impractical number of sequencing runs.
[0014] However, there is a more significant need to develop platforms and methods not for the direct detection of the microorganisms per se, but rather to determine and develop methods of microbiota analysis for screening and detection methods for disease, such as chronic or system disease. There have been microbiological methods in place to directly determine the presence of, e.g., a bacterial infection, since the early days of microbiology and microscopy in the late 1800s. However, there is a lack of platforms and methods to take advantage of and utilize the connection between microbial presence and the diagnosis of diseases, or the propensity to develop a disease.
[0015] Current molecular methods for microbiome analysis primarily fall into two categories. The first, amplicon sequencing, involves the targeted amplification and sequencing of conserved marker genes, e.g., the 16S rRNA gene, to profile a sample's composition. The second, shotgun metagenomics, involves sequencing all DNA in a sample, including from microbial and host sources. While both approaches provide valuable data, they possess significant drawbacks for widespread diagnostic use. Amplicon sequencing can lack the taxonomic resolution to distinguish between closely related species, while shotgun metagenomics generates vast amounts of data that are computationally intensive to analyze and requires deep, costly sequencing to accurately quantify low-abundance organisms.
[0016] Crucially, a fundamental limitation shared by both conventional approaches is their reliance on reference databases. Both methods require comparing experimental DNA sequences to databases of known microbial genomes to identify their source. This database dependency inherently limits the accuracy and scope of conventional methods, as reference databases are incomplete, contain errors, and often cannot resolve ambiguous alignments. This results in a significant loss of information and can confound the development of robust diagnostic signatures. Attorney Docket No. 24831-100810
[0017] While current technologies rely on identifying the DNA of the microbes of concern, the present invention does not rely on this. Instead, the present invention identifies a piece of the DNA, utilizing an amplicon, which is a piece of DNA (or RNA), which is distinctive for or identifies the microbe. In other words, the present invention does not need to identify full DNA sequences, but rather small DNA pieces. Because of this difference, it is not necessary to identify the underlying microbes per se. The power of the present invention is based on the methods it utilizes for training a machine learning algorithm to correlate patient microbiome sequence data with a disease state, where this is done without the need to identify the underlying microbes. The present invention is providing the diagnostics using subtle patterns of small DNA sequences, i.e. the presence, lack, or relative amounts, of these sequences in the sample, as indicators of the disease state or propensity or risk for contracting the disease state.
[0018] What the present invention provides is distinguished from directly identifying and quantifying microbes, e.g., bacteria, to determine an infection of the microbes. Instead, the present invention does not need to determine the presence and relative quantity of the microbes that may be present as a direct measure of a disease, e.g., an infection, but instead determines the presence and quantity of the microbes as a proxy for, or marker for, as a diagnostic method for identifying a systemic disease state, or the propensity for developing that disease state. In other words, the present invention is not interested in identifying the pathogenic microbes for the sake of the microbes themselves, but goes beyond to determine what their presence and / or quantity means as a determinant or predictor of a disease state, e.g., a systemic or chronic disease state. Additionally, rather than taxonomical classification, unique genetic features in microbiome bacteria are utilized which requires no prior knowledge of bacterial taxonomy. That is, the present invention does not associate taxonomical bacterial composition with a disease state, but rather accurately predicts a disease state based upon unique bacterial genetic feature associations with the disease state.
[0019] The present invention goes beyond the limitations of methods for directly determining a disease state or condition from the types and quantities of pathogenic microorganisms. Instead, the present platforms and methods bridge the gap to detecting or determining the risk for developing a disease or condition. The present invention goes beyond the limits of current diagnostic technologies. Therefore, the present invention provides powerful methods for the early detection of chronic and systemic disease states or the risk of developing these disease states. Attorney Docket No. 24831-100810
[0020] SUMMARY
[0021] The present invention provides systems and methods for the detection of systemic disease from bacterial content samples obtained from a subject. Also provided are methods for training a machine learning algorithm to correlate patient microbiome sequence data with a disease state. What is significant is that these systems and methods utilize a high-resolution, databaseindependent, high-throughput microbial profiling platform to diagnose systemic disease in patients or to identify those patients at risk of developing systemic disease. Also provided are kits for carrying out the methods for the detection of systemic disease.
[0022] The systems and methods disclosed herein address the limitations of conventional microbiome analysis by combining a novel long-amplicon sequencing strategy with a databaseindependent computational framework. An aspect of the invention is a method for generating information-rich amplicons spanning microbial genomic regions of varying conservation. In one embodiment, such an amplicon encompasses the 16S rRNA gene, the hypervariable Internal Transcribed Spacer (ITS) region, and a portion of the 23 S rRNA gene. These amplicons are then computationally decomposed into a comprehensive set of their constituent sequence fragments. In one embodiment, these fragments may be but are not limited to, short, overlapping subsequences of a fixed length, commonly known as k-mers. The presence and frequency of each unique fragment within a sample constitute a high-dimensional feature vector, a process that is fundamentally database-independent.
[0023] This feature vector, derived from the sequence fragments, serves as the direct input for a machine learning model. The model is trained to identify complex patterns and signatures within the feature data that are predictive of a specific disease state or condition. Significantly, this predictive model is developed without an intermediate step of taxonomic classification or sequence alignment to a reference database. In some embodiments, the method uses a long-amplicon design with a database-independent computational analysis that is capable of discovering novel biomarkers.
[0024] The present invention includes, but is not limited to, various aspects of the below numbered embodiments. In various further embodiments, these aspects may be combined with each other and / or with various aspects of the present disclosure. Attorney Docket No. 24831-100810
[0025] In an embodiment, the present invention provides a method for diagnosing or predicting the development of a disease state from microbiome sequence data from a prospective patient, comprising the following steps: a. collecting biological specimens and metadata from a first plurality (cohort) of patients having the disease state and from a second plurality (cohort) patient cohort lacking the disease state; b. generating microbiome sequence data from the biological specimens; c. processing the microbiome sequence data to generate features having a quantified relevance to the disease state for each patient from the first and second pluralities (cohorts) of patients; d. associating metadata with the generated features for each patient from the first and second pluralities (cohorts) of patients; e. selecting a subset of the features to generate a reduced feature set; f. training a machine learning algorithm on the reduced feature set to create a classification model that classifies a patient status as having the disease state, lacking the disease state, or being at risk for developing the disease state; g. obtaining microbiome sequence data and metadata from a prospective patient; h. quantifying the features in the reduced feature set from the microbiome sequence data of the prospective patient; and i. applying the classification model to the quantified features in the reduced feature set from the prospective patient to determine whether the prospective patient has or lacks the disease state, or is at risk for developing the disease state.
[0026] In an embodiment, the present invention provides a method, wherein the reduced feature set of step e comprises further expanding the reduced feature set to incorporate sequence set features correlative with the features of the reduced feature set.
[0027] In an embodiment, the present invention provides a method, wherein the sequence set features have a Pearson’s correlation coefficient of at least about 0.97 with the features of the reduced feature set. Attorney Docket No. 24831-100810
[0028] In an embodiment, the present invention provides a method, wherein the features can be matched and compared across patients from the first and second pluralities (cohorts) of patients.
[0029] In an embodiment, the present invention provides a method, further comprising applying data transformations to calibrate, normalize, or quantize the features for comparison across patients from the first and second pluralities (cohorts) of patients.
[0030] In an embodiment, the present invention provides a method, wherein the microbiome sequence data is a 16S rRNA gene and flanking upstream and downstream genomic regions, in part or in whole.
[0031] In an embodiment, the present invention provides a method, wherein the microbiome sequence data begins in or upstream of the 16S rRNA gene and extends past the end of the 16S rRNA gene as a contiguous amplicon sequence.
[0032] In an embodiment, the present invention provides a method, wherein the microbiome sequence data comprises one or more of 16S, ITS, and 23 S sequences.
[0033] In an embodiment, the present invention provides a method, wherein the microbiome sequence data comprises a 16S-ITS-23S amplicon.
[0034] In an embodiment, the present invention provides a method, wherein the quantified relevance to the disease state of each feature is defined to be the number of occurrences of the feature in the microbiome sequence data.
[0035] In an embodiment, the present invention provides a method, wherein the length of each feature is approximately 1 to 250 nucleotides.
[0036] In an embodiment, the present invention provides a method in which the biological specimens are fecal samples, blood samples, CSF samples, urine samples, saliva samples, other Attorney Docket No. 24831-100810 internal or external bodily fluids, skin swabs, gum swabs, vaginal swabs, or swabs of specific internal or external anatomical features.
[0037] In an embodiment, the present invention provides a method, further comprising: obtaining fecal immunochemical test data from one or more of the plurality (cohort) of patients having the disease state and the plurality (cohort) of patients lacking the disease state; and training the machine learning algorithm on the reduced feature set and the fecal immunochemical test data to create the classification model.
[0038] In an embodiment, the present invention provides a method, further comprising: obtaining fecal immunochemical test data from the prospective patient; and applying the classification model to the quantified features in the reduced feature set from the prospective patient and the fecal immunochemical test data from the prospective patient to determine whether the prospective patient has or lacks the disease state, or is at risk for developing the disease state.
[0039] In an embodiment, the present invention provides a method, in which the disease state is a neurodegenerative disease, an Alzheimer’s Disease, Parkinson’s Disease, Amyotrophic Lateral Sclerosis (ALS), Multiple Sclerosis (MS), Lewy Body Dementia, Frontotemporal Dementia, Spinocerebellar Ataxia, autoimmune disease, Celiac Disease, Crohn’s Disease, Ulcerative Colitis, Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis, Type 1 Diabetes, Hashimoto’s Thyroiditis, Graves' Disease, Psoriasis, Sjogren's Syndrome, Systemic Lupus Erythematosus (SLE), Myasthenia Gravis, Vasculitis, Pemphigus Vulgaris, Dermatomyositis, Guillain-Barre Syndrome, digestive disorder, Diverticulitis, Pancreatitis, Irritable Bowel Syndrome (IBS), Gastroesophageal Reflux Disease (GERD), Peptic Ulcer Disease, Non-Alcoholic Fatty Liver Disease (NAFUD), metabolic disorders, Type 2 Diabetes, Obesity, Hyperthyroidism, Hypothyroidism, cardiovascular disease, Coronary Artery Disease, Hypertension (High Blood Pressure), Congestive Heart Failure, Stroke, Atherosclerosis, Renal (Kidney) disease, Chronic Kidney Disease (CKD), Polycystic Kidney Disease, Nephrotic Syndrome, Cancer, Lung Cancer, Breast Cancer, Prostate Cancer, Colon Cancer, Colorectal Cancer (CRC or CA), Early Onset Colorectal Cancer, Leukemia, Lymphoma, Pancreatic Cancer, Ovarian Cancer, Melanoma, Attorney Docket No. 24831-100810
[0040] Bladder Cancer, Liver Cancer, Kidney (renal cell and renal pelvis) Cancer, mental health disorder, Depression, Anxiety Disorders, Bipolar Disorder, Schizophrenia, Obsessive-Compulsive Disorder (OCD), Post-Traumatic Stress Disorder (PTSD), substance use disorder, Alcohol Use Disorder, Opioid Use Disorder, Nicotine Dependence, Chronic Obstructive Pulmonary Disease (COPD), Asthma, Fibromyalgia, Gout, Osteoarthritis, and Osteoporosis.
[0041] In an embodiment, the present invention provides a method, in which the disease state is Colorectal Cancer (CRC or CA).
[0042] In an embodiment, the present invention provides a method of training a machine learning algorithm to correlate (associate) patient microbiome sequence data with a disease state: obtaining sequence data for a first plurality (cohort) of patients having a diagnosed disease state and for a second plurality (cohort) of control patients lacking the disease state, wherein the sequence data of the first and second pluralities (cohorts) of patients comprises respective computer-readable microbiome nucleotide sequences from biological samples collected from the respective patients from the first and second pluralities (cohorts) of patients; identifying sequence features from the microbiome nucleotide sequences which correlate positively or negatively with the disease state; generating machine learning training data comprising: i) at least a subset of the identified sequence features, ii) for each of the identified sequence features, their property of correlating positively or negatively with the disease state, and iii) retrospective patient data comprising computer-readable microbiome nucleotide sequences from biological samples collected from a plurality (cohort) of retrospective patients having the disease state and / or a plurality (cohort) of retrospective patients lacking the disease state; and training the machine learning algorithm with the machine learning training data to predict the presence or absence of the disease state in the retrospective patient data.
[0043] In an embodiment, the present invention provides a method, wherein training the machine learning algorithm produces a model capable of predicting the presence or absence of the disease state in prospective patients having no known disease state. Attorney Docket No. 24831-100810
[0044] In an embodiment, the present invention provides a method, wherein the method includes no taxonomic identification of bacterial strains in the microbiome nucleotide sequences.
[0045] In an embodiment, the present invention provides a method, wherein the microbiome nucleotide sequences comprise bacterial nucleotide sequences.
[0046] In an embodiment, the present invention provides a method, wherein the bacterial nucleotide sequences comprise one or more of 16S, ITS, and 23 S sequences.
[0047] In an embodiment, the present invention provides a method, wherein the bacterial nucleotide sequences comprise a 16S-ITS-23S amplicon.
[0048] In an embodiment, the present invention provides a method, wherein the training data includes no taxonomic identification bacterial strains from the 16S-ITS-23S amplicons.
[0049] In an embodiment, the present invention provides a method, wherein the sequence data and retrospective patient data are proportional to the bacterial populations in the underlying biological samples.
[0050] In an embodiment, the present invention provides a method, wherein the machine learning training data further comprises retrospective patient data comprising computer-readable microbiome nucleotide sequences from biological samples collected from a plurality (cohort) of retrospective patients lacking the disease state.
[0051] In an embodiment, the present invention provides a method, in which the disease state is a neurodegenerative disease, Alzheimer’s Disease, Parkinson’s Disease, Amyotrophic Lateral Sclerosis (ALS), Multiple Sclerosis (MS), Lewy Body Dementia, Frontotemporal Dementia, Spinocerebellar Ataxia, autoimmune disease, Celiac Disease, Crohn’s Disease, Ulcerative Colitis, Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis, Type 1 Diabetes, Hashimoto’s Thyroiditis, Graves' Disease, Psoriasis, Sjogren's Syndrome, Systemic Lupus Erythematosus Attorney Docket No. 24831-100810
[0052] (SLE), Myasthenia Gravis, Vasculitis, Pemphigus Vulgaris, Dermatomyositis, Guillain -Barre Syndrome, digestive disorder, Diverticulitis, Pancreatitis, Irritable Bowel Syndrome (IBS), Gastroesophageal Reflux Disease (GERD), Peptic Ulcer Disease, Non-Alcoholic Fatty Liver Disease (NAFLD), metabolic disorders, Type 2 Diabetes, Obesity, Hyperthyroidism, Hypothyroidism, cardiovascular disease, Coronary Artery Disease, Hypertension (High Blood Pressure), Congestive Heart Failure, Stroke, Atherosclerosis, Renal (Kidney) disease, Chronic Kidney Disease (CKD), Polycystic Kidney Disease, Nephrotic Syndrome, Cancer, Lung Cancer, Breast Cancer, Prostate Cancer, Colon Cancer, Colorectal Cancer (CRC or CA), Early Onset Colorectal Cancer, Leukemia, Lymphoma, Pancreatic Cancer, Ovarian Cancer, Melanoma, Bladder Cancer, Liver Cancer, Kidney (renal cell and renal pelvis) Cancer, mental health disorder, Depression, Anxiety Disorders, Bipolar Disorder, Schizophrenia, Obsessive-Compulsive Disorder (OCD), Post-Traumatic Stress Disorder (PTSD), substance use disorder, Alcohol Use Disorder, Opioid Use Disorder, Nicotine Dependence, Chronic Obstructive Pulmonary Disease (COPD), Asthma, Fibromyalgia, Gout, Osteoarthritis, and Osteoporosis.
[0053] In an embodiment, the present invention provides a method, in which the disease state is Colorectal Cancer (CRC or CA).
[0054] In an embodiment, the present invention provides a method, wherein the machine learning training data further comprises: iv) fecal immunochemistry test collected from at least one of the plurality (cohort) of retrospective patients having or lacking the disease state.
[0055] In an embodiment, the present invention provides a method, wherein the machine learning training data further comprises metadata from at least one of the plurality (cohort) of retrospective patients having or lacking the disease state.
[0056] In an embodiment, the present invention provides a system comprising; a computing device operable to execute computer-readable instructions, the computer-readable instructions being configured to perform the steps of: Attorney Docket No. 24831-100810 obtaining sequence data for a first plurality (cohort) of patients having a diagnosed disease state and for a second plurality (cohort) of control patients lacking the disease state, wherein the sequence data of the first and second pluralities (cohorts) of patients comprises respective computer-readable microbiome nucleotide sequences from biological samples collected from the respective patients from the first and second pluralities (cohorts) of patients; identifying sequence features from the microbiome nucleotide sequences which correlate positively or negatively with the disease state ; generating machine learning training data comprising: i) at least a subset of the identified sequence features, ii) for each of the identified sequence features, their property of correlating positively or negatively with the disease state, and iii) retrospective patient data comprising computer-readable microbiome nucleotide sequences from biological samples collected from a plurality (cohort) of retrospective patients having the disease state and / or a plurality (cohort) of retrospective patients lacking the disease state; and training a machine learning algorithm with the machine learning training data to predict the presence or absence of the disease state in the retrospective patient data.
[0057] In an embodiment, the present invention provides a kit for diagnosing or predicting the development of a disease state from microbiome sequence data from a prospective patient, comprising: a sample collector for obtaining biological specimens from a prospective patient and instructions for obtaining the biological specimens; wherein the collected biological specimens are useful for one or more of: a. generating microbiome sequence data from the biological specimens; b. processing the microbiome sequence data to generate features having a quantified relevance to the disease state for each patient; c. associating metadata with the generated features for each patient; d. selecting a subset of the features to generate a reduced feature set; e. training a machine learning algorithm on the reduced feature set to create a classification model that classifies the patient status as having the disease state, lacking the disease state, or being at risk for developing the disease state; Attorney Docket No. 24831-100810 f. obtaining microbiome sequence data and metadata from a prospective patient; g. quantifying the features in the reduced feature set from the microbiome sequence data of the prospective patient; and h. applying the classification model to the quantified features in the reduced feature set from the prospective patient to determine whether the prospective patients has or lacks the disease state, or is at risk for developing the disease state.
[0058] In an embodiment, the present invention provides a method for early detection of colorectal cancer, comprising: a. obtaining bacterial 16S-ITS-23 S sequence data from one or more bacterial microbiomes of a patient, wherein the 16S-ITS-23S sequence data comprises at least substantially complete 16S sequences for substantially all constituent bacteria of the bacterial microbiome; b. identifying the abundances (in some cases, the presence of absence) of unique features (i.e., nucleotide sequences) within the 16S-ITS-23 S sequence data that have predictive value (i.e., which may correlate positively or negatively with colorectal cancer) with respect to colorectal cancer risk or presence in the patient, wherein the features are associated with bacterial genera; and c. classifying based upon the identified features, by a classifier algorithm, whether the patient has colorectal cancer, whether the patient does not have colorectal cancer, or whether the patient is at risk for developing colorectal cancer.
[0059] BRIEF DESCRIPTION OF THE DRAWINGS
[0060] FIGs. 1A and IB show the addition of training sample data yields continuous increases in sensitivity and specificity on 100 blinded samples.
[0061] FIG. 1A shows the percent correct assignment of 100 blinded samples (y-axis), consisting of an unknown number of cancer (“CA”), advanced adenoma (“AA”), and normal controls. Samples were assigned to groups using classifiers determined by our machine learning algorithm after training on four training datasets: (i) ‘First Dataset’, (ii) ‘More Controls’, (iii) ‘More Controls and CA’, and (iv) ‘+FIT (fecal immunochemical test) data’. Training of our machine learning algorithms on the ‘First Dataset’ used fecal samples from 35 colorectal cancer, 35 advanced adenoma, and 178 non-cancer age-similar controls. In the ‘More Controls’ training dataset, identification of blinded samples was re-assessed after additional training using 43 additional Attorney Docket No. 24831-100810 control samples. After further ML training on 22 cancer and 35 additional controls in the ‘More Controls and CA’, sensitivity and specificity were again re-assessed. Finally, FIT data for each sample was added in the final analysis (+FIT). Changes in specificity of detecting normal controls, sensitivity of detecting CA, and sensitivity of detecting AA were plotted in dotted, dashed, and solid lines, respectively.
[0062] FIG. IB shows the increasing percentages of correct calls of the 100 blind samples after each machine learning training as compared to the results obtained by a currently marketed diagnostic product.
[0063] FIGs. 2A, 2B, 2C, and 2D show the location and frequency of colorectal cancer screen features. The starting point of each feature is shown on the X-axis, the number of features observed is shown on the Y-axis. The features were either positively or negatively correlated with colorectal cancer, advanced adenoma or normal samples. The starting point was measured from the start of primer 27f in the 16S rRNA gene. The 23 S rRNA gene is the final 600 bases on the right of the plot. Features in the Internally Transcribed Spacer (ITS) were plotted across 600 bases, although the ITS length varies.
[0064] FIG. 2A shows the frequency of each feature in the 16S-ITS-23S rRNA amplicon, i.e., the Titan-1™ Amplicon, of the present invention. The common primer positions used in 16S studies are shown in vertical lines (e.g., ‘27f , ‘515’, ‘800’, and ‘ 1492r’) approximately to scale. The 16S gene is represented by a striped arrow, with lighter variable regions (labeled VI to V9) and darker conservative regions (labeled Cl to C9), shown approximately to scale. The Internally ITS region between the 16S and 23 S genes is shown in the center between the 16S and 23 S regions, scaled to 600 bases, although many are shorter. The amplicon includes the first -600 bases of the 23 S rRNA gene, represented by the rightmost arrow.
[0065] FIG. 2B shows the 16S rRNA CRC feature count for the conserved and variable regions. The start and end position of each region is listed at left, counting from the start of the 27f primer. Columns detail the number of features, region size, and feature density. The 5 highest feature densities are all highlighted.
[0066] FIG. 2C shows a comparison of 16S rRNA conserved or variable total features, region size, and feature density.
[0067] FIG. 2D shows comparison of features, region size, and feature density of the 16S, ITS, and 23 S regions of the amplicon. Attorney Docket No. 24831-100810
[0068] FIGs. 3 A, 3B, 3C, and 3D show CRC-specific sequence features contribute strong positive and negative correlations with CRC, demonstrating differential abundance in CRC versus healthy samples. These correlations provide evidence of sample classification by disease status based on bacterial sequence features. The figures, which are in the form of bar graphs, show a small subset of bacterial sequence features detected as part of the ML / Al algorithm for CRC screening sort with CRC and non-CRC (Normal). The y-axis counts the number of samples, the shadings indicate the sample type. The sequence feature abundance is shown on the x-axis, using a log2 scale. Samples that did not have the signal are shown at ‘O’ Feature Abundance. The taxonomy / name summarizes the best taxonomy available for that feature after mapping the NCBI database.
[0069] FIG 3A shows data for fusobacterium;
[0070] FIG. 3B shows data for an uncultured bacterium partial 16S rRNA gene;
[0071] FIG. 3C shows data for an oral bacterium (labeled as Oral Bacterium 2); and
[0072] FIG. 3D shows data for a gut bacterial pathogen (labeled as Gut Bacterial Pathogen 1).
[0073] FIG. 4 shows that the bacterial profile is stable over 14 Months. Fecal samples from a single adult male, age 56 at the first time point, were sampled every few months over a 14-month period. Species above 1% relative abundance are shown. There was no significant change in health status, no antibiotic use, and no major change in diet over the 14-month period.
[0074] FIG. 5 shows a projection reaching 96% or greater sensitivity and specificity. An exponential curve was used to estimate that training the ML algorithm on 50 additional CA, 65 AA, and 200 control samples will achieve 96% or better sensitivity and specificity for CRC and AA. Existing commercial technology on the 100 blinded sample set provides a 76% sensitivity and 82% specificity. The current performance on these same samples is 72% sensitivity and 90% specificity. The present invention with additional samples will provide a projected performance of 95% sensitivity and 96% specificity.
[0075] FIG. 6 shows a principal component analysis (PCA) plot of bacterial biomarker sequence features from the fecal sample dataset of 101 CRC, 59 Advanced Colorectal Polyps (ACP), 57 non-advanced Colorectal Polyps, and 195 Controls with no lesion. Each point represents a fecal sample from an individual, with markers shaped according to diagnostic categories: CRC (circles), Advanced Colorectal Polyps (triangles), non-advanced Colorectal Polyps (squares), and Controls with no legions (crosses). The analysis provides evidence that bacterial biomarker sequence features stratify fecal samples by colorectal cancer risk. Attorney Docket No. 24831-100810
[0076] FIG. 7 shows a table of classifier algorithm results applied to 100 third-party blinded fecal samples not included in the training dataset. Sensitivity and specificity for CRC and ACP detection were reported by the third party, comparing the performance of the present invention (“Bacterial Biomarkers (Classifier Algorithm)” with a market leading fecal CRC test (“Existing Market Fecal CRC Assay”) on the same 100 fecal samples. For additional context, results from a CRC clinical trial are further summarized in FIG. 7.
[0077] FIG. 8A shows a radial target plot showing a representative microbiome profile from a subject prior to antibiotic treatment, displaying a highly diverse microbiome with many different species and strains. The second ring from the center shows the phylum-level diversity.
[0078] FIG. 8B shows a quantitative microbiome score of 100 / 100 displayed on a curved dial (gauge) chart.
[0079] FIG. 8C shows a resistome score of 1.28 plotted on a linear scale chart, with a typical reference range provided for comparison. The score reflects the level of antibiotic-resistant bacteria and estimates prior antibiotic use for the subject.
[0080] FIG. 8D shows an example medical report generated, which can be displayed on a graphical user interface or as a printed report combining the microbiome profile, with microbiome score (FIG. 8B), and resistome score (FIG. 8C) positioned above the target plot (FIG. 8A).
[0081] FIG. 8E shows an alternative medical report displayed in which the microbiome score and resistome scores are placed below the target plot. These medical reports provide complicated microbiome test results in a simplified, at-a-glance format that is understandable to both healthcare provider and patients.
[0082] FIG. 9A shows a target plot showing a microbiome profile from the sample subject as in FIG. 8, assessed after two rounds of antibiotic treatment (Nitrofurantion Mono-MCR lOOmg twice a day for 5 days). The microbiome exhibits reduced diversity, marked depletion of SCFA producers, and an upsurge in bacteria that are known to carry antibiotic resistance genes.
[0083] FIG. 9B shows a quantitative microbiome score of 51 / 100 displayed on a dial (gauge) chart.
[0084] FIG. 9C. Shows a resistome score of 5.63 plotted on a scale chart, with a typical reference range provided for comparison. The score reflects the level of antibiotic-resistant bacteria and estimates prior antibiotic use for the subject. Attorney Docket No. 24831-100810
[0085] FIG. 9D. Shows an example medical report generated, which can be displayed on a graphical user interface as a printed report combining the microbiome profde (FIG. 9A), microbiome score (FIG. 9B), and resistome score (FIG. 9C) positioned above the target plot.
[0086] FIG. 9E. Shows an alternative medical report displayed in which the microbiome score and resistome scores are placed below the target plot. These medical reports provide complicated microbiome test results in a simplified, at-a-glance format that is understandable to both healthcare provider and patients.
[0087] DETAILED DESCRIPTION
[0088] There is strong evidence that human-associated bacteria can be drivers of chronic disease, including multiple cancers, autoimmune diseases, inflammatory gut disorders, metabolic disease, and diabetes. To explore putative associations between chronic disease and bacterial content, a platform controlling all steps from bacterial lysis through analysis was developed. The platform’s sub strain-level, high-throughput bacterial profiling data were optimized for a databaseindependent custom machine learning (ML) algorithm, useful for hypothesis-free discovery of chronic disease / bacterial relationships.
[0089] Colorectal cancer (CRC, also referred to as cancer or CA) was selected as an area of exploration for the platform. Over a decade of evidence correlates the presence of oral bacteria such as Fusobacterium nucleatum with biofilm-related colon tissue invasion and chronic inflammation resulting in CRC. The present invention demonstrates application of the platform to colorectal cancer initiation and progression.
[0090] The following sections provide a detailed walkthrough of the inventive method, using the development of a diagnostic for CRC as a specific, non-limiting example.
[0091] A. Specimen Collection
[0092] In some embodiments, a total of 447 unique stool samples were collected from academic and commercial collaborators. Each sample was assigned to a category based on findings from a colonoscopy, the current gold standard for diagnosis. The categories included:
[0093] CRC (Colorectal Cancer): A patient diagnosed with colorectal cancer at any stage. Attorney Docket No. 24831-100810
[0094] • AA (Advanced Adenoma): A patient with a precancerous lesion determined to be an advanced adenoma, defined as an adenoma that is 1cm or larger in size, or contains villous features, high-grade dysplasia, or a sessile serrated morphology.
[0095] • EA (Early Adenoma): A patient with a precancerous adenoma that does not meet the criteria for an advanced adenoma.
[0096] • Benign Finding: A patient with a finding, such as a non-adenomatous polyp (e.g., hyperplastic polyp), that was determined to be benign and non-precancerous.
[0097] • No Lesion: A patient for whom no lesions or other abnormalities were detected during colonoscopy (Healthy Control).
[0098] In certain embodiments, in addition to the diagnostic category, the samples were provided with additional metadata, the availability of which varied by source.. Suchmetadata may include but is not limited to age, gender, calprotectin levels, and Fecal Occult Blood test (FOBT) or Fecal Immunochemical Test (FIT) results .
[0099] B. Microbiome Data Generation
[0100] In some embodiments, the biological samples were processed with a kit for carrying out the methods disclosed herein to prepare each sample for DNA sequencing. Prepared samples were sequenced on a Pacific Biosciences Sequel lie instrument. In some cases, technical replicates were performed by either preparing the sample multiple times or sequencing the prepared sample multiple times.
[0101] In one embodiment, only replicates containing 15,000 circular consensus HiFi reads after processing with Pacific Biosciences’ SMRT Analysis software were retained for further analysis. Table 1 below provides the number of unique samples and total replicates for each category. Attorney Docket No. 24831-100810
[0102] C. Feature Generation and Initial Selection
[0103] In some embodiments, feature extraction from the raw sequencing data and an initial feature selection were performed using our proprietary software. The software comprises three major modules. The first is an efficient counting of features, such as sequences of nucleotides or k-mers, which supports on-the-fly filtration of low-abundance features in order to limit the data size. The second is a normalization module that takes a matrix of samples and feature counts and transforms the counts to reduce sample-to-sample variation from technical sources. The third (final module) is a statistical test module that selects potentially relevant features from the normalized matrix.
[0104] In some embodiments, a classifier algorithm was trained to identify samples from CRC and AA patients as “positive” and samples from patients with no lesions or earlier stage neoplasia as “negative” by employing our software multiple times as described below.
[0105] Markers for Advanced Adenoma and Colorectal Cancer vs. All Other Categories
[0106] In certain embodiments, the samples from the CRC and AA groups were directly compared against the samples from all other groups combined. In one embodiment, 125 replicates from the CRC+AA group and 125 replicates from the other groups were randomly selected. In some embodiments, the feature counting proceeded one sample at a time; after processing every four samples, the software discarded any feature that has only a single count in the matrix. At the end of this initial processing, any feature with fewer than 14,300 total counts — representing less than 0.05% of the total feature counts across all samples — was also discarded. Such low-abundance markers cannot be reliably quantified and were safely removed as candidates.
[0107] In one embodiment, a value of 25 was utilized for k-mer size. Other experiments have shown the results to be relatively insensitive to this parameter, provided it is large enough to produce non-repetitive sequences from our amplicon. After filtration, 58,056,035 candidate features remained across the 250 samples.
[0108] In some embodiments, the normalization module was utilized with the following approaches: 1) each feature count was converted to “parts per million” (PPM) using the total Attorney Docket No. 24831-100810 number of reads in its respective sample, and 2) a pseudo-count of 1 was added to each PPM value, and a log-base-2 transformation was applied to reduce the dynamic range and help stabilize the variance.
[0109] In some embodiments, the feature selection module was applied. This module supports a substantial number of univariate statistics and selection methods, including but not limited to the following statistics: the Threshold Number of Misclassifications (TNoM), defined as the minimum number of incorrect label assignments achieved by an optimal numeric threshold on a normalized count of the feature, and Welch’s T-test. Examples of other tests implemented in the software include, without limitation the Wilcoxon Rank-Sum test (with a zero-inflation modification) and the Kolmogorov-Smirnov test. Since many statistics can be influenced by outliers, the module also supports bootstrapping, where the test is repeatedly applied to subsets of the samples and the median result is retained.
[0110] Once a statistic was calculated for each feature, a criterion for significance was established. Addressing the multiple comparisons problem was critical for this type of data; therefore, simplistic approaches like a fixed p-value threshold of 0.05 were avoided. Instead, for the t-test, a permutation-based False Discovery Rate (FDR) control method was applied. The sample labels were permuted 500 times and the t-statistic was calculated for 20,000 random features under each permutation. This process generated a distribution of t-statistics under the null hypothesis (i.e., when there is no true association with disease state). From this null distribution, a critical t-statistic value corresponding to a target FDR of 0.001 was established. Any features from the actual data with a t-statistic exceeding this critical value were retained. For the TNoM test, the mean and standard deviation of the TNoM statistic across were calculated for all features and those that were more than 3 standard deviations lower than the mean were retained, as a lower TNoM value indicates better classification performance.
[0111] Markers Between Other Subcategories
[0112] In some embodiments, additional markers were developed to differentiate between other specific categories.
[0113] The same feature selection techniques described previously were applied to find candidate features, such as, for example, sequence of nucleotides or k-mers for the following comparisons: Attorney Docket No. 24831-100810
[0114] AA (58 samples) vs. EA (57 samples)
[0115] AA (58 samples) vs. No Lesion (100 samples)
[0116] CRC (101 samples) vs. No Lesion (100 samples)
[0117] The result of the above process, after removing duplicates found across the different analyses, was a set of 5,675 unique features for further analysis.
[0118] D. Gathering Metadata and Features for All Replicates
[0119] In some embodiments, the feature selection analysis described above was performed on subsets of the data, both for computational efficiency and to allow for assessment on unseen samples. To proceed with classifier development, the 5,675 candidate features were quantified in all 879 study replicates, utilizing another proprietary software tool, SigQuant. In one embodiment, the software takes as input a list of signature elements, where each element is a list of one or more features, and the sequencing data associated with a sample. In the initial stage, each signature element consisted of a single candidate feature. In later analysis stages, multiple highly-correlated features can be combined into multi-feature signature elements for enhanced robustness. The software counts the features in each signature element, provides a final count for the element, e.g., the sum or average of the constituent features, and then performs normalization. As before, this normalization step involves converting the count to PPM, adding a pseudocount, and applying a log-base-2 transformation.
[0120] E. Feature Set Reduction
[0121] Following quantification with SigQuant, a matrix was produced comprising 879 sample rows and 5,675 biomarker columns. Metadata including but not limited to age, gender, calprotectin levels and FOBT or FIT results were utilized. In some embodiments, the candidate feature set was reduced and the "signature elements" was finalized for the SigQuant analysis. The feature list was the primary focus of the feature reduction efforts, with the merit of the auxiliary metadata to be determined later. Attorney Docket No. 24831-100810
[0122] For feature reduction and other machine learning tasks, the Weka Workbench (version 3.8) from the University of Waikato was employed. In certain embodiments, computationally efficient methods, including but not limited to Correlation-based Feature Subset Selection (CFS) and Classifier-based Attribute Evaluation, were utilized, reducing the feature set to approximately 1,000 nucleotide sequence features, e.g., k-mers, in this case. The feature set was further reduced by calculating all pairwise Pearson correlations and, for any pair of features with a correlation exceeding 0.97, retaining only a single representative feature. After these initial reductions, more complex strategies, including but not limited to Random Forest attribute ranking and Classifier Subset Evaluation were employed. The latter exemplary analysis uses a greedy hill-climbing heuristic with backtracking to find effective feature combinations for simple classifiers like logistic regression. Through successive rounds of these techniques, for instance, evaluating the classification performance at each stage using a Random Forest's out-of-bag error estimate, the feature set was distilled to 63 seed features.
[0123] The seed features were then expanded into robust signature elements. For each of the 63 seed features, other features from the original set, such as, for example, sequence of nucleotides or k-mers having a Pearson correlation coefficient above 0.98 were identified. In one embodiment, 20 such highly-correlated features were selected to form a "signature set" for each seed. The quantitative value for each signature set was calculated using an adaptive statistical method designed to be highly robust to outlier feature counts. In some embodiments, the method assessed the overall signal strength of the features within the set. For signature sets with a moderate-to-high signal, a robust outlier detection scheme based on the Median Absolute Deviation (MAD) was used to establish a range of trusted values. For sparse, low-signal sets, an alternative scheme based on the mean and standard deviation is employed to identify anomalous high-value counts. Any feature count outside the established range was labeled as an outlier and was replaced by a statistically imputed value (e.g., the median or zero). The final value for the signature set was then calculated as an aggregate, such as the sum or average, of the trusted and imputed counts, ensuring a stable and reliable measurement.
[0124] Once the signature sets were produced, SigQuant was used to quantify the 63 fully-formed signature sets in our 879 samples. FIGs. 3A-D provide binned histograms of four such features, with separate bars for the CRC and AA samples (“CA”) and the combined control categories (“No Attorney Docket No. 24831-100810
[0125] Lesion”), showing clear differences between the populations. During the quantification process, reads that contained features from each signature set were retained. Although the disclosed method does not depend on database matching, the reads containing signature features can be mapped to a database to assess what types of bacteria each signature set is counting. Sample reads for each signature set were mapped to a database ofl6S-ITS-23S and full-length 16S sequences. For each signature set, the predominant genus, phylum, species and the name of the deepest taxonomic level at which at least 50% of the reads share the same label could be identified. In some embodiments, such analysis may identify a combination of CRC-related bacteria from the literature, new potential associations, and bacteria that have yet to be precisely identified.
[0126] The approximate location of each feature relative to the beginning of the read sequence was recorded, providing insight into areas of the 16S-ITS-23S amplicon that contained diagnostically useful data. FIGs. 2A-D provide a depiction of the location information. The signature set of features lay across the full span of the 16S-ITS-23S amplicon, validating the need for the extended-length design.
[0127] F. Classifier Training
[0128] In certain embodiments, after feature reduction and quantification of the 63 signature sets across all replicates, classifier training was performed. For this task, samples were divided into two classes: a "positive" class comprising the CRC and Advanced Adenoma (AA) samples, and a "control" class comprising the Early Adenoma (EA), Benign, and No Lesion samples.
[0129] In some embodiments, metadata, including but not limited to age, gender, calprotectin levels, FIT / FOBT result, and (for a subset of samples) Cologuard test results were considered alongside the microbiome features.
[0130] A variety of classification algorithms were evaluated, including but not limited to Random Forests, Adaboost with decision stump and logistic base learners, logistic model trees, and J48 decision trees. Model performance was rigorously assessed using multiple validation strategies, including out-of-bag error estimates and k-fold cross-validation. Attorney Docket No. 24831-100810
[0131] In one embodiment, a Random Forest implemented in Weka, configured with 100 trees, a maximum tree depth of 8, and a bag size of 63% was selected as a model. Two versions of the model were generated: one that utilizes both the microbiome signature sets and the patient's FIT result, and another based purely on the microbiome signature sets.
[0132] G. Microbiome Profiling Platform for Detection of Colorectal Cancer and Advanced Pre- Cancerous Lesions Using Blinded Test Set.
[0133] In cancer biology, there is strong evidence that gut bacterial biofilms are important drivers of colorectal cancer (CRC, also referred to as cancer or CA) initiation and progression, in part through secretion of toxins that induce DNA damage, promote inflammation, and co-opt immune response in gut epithelial cells. For the first time, it has been demonstrated that sequence-specific bacterial signatures are present in fecal samples of patients with Advanced Colorectal Polyps (ACP) and CRC which appear to correlate with risk across the spectrum of disease. Sensitive and specific CRC and ACP detection was achieved using bacterial sequence biomarkers selected by the platform disclosed herein, which provides bacterial substrain-level data coupled with databaseindependent and customized machine learning (ML) analysis, enabling hypothesis-free discovery of CRC-to-fecal bacterial relationships.
[0134] Here, high-resolution bacterial profiling was evaluated to determine whether it can detect CRC-related bacterial substrain signals that could serve as a novel diagnostic for CRC and ACP, and to explore whether bacterial biomarkers in fecal samples could be used to monitor the initiation of precancerous lesions and progression of colorectal polyps to CRC.
[0135] Fecal samples comprising 101 CRC, 59 ACP, 57 non-ACP, and 195 controls (no lesions) were prepared using the kits that carry out the methods disclosed herein. These samples were subsequently sequenced using the PacBio Sequel lie as previously described. Sequence data were used to train the machine learning platform (ML), and CRC-related bacterial biomarker sequence features were filtered and curated to include only the most broadly applicable features from the training dataset. The platform was trained on well-characterized samples, demonstrating cross- validation accuracy of 84% for separation of CRC / ACP from other samples. There was a gradation of signal, where precancerous lesions span the boundary of normal and CRC, consistent with a Attorney Docket No. 24831-100810 continuum of disease progression. The bacterial biomarkers selected de novo by the ML represented a small, specific subset of the total bacterial community. Most common gut commensals were excluded, while the model preferentially weighted the presence of atypical taxa, including oral bacteria and opportunistic pathogens previously implicated in CRC, including Fusobacterium nucleatum animalis clade 2, Escherichia coli substrains, and other species known to participate in biofilm formation and tissue invasion. Notably, the ML identified previously undescribed bacterial biomarkers negatively associated with CRC.
[0136] These features were used to develop a classifier algorithm, which was separately applied to 100 blinded fecal samples provided by a third party. The blinded samples, which included CRC, ACP, non-ACP, and normal controls, were not part of the training dataset. In some embodiments, a blinded set of 100 new frozen stool samples with associated metadata was utilized to test the accuracy of the method disclosed herein. The test set contained a mixture of all sample subtypes. In certain embodiments, the kit carrying out the method disclosed herein was utilized and submitted the samples for sequencing on the Sequel lie, in the same manner as the training data.
[0137] Principal component analysis (PCA) of bacterial biomarker sequence features from the sample dataset revealed stratification of samples by diagnostic category (FIG. 6). In FIG. 6, each point represents a fecal sample from an individual, with markers shaped according to diagnostic categories: CRC (circles), Advanced Colorectal Polyps (triangles), non-advanced Colorectal Polyps (squares), and Controls with no legions (crosses). The distinct clustering observed indicates that bacterial biomarker sequence features separate individuals according to colorectal cancer risk.
[0138] H. Prospective Data Feature Quantification
[0139] In some embodiments, a strength of the method disclosed herein lies in the simplicity and efficiency of handling prospective samples. Concerting the microbiome sequencing data into feature set measurements suitable for input into the classifier algorithm is accomplished by running each sample’s sequencing data through SigQuant with the selected signature sets. In certain embodiments, a version of the data was prepared that included both FIT results and the microbiome signature measurements. Attorney Docket No. 24831-100810
[0140] I. Classifying the Blinded Test Data
[0141] In some embodiments, the two models described from the classifier training process were applied to the blinded data with and without FIT information. The results are summarized in Figure 1, which provides sensitivity and specificity for each model. On the blinded test set without using FIT, the latest model achieved 72% CRC sensitivity, 40% AA sensitivity, and 74% specificity. With the addition of FIT, the model achieved 72% CRC sensitivity, 36% AA sensitivity, and 90% specificity.
[0142] For comparison, a fully commercialized non-invasive CRC test achieved a sensitivity of 76% for CRC, 48% for AA, and a specificity of 82% on the same samples, lower than the typically advertised numbers. Earlier versions of the process based on fewer training samples were evaluated against the same blinded set. FIG. 5 shows the progression of sensitivity and specificity as the number of samples in each category increased. The plot indicates that the method disclosed herein continues on a strongly increasing trajectory. Accordingly, it is anticipated that accuracy will continue to improve as additional samples are incorporated, indicating the strength of the multivariate microbiome marker approach paired with the ensemble classifier versus older non- invasive CRC tests that are based on simple thresholds and small, fixed sets of biomarkers.
[0143] Classification of 100 third-party blinded fecal samples not included in training demon started that the bacterial biomarker classifier algorithm disclosed herein achieved sensitivity and specificity for CRC and ACP detection as independently reported. Performance was compared to a market-leading fecal CRC test using the same blinded samples, with results from the Freenome multi-omics CRC clinical trial provided as additional context (FIG. 7).
[0144] This study found that high-resolution bacterial sequencing enabled robust CRC diagnostics. Specifically, substrain-level bacterial sequence CRC signatures were present in fecal samples of patients with advanced colorectal polyps and colorectal cancer. Blinded sample accuracy matched current fecal CRC diagnostic sensitivity and specificity, demonstrating for the first time that high resolution bacterial sequence signatures can have comparable sensitivity and specificity to human -omics blood mutational / epigenetic markers. Additionally, the study showed biomarker signals that reflected disease progression. Such biomarkers enabled stratification of fecal samples across disease progression, from controls (no lesion) to advanced colorectal polyps, to CRC, enabling accessible and dynamic platform screening for earlier CRC risk assessment. This unique approach Attorney Docket No. 24831-100810 uncovered both known and novel microbial signatures. Database independent, hypothesis free analysis, unique to the methods disclosed herein, independently reinforced over a decade of CRC association with biofilms containing genotoxic pathogens and oral bacteria, while adding capability to improve discovery of previously uncharacterized bacteria with strong CRC association. Lastly, this study confirmed that early microbial signals may inform novel therapeutic intervention strategies. Detection of CRC-related bacteria in precancer suggested new opportunities for monitoring early intervention and risk reduction strategies at the microbial biofilm stage.
[0145] J. Further implementations
[0146] The machine learning (ML) algorithm was trained on sequence data from well characterized reference CRC, advanced adenoma (AA), and healthy control fecal samples. CRC and AA-associated sequence features discovered de novo by the ML included the recently described CRC-associated Fusobacterium nuceatum animalis clade 2, Escherichia coli strains, as well as other known CRC- biofilm related oral and gut bacteria. The platform independently discovered known CRC-associated bacteria using orders of magnitude less data and time than required for whole genome sequencing methods. The ML also discovered poorly described bacterial strains that are common in the population but negatively associated with CRC.
[0147] To test utility of the platform for CRC diagnosis, the most robust CRC-related bacterial sequence features discovered by the ML were combined into classifiers, which were used to characterize 100 blinded fecal samples, with accuracies comparable to currently available testing product results. Uniquely, as additional CRC and controls were added to the ML training dataset, CA / AA sensitivity and specificity for the blinded samples improved.
[0148] These results suggest that early detection of bacteria involved in biofilm formation may have utility as a novel CRC diagnosis and prevention method, identifying the initiation and progression of CRC-related biofilms potentially years before CRC can be detected using traditional approaches relying on human ‘omics’. Routine fecal sample monitoring combined with non-invasive methods targeting the establishment and expansion of AA- and CRC-associated biofilms could be an effective CRC prevention tool. Attorney Docket No. 24831-100810
[0149] A high -resolution, database independent, high throughput fecal microbial profiling platform was developed to generate novel insights into the relationship between bacteria and chronic disease. In one embodiment, the present invention delivers strain-level bacterial profiling, using thousands of times less sequencing than current shotgun bacterial profiling methods. In some aspects, this invention enables large-scale genomic data collection based on amplicon technology for utilization in the present invention to provide systems and methods for the detection of systemic disease from bacterial content samples obtained from a subject. This data is used as the basis for providing methods for training a machine learning algorithm to correlate patient microbiome sequence data with a disease state.
[0150] Multiple chronic human diseases have been linked to gut bacterial content, and for instance, there is evidence that oral bacterial biofilm formation in the colon contributes significantly to the development and progression of CRC. Bacterial involvement is currently being studied in breast, pancreatic, and other cancers. Individual bacteria from humans with CRC, including strains of E. coll. Fusobacterium, Bacteroides fragilis, are also associated with increased colon tumor formation in animal models, but recent work indicates that complex microbial consortia may be required for inception of tumorigenesis.
[0151] CRC therefore was selected as proof of principle for the present invention to test which data and analytical tools were both necessary and sufficient for finding associations between bacterial content and chronic disease. Methods of bacterial profiling to accurately detect AA and CRC have been explored for more than a decade, but lack of success indicated that practical technology has had not been previously developed to achieve this goal. One key advantage of the present invention is that because it is trained de novo to identify relationships between bacterial DNA sequence features and disease, the technology platform should independently re-discover bacteria that are known to be associated with AA and CRC, as well as identify novel sequence features from bacteria with undiscovered associations with CRC.
[0152] Our laboratory developed test (LDT) platform was used to profile well-characterized CRC, AA, and control fecal samples. Samples were prepared and sequenced using our platform lysis and multiple amplicon amplicon sequencing variants (ASV) bacterial strain fingerprint patterns as described previously. Sequence data were analyzed using machine learning algorithm (ML)-based feature selection for each of four ML training datasets. The training datasets started with 35 CRC, Attorney Docket No. 24831-100810
[0153] 35 AA, and 178 control samples. In subsequent ML training, more control samples were added, more CRC samples were added, and finally, FIT data was added to the training datasets, to track the rate of improvement as the ML accessed more samples, to ascertain approximately how many samples would be needed to determine what constitutes bacterial profiles corresponding to CRC, AA and normal fecal samples. In each dataset, robust features (as described in materials and methods) identified across multiple samples in the training dataset were selected and combined using ML into classifiers for identifying AA and CA-positive samples. The classifiers resulting from the ML training datasets were used to screen 100 blinded samples, results of which were not shared with us prior. FIG. 1 details the continuous increases in sensitivity and specificity for CRC screening for the 100 blinded CRC / AA / Control samples as more CA positive, control samples, and corresponding FIT data were added to the training dataset.
[0154] Database independence is a requirement for CRC screening. Uniquely, the platform of the present invention is database independent, and therefore does not require taxonomic assignment of bacterial sequences for feature identification, combination of features into classifiers, or subsequent analysis / diagnosis of new samples. Avoiding taxonomic assignments is key to the accuracy of the platform. Taxonomy reliance presents issues that can obscure high resolution data and prevent successful CRC diagnosis, because important distinctions between bacteria required to find association with CRC are frequently not assigned differential taxonomies. In addition, taxonomies can change over time, taxonomies can be incorrectly assigned, and old assignments can get ‘more incorrect’ over time. Most importantly, most bacteria are not represented in any database, with no assigned taxonomy. We show how database independence is a basic function of our platform, a result of the combination of the 16S-ITS-23S bacterial amplicon, sequencing the amplicon in a single read, identification of single read amplicon sequence variants that are 100% correct representatives of the original genome, and, because bacterial genomes typically have multiple copies of the amplicon, a unique combination of amplicons generates a bacterial ASV amplicon fingerprint unique to that strain. Database independence is a major advantage because analysis of CRC data is free from the intrinsic errors, limitations, bias, and omissions of all analysis pipelines relying on the taxonomies of sequenced bacteria.
[0155] Signals for CRC / AA biofilm detection span the entire 16S-ITS-23S amplicon. The ML identified and combined sequence features that differentiated AA and CA-positive fecal samples to build database independent classifiers for screening the 100 blinded, unknown samples. Attorney Docket No. 24831-100810
[0156] Investigating the sequence features and feature combinations in the classifiers can provide context for why the ML is able to identify CRC- and AA- related bacterial biofilms associated with lesions, and may provide insight into CRC-related disease mechanisms. To explore the ML-generated classifiers, sequence features identified by the ML were plotted by starting location in FIG. 2A. Within the 16S and 23 S regions of our amplicon there were relatively long stretches of sequence with few features, with intermittent spikes of highly discriminant features concentrated in specific regions. FIG. 2A depicts the location of four commonly used amplicon primer sites bracketing known 16S rRNA gene variable sites that are commonly used as amplicon targets for identifying bacteria, so different commonly used amplicons can be easily correlated with feature density in the Figure. FIG. 2A shows the locations of 16S rRNA gene conserved and variable regions, shown approximately to scale, as well as the ITS (Internally Transcribed Spacer) and the first -600 bases of the 23 S gene that is included in our amplicon. The ITS is highly variable in size and sequence content within and across bacterial genomes. FIG. 2B details CRC and AA discriminant feature density within the 16S gene regions. The 16S V2 region has the highest density, but the more conserved regions C5, C6, and C7 had higher density than any of the other 8 variable regions. Overall, as seen in FIG. 2C, the total feature density is 16% higher in the conserved regions than in the variable regions. Unsurprisingly, FIG. 2D shows the highly variable ITS region from -1600 to -2000 had the highest feature density as compared to the 16S and the partial 23 S gene within the amplicon.
[0157] Recent evidence demonstrates that strain-level resolution is a minimum requirement for accurate CRC screening; the ability to differentiate substrains / clades of Fusobacterium more likely to be associated with CRC is a primary example. Since the Fusobacterium nucleatum anamalis clade2 (Fna C2) / CRC clade-level association is well-studied and very specific, it was investigated whether our database-independent amplicon sequence variation was sufficient for Fna C2 identification. A detailed analysis of the 90 CRC-related screening classifiers used in the automated CRC / AA screen revealed a total of 5 features that detected all 118 Fusobacterium isolates sequenced and published in Zepeda-Rivera et al. Importantly, there was one feature, designation ‘aaca74’, that selectively detected Fna C2, demonstrating that our ML independently discovered the CRC -Fna C2 association, which is integrated into the classifiers used to screen samples. Attorney Docket No. 24831-100810
[0158] Bacterial features identified by the methods disclosed herein demonstrate that Fusobacteria clade-level association is unlikely to be the only example of highly specific, substrain-level differentiation of CRC-related bacteria. The ML was also able to identify bacterial signals that were expected, like Fusobacterium (FIG. 3A), but also other oral bacteria (FIG. 3C) and gut bacterial pathogens (FIG. 3D). One interesting result was that the ML identified bacterial sequence features, and therefore bacteria, rarely seen in fecal CRC-positive samples. FIG. 3B depicts a sequence feature mapping to ‘uncultured bacterium partial 16S rRNA gene’ sequence in the National Center for Biotechnology Information (NCBI) database. This feature serves as an example of a sequence feature negatively correlated with CRC- positive samples.
[0159] FIG. 3 shows CRC-specific sequence features contribute strong positive and negative correlations with CRC. A small subset of bacterial sequence features detected as part of the ML / Al algorithm for CRC screening sort with CRC and non-CRC (Normal). The y-axis counts the number of samples, the colors indicate the sample type. The sequence feature abundance is shown on the x-axis, using a log2 scale. Samples that did not have the signal are shown at ‘0’ Feature Abundance. The taxonomy / name summarizes the best taxonomy available for that feature after mapping the NCBI database.
[0160] Features were selected and used to build classifiers for screening samples in a databaseindependent manner, without reference to bacterial taxonomy. However, comparison of taxonomies associated with important features can provide insight into why the ML is sorting each of the blinded samples into CA, AA and normal. Investigation in silico of the taxonomic association of a subset of features demonstrated that their utility was at least partially based on the relative abundance of the feature with respect to the CRC vs. normal samples. Furthermore, the taxonomies associated with the features frequently reflected bacterial taxa that have been consistently associated with CRC in the literature. For example, the feature in FIG. 3D mapped to a bacterial ASVs including C. difficile and Peptostreptococcus, both of which have been strongly linked to CRC. For the unknown bacterial feature associated with normal controls in FIG. 3B, the follow up investigation reinforced the original finding that it is currently unknown, even though it is found in half of the training samples. Oral Bacterium 2 in FIG. 3C contains features that map to Streptococcus, different biotypes of which have strong CRC association. Other features mapped to Bacteroidies fragilis, also a strong driver of CRC. These findings are consistent with previous Attorney Docket No. 24831-100810 strain level analysis of common pathogens, which suggests that there are far more strains of familiar human-associated species than are sequenced genomes. Our ability to identify signals from completely unknown bacteria that are strongly associated with CRC, and use those features to successfully screen blinded samples is proof that the combination of strain-level resolution and database independence of our platform is a powerful and necessary feature for accurate CRC and AA screening. The finding that there are multiple features and classifiers that combine to generate signals required for accurate CRC screening is consistent with the hypothesis that multiple bacterial biofilm consortia drive colon cancer. In FIG. 3, similar to Fusobacteria, there was no single bacterial taxon or single complex feature classifier that was present only in CA, AA or normal controls.
[0161] Bacterially generated metabolites that are associated with CRC were identified. Bacterial metabolites generated by the fecal microbiome have access to the local environment, but unlike bacteria, can pass into human cells, and into the bloodstream, signaling the immune system, supporting or disrupting homeostasis, and have effects far from the gut environment. Comparison of multiple bacterial metabolites in fecal samples identified a significant drop in PE DHC in both AA and CA samples as compared to controls. As with bacterial strain association, the drop was not exactly correlated with disease, but on average the difference was significant.
[0162] Bacterial profiles in the gut are generally stable over long time periods. It is important to understand stability and reproducibility of our platform, from sample collection through lysis, PCR, sequencing and analysis. It is reported that the microbiome in a healthy adult is generally stable, a key insight that enables measurement of meaningful changes over time. Our platform was validated and certified as a laboratory developed test. As an illustration of LDT results stability over time, longitudinal profiling of an individual using the platform was used to demonstrate that the same bacterial profile can be consistently measured over long periods in different samples from a healthy individual. As a demonstration, in FIG. 4, a gut microbial profile was monitored over 14 months, where no major changes to diet, health, or lifestyle occurred. As expected, the bacterial profile was remarkably stable over time. This result indicates that our assay is able to produce consistent results over time from different samples from the same individual. Multiple samples of the same fecal material produced nearly identical results as well. The stability of the microbiome over time suggests that there is a mechanism for maintaining bacterial strains and their relative Attorney Docket No. 24831-100810 abundance that is strongly conserved that may be important for long term health. Conversely, disruptions to the stable state, as seen after antibiotic use, may have negative consequences. FIG. 4 illustrates that measurements of bacterial representation produced by our platform are stable over many months if the microbiome is stable. Conversely, this also indicates that changes to the profile are meaningful.
[0163] FIG. 4. shows a bacterial profile stable over 14 months. Fecal samples from a single adult male, age 56 at the first time point, were sampled every few months over a 14 month period. Species above 1% relative abundance are shown. There was no significant change in health status, no antibiotic use, and no major change in diet over the 14-month period.
[0164] The CRC screening preformed with the present invention is unique in that it improves with more data, which can be leveraged to increase accuracy. Trends for increased CRC screening sensitivity and specificity were confirmed with the addition of samples over four rounds of sample acquisition and ML training. In FIG. 1, the addition of cancer samples resulted in increased sensitivity for CRC (improved cancer detection, fewer false negatives), whereas increases in control samples and addition of FIT data increased specificity (fewer false positives). In FIG. 5 we use those data to project how additional samples may increase accuracy of screening to 96% or greater.
[0165] FIG. 5 shows a trendline to reach 96% or greater sensitivity and specificity. An exponential curve was used to estimate that training the ML algorithm on 50 additional CA, 65 AA, and 200 control samples can be expected to achieve 96% or better sensitivity and specificity for CRC and AA.
[0166] Results to date with a limited number of samples demonstrate that as samples were added to the training dataset, the ability to correctly screen 100 blinded samples increased rapidly. For example, the addition of 22 cancer samples resulted in a sensitivity gain of 16% for cancer detection in the blinded sample set, evidence that further addition of cancer samples should yield significant additional sensitivity gains. The addition of 78 normal controls added 20% to specificity in the blinded sample set, and the addition of FIT data added a further 16%. Interestingly, FIT results did not increase sensitivity for CRC. It appears that a low FIT score is contributing to correct differentiation of false positives, but that the bacterial content is providing high levels of CRC sensitivity. One key consideration for discussion is the AA sensitivity. There Attorney Docket No. 24831-100810 were only 35 samples available for this pilot experiment, but it was a sufficiently large sample to obtain AA accuracy similar to a currently available test for the 100 blinded samples. Based on the increase in sensitivity and specificity seen for CRC as more samples were added, it is reasonable to project similar improvements for AA as more samples are added to the training dataset in future experiments.
[0167] Findings:
[0168] This study was designed to differentiate bacterial populations associated with well- characterized CRC, AA and normal fecal samples using machine learning algorithms trained on increasing quantities of high resolution amplicon sequencing data. This amplicon includes the 16S gene, a well-studied mix of conserved and variable regions that typically enables bacterial taxonomic differentiation to the genus and sometimes the species level. The inclusion of the high variability of the ITS region in the amplicon enables differentiation of closely related strains. During training, the ML selected features with significant differences between the AA / CA and normal samples, and combined those features into classifiers. The database independent feature selection yielded features and classifiers with high discriminant power across the 16S, ITS, and 23 S regions of the amplicon, each of which was derived without the need for taxonomic identification, which could also be used to screen blinded samples without the need for taxonomic identification. The classifiers were used to screen a set of 100 blinded samples, and the resulting sensitivity and specificity for CA and AA reported by Exact Sciences was similar to the results with a currently available testing product. Features selected by the ML to differentiate bacterial groups were database independent, with significance calibrated to the differences between the CRC-related cases and controls, rather than established taxonomic conventions like species or strain names, or relatively arbitrary taxonomic conventions like genomic average nucleotide identity, or housekeeping gene differences, which can help with taxonomic assignment, but which may not be important in CRC or other disease states. For example, it has long been known that Fusobacterium nucleatum (Fn) is frequently associated with CRC, but only in about half of cases. A clade-level study recently demonstrated that Fusobacterium nucleatum anamalis clade 2 ( Fna C2) is enriched in CRC tumors, and other Fn strains are less likely to be associated with CRC. However, even this knowledge can have limited practical utility, because there are a limited number of Fna C2 representatives sequenced, which presents a challenge to any CRC screen using taxonomic (database dependent) information. For example, if a Fusobacterium nucleatum is Attorney Docket No. 24831-100810 identified in a fecal sample, whether it is part of the ‘Fn animalis clade 2’ or not requires a comparison of the sample Fn housekeeping genes to housekeeping genes previously used for Fn typing, which in turn requires high coverage of the entire Fn genome from that sample. Obtaining sufficient sequence of a single ~2M base Fn genome from a sample for housekeeping gene assessment, where Fn may be present at only a few percent of the total population, requires extensive sequencing of the sample. For example, lOx coverage of a 2M base genome present at 1% relative abundance requires ~2B bases of sequencing for that sample, followed by assembly of the genome / gene content, followed by housekeeping gene phylogenetic tree comparison and clade determination. The sequencing required for the shotgun / taxonomy-based approach is ~80x more than the -10,000 reads required for obtaining a 100-fold oversampling of the ingerprint for the strain at 1% relative abundance. The time, cost and expertise required for shotgun-based taxonomic assessment is impractical for a routine assay, and an additional complication is that it has not been shown that the ‘Clade 2’ designation is the complete and correct taxonomic division for Fn- related CRC association, as it is based on housekeeping gene comparison rather than CRC disease association. In contrast, the database-independent classifiers derived from the amplicon / ML approach independently discovered C2-specific sequences in the target amplicon that were sufficient for immediate, accurate automated screening. Rather than relying on housekeeping gene comparisons to sort bacteria by function, our database independent method allows division of bacterial function with respect to the disease of interest, generating strain-level resolution, and sometimes beyond the strain level, as required for robust CRC / AA screening.
[0169] One advantage of the ML approach is that it integrates both positive and negative associations for screening at different levels of taxonomic resolution. This database-independent ML approach demonstrates that our amplicon contains sufficient information in the absence of the full genome sequence to identify both known and unknown CRC related bacterial sequence features. In addition to Fusobacteria, the ML identified CRC-associated substrain-level sequences from E. coli, Streptococcus, B. fragilis, and other gut and oral bacteria that have been reported as increased in CRC . Importantly, the ML was able to identify sequence features with strong negative correlations with CRC, that were present in higher levels in control samples. A combination of positive and negatively associated bacterial features were combined by the ML for AA and CRC screening specificity and sensitivity similar to a commercially available product, providing an explanation of why larger numbers of samples in the training dataset leads to increases in Attorney Docket No. 24831-100810 sensitivity and specificity. Strong but non-universal association of bacterial sequence features can include rare or uncharacterized pathogens and commensals that are either positively or negatively correlated with CRC, and these associations will become more significant to ML decision making at higher numbers of samples. As an example, the feature in FIG. 3B primarily sorted with non- CRC (Normal). Approximately half of samples did not have the feature. When looking across a limited number of samples at lower resolution, it is easy to understand how a background commensal like this one that is poorly characterized (not in a database) and relatively rare could be missed in database-dependent analyses, as the sequence data will not map to any known taxonomy and might be discarded. Database independence enables disease-specific discovery of bacterial sequence features, because once the sequence data become a feature as part of the classifiers, it is possible to link sequence-associated ASVs from the Intus Bio sample dataset. Even unknown bacterial features will link to one or more ASVs consisting of -2500 base 16S-ITS-23S amplicon sequences unique to the bacterial genomes in the training dataset. The complete ASVs enable specific sequence-level identification of the strains involved, which can guide strain isolation and characterization where appropriate. For the unknown bacterium in FIG. 3B, further investigation in silico of the taxonomic hierarchy of the sequence reinforced the original finding that it is currently unknown, even though it is found in half of the training samples, and is more common in healthy people. This finding is consistent with analysis of hospital environments, which suggests that there are far more strains of familiar human-associated pathogenic species than are characterized in NCBI. The ability to identify completely uncharacterized bacterial sequences that are strongly associated with CRC is proof that the combination of strain-level resolution and database independence of our platform is a powerful and necessary feature for CRC screening and diagnosis.
[0170] The ML developed classifiers sufficient for accurate CRC screening that contained signals for multiple bacteria, consistent with recent findings indicating that multiple biofilm-specific bacterial consortia may be responsible for tumor instantiation. Individual bacterial features were enriched in CRC or controls, but no single bacterial signals were found in all samples. For example, C. difficile was identified as one of the consortia members, but similar to Fusobacteria, about half the samples lacked the feature. Once again, in the absence of specific instructions, ML considers the simultaneous presence of multiple bacterial signals when building the classifiers. Attorney Docket No. 24831-100810
[0171] A foundational element in the success of the ML to build accurate classifiers is the size and content of our amplicon. The relative lack of CRC-related signals in specific 16S rRNA gene variable regions may explain why previous studies using short read amplicon technology were unable to detect a robust signal for CRC-related bacteria. Interestingly, a previous study attempted to classify CRC using data from the 16S V4 region (Baxter, et al.). The V4 region only contains about 2% of the features differentiating CRC from normal (FIG. 2). To measure the effects of the amplicon content, the ability of the ML was tested to generate features and classifiers using the published V4 dataset from Baxter, et al., to determine if the decreased data richness from the V4 amplicon could be used by the ML to correctly classify the 100 blinded sample set. Similar to the Baxter, et al., the V4-dataset based ML classifiers also failed to correctly classify unknown CRC fecal samples.
[0172] The amplicon data was combined with machine learning algorithms to analyze a small, well characterized 57 sample CRC, 35 sample AA, and 256 control training sample set to identify CRC and AA- related sequence features, without reliance on taxonomy or database mapping. The features were combined into classifiers that were used to screen a blinded 100 sample test dataset provided by Exact Sciences. Classifier performance using bacterial sequences was equivalent to results from a currently available testing product on a separate, well-characterized, blinded 100 fecal sample set. Uniquely, sensitivity and specificity of machine learning CRC diagnosis consistently improve with additional samples, which can be used to increase accuracy by adding more samples to the training dataset. Future work is planned to leverage the combination of high data quality, database independence, and machine learning-based feature selection to improve CRC and AA precancerous lesion detection, exceeding the benchmarks of existing CRC screening devices.
[0173] Materials and Methods:
[0174] Sample Collection: Well-characterized, anonymized AA, CA and Control samples for machine learning training were provided by Exact Sciences Corporation (Redwood City, CA, USA). Summary of diagnosis and pathology in Table NNN (supplementary file A, Table 1, ‘Training Sample Dataset 1’).
[0175] Obtaining microbiome sequence data from a biological sample may be accomplished by known methods. Releasing genetic material from bacterial cells in the biological sample Attorney Docket No. 24831-100810 typically includes subjecting the sample or bacterial cells to lysis conditions. Advantageously, proportional lysis methods such as those described in US 10,774,322, issued September 15, 2020, or US 11,149,246, issued October 19, 2021, are utilized such that the genetic material released from the bacterial cells is representative of the bacterial makeup of the sample. These patents are incorporated by reference herein for their teachings on proportional lysis methods. Preferably, though not exclusively, such proportional lysis techniques are utilized to generate retrospective and prospective patient sequence data. As otherwise described herein, unique bacterial genetic sequences may be determined from rRNA sequences, including but not limited to the 16S-ITS-23S amplicon. Such sequences and others, as well as unique sequence identification and sequencing methods, are taught in, for example, US 10,894,990, issued January 19, 2021, which is incorporated by reference herein for said teachings. PCR amplification may be accomplished by methods which eliminate primer concentration-dependent PCR amplification bias, such as those taught in US 2019 / 0352712, published November 21, 2019, which is incorporated by reference herein for said teachings. While these methods provide advantages, any other appropriate methods may be implemented.
[0176] An additional 100 blinded, anonymized AA, CA and Control samples were provided by Exact Sciences Corporation (Redwood City, CA, USA). Diagnosis and pathology results were only known by Exact Sciences. Results were reported by Exact Sciences to Intus Bio as ‘percentage correct’, rather than on an individual sample basis, for all blinded sample experiments. Metadata for these samples is held in confidence by Exact Sciences, and is unknown to the authors as part of the continuing data collection / algorithm improvement.
[0177] Additional anonymized CA samples were provided by James Kinross, (supplementary file A, Table 2, ‘Additional CRC Training Samples)
[0178] Additional anonymized normal control samples were provided by Intus Biosciences (supplementary file A, Table 3, ‘Additional Control Training Samples)
[0179] Sample Processing: DNA from fecal samples was extracted and barcoded amplicons were prepared as previously described using our Intus Biosciences Complete Kit (Intus Biosciences, Farmington CT). Briefly, samples were transferred to individual wells of a 96 well plate and subjected to lysis and purification as per manufacturer’s instructions. Samples were transferred to a second provided 96-well plate containing barcoded StrainlD amplicon primer sets dried down in Attorney Docket No. 24831-100810 each of 96 wells, one barcode per well. Post PCR, the samples were purified and pooled according to kit instructions for PacBio library SMRTbell preparation (cat# 100-938-900) and sequencing (Sequel lie, PacBio). Sequencing was performed using certified Laboratory Developed Test (LDT) protocols.
[0180] ML feature selection and classifier building: Database-independent machine learning algorithms were used to generate AA / CA results for CRC screening of the 100 blinded samples test dataset.
[0181] The following terms and definitions are used herein.
[0182] As used herein, “patient” and “retrospective patient” generally mean patients having a disease state diagnosed under the appropriate medical guidelines. Sequence data from “patients” and “retrospective patients” is useful for training and validating the machine learning algorithms herein. “Patients” may be organized into cohorts having or lacking the disease state. The term “subject” may be used interchangeably with “patient.” In some embodiments, “retrospective patient” may have been diagnosed as having a disease or condition, at risk of developing a disease or condition, or not diagnosed with a disease or condition and considered healthy. In other embodiments, “retrospective patient” may not yet have been diagnosed with any of these categories. The “patient” or “retrospective patient” may be mammalian, including human and nonhuman mammals. In an embodiment, the “patient” or “retrospective patient” is a human. In another embodiment, the mammalian “patient” or “retrospective patient” is bovine, equine, canine, feline, porcine, or other mammal. In further embodiments, the “patient” or “retrospective patient” may be a non-mammalian subject, including reptile, amphibian, fish, or others. As a person of skill in the art would recognize, the patient or subject may generally be any organism harboring a bacterial microbiome.
[0183] As use herein, “prospective patients” do not have a diagnosed disease state. A biological sample from a “prospective patient” may be screened using a trained machine learning model to determine whether the prospective patient has a disease state or is at risk for developing the disease state. The term “subject” may be used interchangeably with “patient.” The “patient” or “prospective patient” may be mammalian, including human and non-human mammals. In an embodiment, the “patient” or “prospective patient” is a human. In another embodiment, the Attorney Docket No. 24831-100810 mammalian “patient” or “prospective patient” is bovine, equine, canine, feline, porcine, or other mammal. In further embodiments, the “patient” or “prospective patient” may be a non-mammalian subject, including reptile, amphibian, fish, or others. As a person of skill in the art would recognize, the patient or subject may generally be any organism harboring a bacterial microbiome.
[0184] In various embodiments, it may be determined that the patient has a disease state (i.e., the characteristic of having the disease), that the patient lack a disease state (i.e., the characteristic of lacking the disease), or that the patient is at risk for developing the disease (i.e., the characteristic of being at risk for having the disease, but not yet having sufficient indicators for diagnosis under clinical guidelines). A patient “having the disease” may have any of a pre-disease stage (i.e., being at risk for disease), an early stage of a disease, an intermediate stage of a disease, or an advanced stage of a disease, depending upon the clinical guidelines for diagnosis of said disease. It should be appreciated that the present systems and methods, in various embodiments, may differentiate between these stages to determine the current disease stage of a prospective patient.
[0185] As used herein, “microbiome sequence data” and variations of the term such as “microbiome nucleotide sequences” refer to bacterial genetic sequences collected from a biological sample of a patient which are from or representative of a bacterial microbiome of the patient. At least a portion of the “microbiome sequence data” is utilized to identify “features” as described herein. The at least a portion of the “microbiome sequence data” may include any useful portion thereof. In an embodiment, the “microbiome sequence data” includes or comprises the 16S-ITS-23S region (alternatively referred to as the 16S-ITS-23S amplicon). In other embodiments, regions neighboring the 16S-ITS-23S region (such as within about 2,000; 5,000; or 10,000 base pairs) may be included with or without the 16S-ITS-23S region.
[0186] As used herein, “features” identified from microbiome sequence data are portions of the microbiome sequence data which tend to be unique across various microbiome bacterial species and which may be quantified to indicate a difference in microbiome bacterial composition. In some embodiments, the “features” may comprise but are not limited to nucleotide sequence features, k- mers, other contiguous sequences, or amplicon sequence variants (ASVs). The “features” need not be assigned to a certain taxonomical category and may be taxonomy and database independent. Generally, a feature will have an associated property of being negatively or positively correlated Attorney Docket No. 24831-100810 with a disease state. The property of being negatively or positively correlated with a disease state may be but is not necessarily binary, and various degrees of correlation may be implemented.
[0187] As used herein the term “k-mer” refers to a contiguous subsequence of length k nucleotides derived from a longer nucleotide sequence. A k-mer may be generated from microbiome sequence data. Each distinct k-mer represents a nucleotiude sequence feature that may be counted, normalized, and analyzed for association with a disease state. In some embodiments, k is an integer ranging from about 1 to about 250.
[0188] As used herein, “reduced feature set” may be a subset of candidate microbiome sequence features selected from a larger set of sequence features through one or more steps of feature set reduction methods. Reduction steps may include but are not limited to statistical filtering, normalization, correlation-based feature subset selection, classifier-based attribute evaluation, Random Forest attribute ranking, Classifier Subset Evaluation, or other computationally efficient selection methods. In some embodiments, correlation may be quantified using a Pearson correlation coefficient. Features or nucleotide sequence features with correlation coefficients great than about 0.90, 0.92, 0.94, 0.96, 0.98, or 0.99 may be selected into “reduced feature set”. In some embodiments, the “reduced feature set” may comprise of 50, 100, 200, 400, 500, 600, 800, 900, or 1000 features or nucleotide sequence features. In certain embodiments, the “reduced feature set” may comprise from about 50 to 1000 features. The “reduced feature set” may be enriched for features or nucleotide sequence features that are correlated or associated with a disease state or having a predictive value with disease state.
[0189] As used herein, “signature set” refers to a collection of one or more highly correlated nucleotide sequence features that are grouped together to form a biomarker unit. “Signature set” may originate from reduced feature set or the seed feature identified during feature set reduction methods. “Signature set” may be expanded to include additional features that exhibit high correlation with the reduced feature set or the seed feature. In some embodiments, correlation may be quantified using a Pearson correlation coefficient. Features with correlation coefficients great than about 0.90, 0.92, 0.94, 0.96, 0.98, or 0.99 relative to a seed feature may be grouped with the seed feature to form “signature set”. In certain embodiments, the threshold may be selected based on dataset and statistical analysis and may fall within a range of about 0.90 to 1.00. In some embodiments, “signature set” may comprise from about 10 to about 20 features. The “signature Attorney Docket No. 24831-100810 set” need not be assigned to a certain taxonomical category and may be taxonomy and database independent. The “signature set” will have an associated property of being negatively or positively correlated with a disease state. The property of being negatively or positively correlated with a disease state may be but is not necessarily binary, and various degrees of correlation may be implemented.
[0190] In some embodiments, “metadata” for the various patients may be useful for training and predictive capabilities of machine learning models. Such metadata may include sex, age, prior or current medication use, diet, geographical location, race, diagnoses for diseases other than the disease state for the model, calprotectin levels, Fecal Occult Blood Test (FOBT), Fecal Immunochemical Test (FIT), Cologuard test results, or any other useful patient metadata.
[0191] As used herein, the term “biological specimen” generally encompasses any biological specimen which may contain genetic material upon which a machine learning model may be trained to correlate nucleotide sequences with a disease state. Typically, the biological specimen will contain bacterial species or genetic material from bacterial species representative of a microbiome. In some embodiments, the biological specimens are fecal samples.
[0192] In various embodiments, patient data includes “computer-readable microbiome nucleotide sequences.” It is generally contemplated that any computer-readable sequence may be utilized, as would be appreciated by a person of ordinary skill in the art. The computer-readable sequences may contain more sequence data than is necessary or desirable, in which case the methods and systems herein may truncate the sequence or extract particular regions of the sequence, as identified in the sequence data or as identified by recognition of the particular regions. Generally, any appropriate manner in which the sequence data may be accessed, loaded, and utilized is contemplated.
[0193] The term “fecal immunochemical test” refers to a test which detects the presence or absence of, or quantifies the amount of, blood present in stool from a fecal sample. In various embodiments, the “fecal immunochemical test” may encompass detection of any non-genetic material present in the stool from a fecal sample, such as bacterial markers or other biological fragments.
[0194] The systems and methods herein may correlate patient microbiome sequence data with a disease state. This correlation may include identifying the presence of the disease state, the severity Attorney Docket No. 24831-100810 of the disease state, risk for developing the disease state, or any other relevant, quantifiable aspect of the disease state of the patient. Moreover, if the patient is being treated for the disease state after an initial diagnosis, the systems and methods herein may be utilized to monitor the treatment, disease regression or progression, and / or disease remission. Additionally, if a patient having a prior disease state is now in remission of said disease state, the systems and methods may be utilized to monitor for recurrence of the disease state or recurrence of risk for re-developing the disease state.
[0195] REFERENCES
[0196] 1. Coleman, S. et al. High-resolution microbiome analysis reveals exclusionary Klebsiella species competition in preterm infants at risk for necrotizing enterocolitis. Sci. Rep. 13, 1-11 (2023).
[0197] 2. Graf, J. et al. High-Resolution Differentiation of Enteric Bacteria in Premature Infant Fecal Microbiomes Using a Novel rRNA Amplicon. MBio 12, 1-18 (2021).
[0198] 3. Hendricks, S. A. et al. High-Resolution Taxonomic Characterization Reveals Novel Human Microbial Strains with Potential as Risk Factors and Probiotics for Prediabetes and Type 2 Diabetes. Microorganisms 11, (2023).
[0199] 4. Gehrig, J. L. et al. Finding the right fit: evaluation of short-read and long-read sequencing approaches to maximize the utility of clinical microbiome data. Microb. Genomics 8, (2022).
[0200] 5. Dejea, C. M. et al. Microbiota organization is a distinct feature of proximal colorectal cancers. Proc. Natl. Acad. Sci. U. S. A. I l l, 18321-18326 (2014).
[0201] 6. El Tekle, G. & Garrett, W. S. Bacteria in cancer initiation, promotion and progression. Nat. Rev. Cancer 23, 600-618 (2023).
[0202] 7. Drewes, J. L. et al. Human Colon Cancer-Derived Clostridioides difficile Strains Drive Colonic Tumorigenesis in Mice. Cancer Discov. 12, 1873-1885 (2022).
[0203] 8. Tjalsma, H, Boleij, A., Marchesi, J. R. & Dutilh, B. E. A bacterial driver-passenger model for colorectal cancer: Beyond the usual suspects. Nat. Rev. Microbiol. 10, 575-582 (2012).
[0204] 9. Baxter, N. T., Ruffin, M. T., Rogers, M. A. M. & Schloss, P. D. Microbiota-based model improves the sensitivity of fecal immunochemical test for detecting colonic lesions. Genome Med. 8, 1-10 (2016). Attorney Docket No. 24831-100810
[0205] 10. Hanna, M., Dey, N. & Grady, W. M. Emerging Tests for Noninvasive Colorectal Cancer Screening. Clin. Gastroenterol. Hepatol. 21, 604-616 (2023).
[0206] 11. Betge, J. & Ebert, M. P. Unveiling the culprit: the fusobacterium lineage that populates colorectal cancer. Signal Transduct. Target. Ther. 9, 8-9 (2024).
[0207] 12. Zepeda-Rivera, M. et al. A distinct Fusobacterium nucleatum clade dominates the colorectal cancer niche. Nat. 2024 1-9 (2024) doi: 10.1038 / s41586-024-07182-w.
[0208] 13. Oren, A. & Garrity, G. M. Valid publication of the names of forty -two phyla of prokaryotes. Int. J. Syst. Evol. Microbiol. 71, (2021).
[0209] 14. Qiao, N. et al. After the storm — Perspectives on the taxonomy of Lactobacillaceae. JDS Commun. 3, 222-227 (2022).
[0210] 15. Ferraz Helene, L. C., Klepa, M. S. & Hungria, M. New Insights into the Taxonomy of Bacteria in the Genomic Era and a Case Study with Rhizobia. Int. J. Microbiol. 2022, (2022).
[0211] 16. Blackwell, G. A. et al. Exploring bacterial diversity via a curated and searchable snapshot of archived DNA sequences. PLoS Biol. 19, (2021).
[0212] 17. Sayers, E. W. et al. Database resources of the national center for biotechnology information. Nucleic Acids Res. 50, D20-D26 (2022).
[0213] 18. Liu, Y. et al. Peptostreptococcus anaerobius mediates anti-PDl therapy resistance and exacerbates colorectal cancer via myeloid-derived suppressor cells in mice. Nature Microbiology (Springer US, 2024). doi: 10.1038 / s41564-024-01695-w.
[0214] 19. Boleij, A., Van Gelder, M. M. H. J., Swinkels, D. W. & Tjalsma, H. Clinical importance of streptococcus gallolyticus infection among colorectal cancer patients: Systematic review and meta-analysis. Clin. Infect. Dis. 53, 870-878 (2011).
[0215] 20. Ahmed S, A., Rand R, H. & Fatimah Abu, B. The association of Streptococcus bovis / gallolyticus with colorectal tumors: The nature and the underlying mechanisms of its etiological role. J. Exp. Clin. Cancer Res. 30, 1—13 (2011).
[0216] 21. O’Brien, C. L., Allison, G. E., Grimpen, F. & Pavli, P. Impact of Colonoscopy Bowel Preparation on Intestinal Microbiota. PLoS One 8, 1-10 (2013).
[0217] 22. Zhou, X. et al. Longitudinal profding of the microbiome at four body sites reveals core stability and individualized dynamics during health and disease. Cell Host Microbe 1-21 (2024) doi: 10.1016 / j.chom.2024.02.012. Attorney Docket No. 24831-100810
[0218] 23. Aprile, F. et al. Microbiota alterations in precancerous colon lesions: A systematic review. Cancers (Basel). 13, 1-14 (2021).
[0219] 24. Morrison, A. G., Sarkar, S., Umar, S., Lee, S. T. M. & Thomas, S. M. The Contribution of the Human Oral Microbiome to Oral Disease: A Review. Microorganisms 11, 1-17 (2023).
[0220] 25. Hong, B. Y., Driscoll, M., Gratalo, D., Jarvie, T. & Weinstock, G. M. Improved DNA Extraction and Amplification Strategy for 16S rRNA Gene Amplicon-Based Microbiome Studies. Int. J. Mol. Sci. 25, (2024).
[0221] 26. White, M. T. & Sears, C. L. The microbial landscape of colorectal cancer. Nat. Rev. Microbiol. 22, 240-254 (2024).
[0222] 27. Dziubahska-kusibab, P. J. et al. Colibactin DNA-damage signature indicates mutational impact in colorectal cancer. Nat. Med. 26, (2020).
[0223] 28. Knippel, R. J., Drewes, J. L. & Sears, C. L. The Cancer Microbiome : Recent Highlights and Knowledge Gaps. Cancer Discov 2378-2395 (2021) .
[0224] 29. Diaz-Gay, M. et al. Geographic and age variations in mutational processes in colorectal cancer. Nature 0-1 (2025).
[0225] 30. Graf, J. et al. High-Resolution Differentiation of Enteric Bacteria in Premature Infant Fecal Microbiomes Using a Novel rRNA Amplicon. MBio 12, 1-18 (2021).
[0226] 31. Zepeda-Rivera, M. et al. A distinct Fusobacterium nucleatum clade dominates the colorectal cancer niche. Nature 2024 1-9 (2024).
[0227] 32. Dejea, C. M. et al. Microbiota organization is a distinct feature of proximal colorectal cancers. Proc. Natl. Acad. Sci. U. S. A. I l l, 18321-18326 (2014).
[0228] INCORPORATON BY REFERENCE
[0229] The entire disclosure of each of the patent documents, including certificates of correction, patent application documents, scientific articles, governmental reports, websites, and other references referred to herein is incorporated by reference herein in its entirety for all purposes. In case of a conflict in terminology, the present specification controls. Attorney Docket No. 24831-100810
[0230] EQUIVALENTS
[0231] The invention can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are to be considered in all respects illustrative rather than limiting on the invention described herein. In the various embodiments of the present invention, where the term comprises is used with respect to the recited components or steps of the platforms or methods, it is also contemplated that the platforms and methods consist essentially of, or consist of, the recited components or steps. Furthermore, the order of steps or order for performing certain actions is immaterial so long as the invention remains operable. Moreover, two or more steps or actions can be conducted simultaneously. In the specification, the singular forms also include the plural forms, unless the context clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In the case of conflict, the present specification will control.
[0232] All percentages and ratios used herein, unless otherwise indicated, are by weight.
Claims
Attorney Docket No. 24831-100810CLAIMSWHAT IS CLAIMED IS:
1. A method for diagnosing or predicting the development of a disease state from microbiome sequence data from a prospective patient, comprising the following steps: a. collecting biological specimens and metadata from a first plurality of patients having the disease state and from a second plurality of patients lacking the disease state; b. generating microbiome sequence data from the biological specimens; c. processing the microbiome sequence data to generate features having a quantified relevance to the disease state for each patient; d. associating metadata with the generated features for each patient from the first and second pluralities of patients; e. selecting a subset of the features to generate a reduced feature set; f. training a machine learning algorithm on the reduced feature set to create a classification model that classifies the patient status as having the disease state, lacking the disease state, or being at risk for developing the disease state; g. obtaining microbiome sequence data and metadata from a prospective patient; h. quantifying the features in the reduced feature set from the microbiome sequence data of the prospective patient; and i. applying the classification model to the quantified features in the reduced feature set from the prospective patient to determine whether the prospective patient has or lacks the disease state, or is at risk for developing the disease state.
2. The method of claim 1, wherein the reduced feature set of step e comprises further expanding the reduced feature set to incorporate sequence set features correlative with the features of the reduced feature set.Attorney Docket No. 24831-1008103. The method of claim 2, wherein the sequence set features have a Pearson’s correlation coefficient of at least about 0.97 with the features of the reduced feature set.
4. The method of claim 1, wherein the features can be matched and compared across patients from the first and second pluralities of patients.
5. The method of claim 4, further comprising applying data transformations to calibrate, normalize, or quantize the features for comparison across patients from the first and second pluralities of patients.
6. The method of claim 1, wherein the microbiome sequence data is a 16S rRNA gene and flanking upstream and downstream genomic regions, in part or in whole.
7. The method of claim 1, wherein the microbiome sequence data begins in or upstream of the 16S rRNA gene and extends past the end of the 16S rRNA gene as a contiguous amplicon sequence.
8. The method of claim 1, wherein the microbiome sequence data comprises one or more of 16S, ITS, and 23 S sequences.
9. The method of claim 8, wherein the microbiome sequence data comprises a 16S-ITS-23S amplicon.
10. The method of claim 1, wherein the quantified relevance to the disease state of each feature is defined to be the number of occurrences of the feature in the microbiome sequence data.Attorney Docket No. 24831-10081011. The method of claim 1, wherein the length of each feature is approximately 1 to 250 nucleotides.
12. The method of claim 1 in which the biological specimens are fecal samples, blood samples, CSF samples, urine samples, saliva samples, other internal or external bodily fluids, skin swabs, gum swabs, vaginal swabs, or swabs of specific internal or external anatomical features.
13. The method of claim 1, further comprising: obtaining fecal immunochemical test data from one or more of plurality of patients having the disease state and the plurality of patients lacking the disease state; and training the machine learning algorithm on the reduced feature set and the fecal immunochemical test data to create the classification model.
14. The method of claim 13, further comprising: obtaining fecal immunochemical test data from the prospective patient; and applying the classification model to the quantified features in the reduced feature set from the prospective patient and the fecal immunochemical test data from the prospective patient to determine whether the prospective patient has or lacks the disease state, or is at risk for developing the disease state.
15. The method of claim 1, in which the disease state is a neurodegenerative disease, an Alzheimer’s Disease, Parkinson’s Disease, Amyotrophic Lateral Sclerosis (ALS), Multiple Sclerosis (MS), Lewy Body Dementia, Frontotemporal Dementia, Spinocerebellar Ataxia, autoimmune disease, Celiac Disease, Crohn’s Disease, Ulcerative Colitis, Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis, Type 1 Diabetes, Hashimoto’s Thyroiditis, Graves' Disease, Psoriasis, Sjogren's Syndrome, SystemicAttorney Docket No. 24831-100810Lupus Erythematosus (SLE), Myasthenia Gravis, Vasculitis, Pemphigus Vulgaris, Dermatomyositis, Guillain-Barre Syndrome, digestive disorder, Diverticulitis, Pancreatitis, Irritable Bowel Syndrome (IBS), Gastroesophageal Reflux Disease (GERD), Peptic Ulcer Disease, Non-Alcoholic Fatty Liver Disease (NAFLD), metabolic disorders, Type 2 Diabetes, Obesity, Hyperthyroidism, Hypothyroidism, cardiovascular disease, Coronary Artery Disease, Hypertension (High Blood Pressure), Congestive Heart Failure, Stroke, Atherosclerosis, Renal (Kidney) disease, Chronic Kidney Disease (CKD), Polycystic Kidney Disease, Nephrotic Syndrome, Cancer, Lung Cancer, Breast Cancer, Prostate Cancer, Colon Cancer, Colorectal Cancer (CRC or CA), Early Onset Colorectal Cancer, Leukemia, Lymphoma, Pancreatic Cancer, Ovarian Cancer, Melanoma, Bladder Cancer, Liver Cancer, Kidney (renal cell and renal pelvis) Cancer, mental health disorder, Depression, Anxiety Disorders, Bipolar Disorder, Schizophrenia, Obsessive-Compulsive Disorder (OCD), Post-Traumatic Stress Disorder (PTSD), substance use disorder, Alcohol Use Disorder, Opioid Use Disorder, Nicotine Dependence, Chronic Obstructive Pulmonary Disease (COPD), Asthma, Fibromyalgia, Gout, Osteoarthritis, and Osteoporosis.
16. The method of claim 15, in which the disease state is Colorectal Cancer (CRC or CA).
17. A method of training a machine learning algorithm to correlate patient microbiome sequence data with a disease state: obtaining sequence data for a first plurality of patients having a diagnosed disease state and for a second plurality of control patients lacking the disease state, wherein the sequence data of the first and second pluralities of patients comprises respective computer-readable microbiome nucleotide sequences from biological samples collected from the respective patients from the first and second pluralities of patients; identifying sequence features from the microbiome nucleotide sequences which correlate positively or negatively with the disease state; generating machine learning training data comprising: i) at least a subset of the identified sequence features,Attorney Docket No. 24831-100810 ii) for each of the identified sequence features, their property of correlating positively or negatively with the disease state, and iii) retrospective patient data comprising computer-readable microbiome nucleotide sequences from biological samples collected from a plurality of retrospective patients having the disease state and / or a plurality of retrospective patients lacking the disease state; and training the machine learning algorithm with the machine learning training data to predict the presence or absence of the disease state in the retrospective patient data.
18. The method of claim 17, wherein training the machine learning algorithm produces a model capable of predicting the presence or absence of the disease state in prospective patients having no known disease state.
19. The method of claim 17, wherein the method includes no taxonomic identification of bacterial strains in the microbiome nucleotide sequences.
20. The method of claim 17, wherein the microbiome nucleotide sequences comprise bacterial nucleotide sequences.
21. The method of claim 20, wherein the bacterial nucleotide sequences comprise one or more of 16S, ITS, and 23 S sequences.
22. The method of claim 21, wherein the bacterial nucleotode sequences comprise a 16S- ITS-23S amplicon.
23. The method of claim 22, wherein the training data includes no taxonomic identification bacterial strains from the 16S-ITS-23S amplicons.
24. The method of claim 17, wherein the sequence data and retrospective patient data are proportional to the bacterial populations in the underlying biological samples.Attorney Docket No. 24831-10081025. The method of claim 17, wherein the machine learning training data further comprises retrospective patient data comprising computer-readable microbiome nucleotide sequences from biological samples collected from a plurality of retrospective patients lacking the disease state.
26. The method of claim 17, in which the disease state is a neurodegenerative disease, Alzheimer’s Disease, Parkinson’s Disease, Amyotrophic Lateral Sclerosis (ALS), Multiple Sclerosis (MS), Lewy Body Dementia, Frontotemporal Dementia, Spinocerebellar Ataxia, autoimmune disease, Celiac Disease, Crohn’s Disease, Ulcerative Colitis, Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis, Type 1 Diabetes, Hashimoto’s Thyroiditis, Graves' Disease, Psoriasis, Sjogren's Syndrome, Systemic Lupus Erythematosus (SLE), Myasthenia Gravis, Vasculitis, Pemphigus Vulgaris, Dermatomyositis, Guillain-Barre Syndrome, digestive disorder, Diverticulitis, Pancreatitis, Irritable Bowel Syndrome (IBS), Gastroesophageal Reflux Disease (GERD), Peptic Ulcer Disease, Non-Alcoholic Fatty Liver Disease (NAFLD), metabolic disorders, Type 2 Diabetes, Obesity, Hyperthyroidism, Hypothyroidism, cardiovascular disease, Coronary Artery Disease, Hypertension (High Blood Pressure), Congestive Heart Failure, Stroke, Atherosclerosis, Renal (Kidney) disease, Chronic Kidney Disease (CKD), Polycystic Kidney Disease, Nephrotic Syndrome, Cancer, Lung Cancer, Breast Cancer, Prostate Cancer, Colon Cancer, Colorectal Cancer (CRC or CA), Early Onset Colorectal Cancer, Leukemia, Lymphoma, Pancreatic Cancer, Ovarian Cancer, Melanoma, Bladder Cancer, Liver Cancer, Kidney (renal cell and renal pelvis) Cancer, mental health disorder, Depression, Anxiety Disorders, Bipolar Disorder, Schizophrenia, Obsessive-Compulsive Disorder (OCD), Post-Traumatic Stress Disorder (PTSD), substance use disorder, Alcohol Use Disorder, Opioid Use Disorder, Nicotine Dependence, Chronic Obstructive Pulmonary Disease (COPD), Asthma, Fibromyalgia, Gout, Osteoarthritis, and Osteoporosis.
27. The method of claim 26, in which the disease state is Colorectal Cancer (CRC or CA).Attorney Docket No. 24831-10081028. The method of claim 17, wherein the machine learning training data further comprises: iv) fecal immunochemistry test collected from at least one of the plurality of retrospective patients having or lacking the disease state.
29. The method of claim 17, wherein the machine learning training data further comprises metadata from at least one of the plurality of retrospective patients having or lacking the disease state.
30. A system comprising; a computing device operable to execute computer-readable instructions, the computer- readable instructions being configured to perform the steps of: obtaining sequence data for a first plurality of patients having a diagnosed disease state and for a second plurality of control patients lacking the disease state, wherein the sequence data of the first and second pluralities of patients comprises respective computer-readable microbiome nucleotide sequences from biological samples collected from the respective patients from the first and second pluralities of patients; identifying sequence features from the microbiome nucleotide sequences which correlate positively or negatively with the disease state; generating machine learning training data comprising: i) at least a subset of the identified sequence features, ii) for each of the identified sequence features, their property of corelating positively or negatively with the disease state, and iii) retrospective patient data comprising computer-readable microbiome nucleotide sequences from biological samples collected from a plurality of retrospective patients having the disease state and / or a plurality of retrospective patients lacking the disease state; and training a machine learning algorithm with the machine learning training data to predict the presence or absence of the disease state in the retrospective patient data.Attorney Docket No. 24831-10081031. A kit for diagnosing or predicting the development of a disease state from microbiome sequence data from a prospective patient, comprising: a sample collector for obtaining biological specimens from a prospective patient and instructions for obtaining the biological specimens; wherein the collected biological specimens are useful for one or more of: a. generating microbiome sequence data from the biological specimens; b. processing the microbiome sequence data to generate features having a quantified relevance to the disease state for each patient; c. associating metadata with the generated features for each patient; d. selecting a subset of the features to generate a reduced feature set; e. training a machine learning algorithm on the reduced feature set to create a classification model that classifies the patient status as having the disease state, lacking the disease state, or being at risk for developing the disease state; f. obtaining microbiome sequence data and metadata from a prospective patient; g. quantifying the features in the reduced feature set from the microbiome sequence data of the prospective patient; and h. applying the classification model to the quantified features in the reduced feature set from the prospective patient to determine whether the prospective patients has or lacks the disease state, or is at risk for developing the disease state.
32. A method for early detection of colorectal cancer, comprising: a. obtaining bacterial 16S-ITS-23S sequence data from one or more bacterial microbiomes of a patient, wherein the 16S-ITS-23S sequence data comprises at least substantially complete 16S sequences for substantially all constituent bacteria of the bacterial microbiome; b. identifying the abundances (in some cases, the presence of absence) of unique features (i.e., nucleotide sequences) within the 16S-ITS-23S sequence data that have predictive value (i.e., which may correlate positively or negatively with colorectal cancer) with respect to colorectal cancer risk or presence in the patient, wherein the features are associated with bacterial genera; andAttorney Docket No. 24831-100810 c. classifying based upon the identified features, by a classifier algorithm, whether the patient has colorectal cancer, whether the patient does not have colorectal cancer, or whether the patient is at risk for developing colorectal cancer.
Citation Information
Patent Citations
Probiotic function of human intelectin
US20170136088A1
Use of a gut microbiome as a predictor of animal growth or health
US20170220731A1
Method and system for microbiome analysis
US20170268045A1
Method and system for microbiome-derived diagnostics and therapeutics for conditions associated with microbiome taxonomic features
US20170367640A1
Sequential sequencing
US20180112264A1