Use of organoids representing pre-cancer states

Patient-derived organoids model microbe-associated colorectal signatures to predict CRC risk and identify therapeutic targets, addressing the lack of precision in current methods and elucidating the role of gut microbiota in CRC progression, enabling personalized treatment strategies.

WO2025250931A1PCT designated stage Publication Date: 2025-12-04RGT UNIV OF CALIFORNIA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/031641
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2025-05-30
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current methods lack precision in targeting colorectal cancer (CRC) risk using therapeutics and biomarkers, with NSAIDs being ineffective and poorly tolerated, and the role of gut microbiota in CRC initiation and progression remains poorly understood.

Method used

A method using patient-derived organoids (PDOs) to model microbe-associated colorectal signatures (MACS) for CRC risk prediction and therapeutic targeting, involving the use of a Boolean network to identify invariant gene signatures and validate CRC-associated microbes like Fusobacterium nucleatum, with MACS genes such as Claudin-2, Lgr-5, CEMIP, and IL-8 being upregulated upon Fn infection.

Benefits of technology

The method accurately predicts CRC risk and identifies effective therapeutic agents by modeling early events in CRC initiation, providing personalized treatment strategies for genetically predisposed individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025031641_04122025_PF_FP_ABST
    Figure US2025031641_04122025_PF_FP_ABST
Patent Text Reader

Abstract

Compositions and methods useful as research tools for identifying agents effective in treating colorectal cancer and infections are disclosed herein. Methods include providing a composition comprising a human derived organoid model culture composition comprising microbe-associated colorectal signatures (MACS). Methods further include administering a candidate therapeutic agent to the composition and determining whether a treatment effective response by the agent occurs in the composition.
Need to check novelty before this filing date? Find Prior Art

Description

USE OF ORGANOIDS REPRESENTING PRE-CANCER STATESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the priority benefit of U.S. Provisional Application No. 63 / 653,557, filed May 30, 2024. which is incorporated herein by reference in its entirety.GOVERNMENT SPONSORSHIP

[0002] This invention was made with government support under grant TR002968 aw arded by the National Institutes of Health. The government has certain rights in the invention.TECHNICAL FIELD

[0003] The present invention relates to a method for modeling cancer risk and disease using patient-derived organoids.BACKGROUND

[0004] There is no existing art of precision targeting of colorectal cancer (CRC) risk with therapeutics and its assessment with biomarkers that can be tested and validated in computationally vetted patient-derived organoids. The current practice is to monitor polyp burden and to offer NSAIDs. The latter is rarely effective and poorly tolerated.

[0005] Colorectal carcinoma (CRC) represents the third most prevalent cancer worldwide1. The global incidence of CRCs is predicted to increase by 60%, accounting for ~1.1 million deaths by 20302. According to the American Cancer Society, about 5% of people who develop CRC have inherited mutations associated with CRC-related cancer predisposition syndromes. The most common inherited syndromes linked with colorectal cancers are Lynch syndrome (hereditary non-polyposis colorectal cancer, or HNPCC)3and familial adenomatous polyposis (FAP)4 5. The mismatch repair genes MLH1 (mutL homolog 1), MSH2 (mutS homolog 2), MSH6 (mutS homolog 6), and PMS2 (postmeiotic segregation 2), as well as the gene EPCAM (epithelial cellular adhesion molecule), are associated with the Lynch Syndrome6. On the other hand, in FAP, there is loss of function of the adenomatous polyposis coli (APC) gene which is a negative regulator of WNT / 0- Catenin, a proliferative signaling pathway whose upregulation is associated with cancer development. MUTYH-associated polyposis and certain hamartomatous polyposisconditions, including Peutz-Jeghers Syndrome (PJS) and Juvenile Polyposis Syndrome (JPS), have also been associated with an increased risk of CRC7’8.

[0006] Beyond genetic CRC predispositions and other non-modifiable risk factors, approximately 70-75% of CRC cases are associated with modifiable risk factors such as lifestyle, diet, and microbial populations9. Omics-technology has been used to investigate the association of the gut microbiota with cancer initiation. Recently, Cao et al. have shown that a group of commensals together can cause CRC by generating DNA damage from the microbe-derived metabolites10. A large cohort of fecal metagenomic and metabolomic studies showed the shift in microbial population and their metabolite profile is linked to CRC11. It is also possible that a compromised gut barrier allows microbes and toxins to cross the gut epithelium and trigger inflammation, which may fuel the initiation and progression of CRCs12’15. Microbial biofilms have also been implicated in driving CRCs in both the major syndromes. Lynch and FAP, i.e., genetically engineered mouse models of these CRC syndromes show that specific bacteria could work together to induce colon inflammation and tumor formation"16,17. The transplantation of conventional microbiota increased microsatellite instability in untransformed intestinal epithelium of Msh2-Lynch mice, indicating that the microbial composition influences the rate of mutagenesis in MSH2-deficient crypts18. Pathogenic bacteria such as enterotoxigenic Bacteroides fragilis and pks+ E. coli have been linked to colitis-associated CRCs and Fusobacterium nucleatiim (Fn) has been linked to sporadic CRCs19’23. Other microbes that are found in CRC include Clostridium difficile, Enterococcus faecalis, Helicobacter pylori, and Streptococcus gallolyticus (formerly known as S. bovis-type I)19. Despite increasing evidence of these associations, how microbes fuel CRCs remains poorly understood24. Furthermore, the individual and combined contribution(s) of microbes and / or host genetics tow ards the risk of CRCs remains to be defined.SUMMARY OF THE INVENTION

[0007] In one aspect, a composition and method for modeling a biological process in a human organoid that can be used to detect microbe-associated colorectal signatures (MACS) that is a predictor of colon cancer risk. The invention provides a model and system for screening drugs against the MACS biomarker. The invention provides for use of Patient-Derived Organoids (PDOs) for testing for biomarkers of colorectal cancer (CRC) risk and screening for therapeutic targets. The invention provides that a MACSsignature can be tested by RNA seq or qPCR or MACS proteins can easily be monitored by IHC on polyp tissues to study CRC risk in patients and use that to guide colonoscopy intervals. The invention provides that PDOs can be derived from various patients at risk for CRC (IBD, hereditary syndromes, or patients with a diagnosis of CRC). The invention provides that computationally vetted PDO models can be used to screen therapeutics, toxins, CRC associated microbes and engineered microbes or other nutritional products for their ability to induce or suppress MACS genes (and therefore, risk of CRC).

[0008] In another aspect, a method of identifying a therapeutic agent effective for treating a colorectal cancer is provided. In another aspect, a human MACS organoid model and system are provided.

[0009] In aspects, a Boolean network is built using a publicly available transcriptomic dataset from healthy and adenoma affected patients to identify invariant Microbe-Associated Colorectal Cancer Signatures (MACS). The CRC-associated microbe, Fusobacterium nucleatum (Fn), has been used as a model bacterium. MACS-associated genes were validated transcriptionally by qRT-PCR, RNA seq and translationally by ELISA, IF and IHCs using tissues and colon-derived organoids from genetically predisposed mice (CPC-APCMin+ / - mice) and patients (FAP, Lynch Syndrome, PJS, and JPS).

[0010] In aspects, MACS are comprised of 4 key components: upregulation of Claudin-2 (leakiness), Lgr-5 (sternness), CEMIP (epithelial-mesenchymal transition) and IL-8 (inflammation). MACS were induced upon Fn infection, but not in response to infection with other enteric bacteria or probiotics. MACS induction upon Fn infection was higher in CPC-APCMin+ / -organoids compared to WT controls. The degree of MACS expression in the patient-derived organoids (PDOs) generally corresponded with the known lifetime risk of CRCs.

[0011] In aspects, computational prediction followed by validation in the organoid-based disease model identified the early events in CRC initiation. MACS reveals that the CRC-associated microbes induce a greater risk in the genetically predisposed hosts. MACS can be used for risk prediction or targeted for cancer prevention.

[0012] In some aspects, this application describes a method of treating colorectal cancer comprising: isolating a sample from a patient, screening the sample for microbe- associated colorectal signatures (MACS), identifying an appropriate treatment, and treating the patient with the identified appropriate treatment. In some aspects, the MACScomprises the increased expression of at least one gene. It some aspects, the MACS comprises the increased expression of a gene selected from the group consisting of Claudin-2, Lgr-5, CEMIP, and IL-8. In some aspects, the MACS comprises the increased expression of all of Claudin-2. Lgr-5, CEMIP, and IL-8. In some aspects, the sample is comprised of stem cells isolated from colonic crypts in a human patient. In some aspects, the sample is obtained from the blood of a human patient. In some aspects, the human patient has a CRC-related cancer predisposition syndrome. In some aspects, the treatment comprises the administration of an effective amount of a therapeutic compound. In some aspects, the method additionally comprising the isolation of colonic tissue from a human subject and preparation of an organoid from said tissue. In some aspects, the method additionally comprises administering a therapeutic compound to the organoid. In some aspects, the method additionally comprises monitoring the impact of the therapeutic compound on the MACS.

[0013] In another aspect, the application describes a composition comprising a patient-derived organoid, wherein the organoid exhibits one or more MACS. In some aspects, the MACS comprises increased expression of at least one gene selected from the group consisting of Claudin-2, Lgr-5, CEMIP, and IL-8. In some aspects, the organoid is derived from the colonic tissue of a patient. In some aspects, the patient has a CRC-related cancer predisposition syndrome. In some aspects, the composition additionally comprises a therapeutic compound.

[0014] In a third aspect, the application describes a method of preparation of an organoid, comprising isolation of a colonic tissue sample from a human patient, digestion of the tissue sample, and culturing the digested tissue sample in stem-cell enriched conditioned media. In some aspects, the digestion is with collagenase type I, and wherein the stem-cell enriched conditioned media contains WNT 3a, R-spondin and Noggin. In some aspects, the method additionally comprises the administration of a therapeutic compound. In some aspects, the colonic tissue sample comprises stem cells isolated from colonic crypts.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 shows study design. A database including 1567 gene expression data from both 1406 human samples and 88 mouse samples was used to build up acomputational model of colorectal carcinoma (CRC) via a Boolean Network Explorer (BoNE).

[0016] Figure 2 shows a Boolean Network map of CRC, and Establishment and Evaluation.

[0017] Figure 3 shows validation of the invariant gene signatures using human CRC tissues and organoids.

[0018] Figure 4 shows validation of the BONE predicted targets Claudin2 and CEMIP by IF and IHC.

[0019] Figure 5 shows that cancer-adjacent polyps and colon cancer-associated microbe recapitulate the gene expression changes predicted in BONE and observed during CRC initiation and progression.

[0020] Figure 6 shows the Boolean Network Explorer (BoNE): A tool for clustering and visualization of the Boolean implication network. A

[0021] Figure 7 shows that the C#l-2-3-4-5 path separates healthy vs adenoma samples.

[0022] Figure 8 shows the heatmap of genes present in the C#l-2-3-4-5 path in the test cohort.

[0023] Figure 9 shows the heatmap of the gene expression values using genes present in the C#l-2-3-4-5 path in validation cohort #1 (accession: phs001384.vl .pl). Key genes in the clusters are presented on the left of the heatmap. The expression level ranged from -1 to +1; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C#l-2-3-4-5 path is shown in the bar plot above the heatmap.

[0024] Figure 10 shows the heatmap of the gene expression values using genes present in the C#l-2-3-4-5 path in validation cohort #2 (GSE117606, GSE117607). Key genes in the clusters are presented on the left of the heatmap. The expression level ranged from -1 to +1; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C#l-2-3-4-5 path is shown in the bar plot above the heatmap.

[0025] Figure 11 shows the heatmap of the gene expression values using genes present in the C 1-2-3-4-5 path in validation cohort #3 (GSE77953). Key genes in theclusters are presented on the left of the heatmap. The expression level ranged from -1 to +1; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C#l-2-3-4-5 path is shown in the bar plot above the heatmap.

[0026] Figure 12 shows the analysis of RT-qPCR on various samples based on DCT values. Genes included in the analysis are the core MACS genes: PRKAA2, ChgA, Lgr-5, CLDN-2, IL-8 and CEMIP.

[0027] Figure 13 shows in the top the bar plot and violin plot using the C#1 -2-3-4- 5 path (left) or C#4 (right) can segregate pediatric (PED) colon samples collected from non-involved (NI) regions (n=3) from colon samples collected from polyp (P) region (n=4). Samples were harvested from genetically predisposed CRC pediatric patients. Analysis of the gene expression was done by RNA-seq. The bottom shows the heat map showed the expression of the invariant genes from non-involved and polyp areas. Core MACS genes are presented on the left of the heatmap.

[0028] Figure 14 shows the bar and violin plots using various gene signatures show separation of FAP non-involved (NI) from FAP polyp (P) samples that are obtained from patient derived organoids. Analysis of the samples was done by RNA-seq.

[0029] Figure 15 shows in the top row bar plots and violin plots showing the separation of samples from uninfected organoids and organoids infected with pks+ E. coli (GSE140929). In the middle row are shown bar plots and violin plots showing the segregation of genes of uninfected primary mouse epithelial cells and epithelial cells infected with pks+ E. coli. In the bottom row is shown the combined data of A & B. The sample ranks were based on genes in theC#l-2-3-4-5 path (left column) and C#4 genes (right column).

[0030] Figure 16 shows that the C#l-2-3-4-5 path cannot separate uninfected compared to infected using probiotics Lactobacillus (left) and enteric pathogens such as E. coli KI 2 and E. coli 0157 strains (middle), and Shigella and mutants (right).

[0031] Figure 17 shows Barrett’s esophagus organoid images.

[0032] Figure 18 shows normal esophagus organoid images.

[0033] Figure 19 shows Barrett’s gastroesophageal junction organoid images.

[0034] Figure 20 shows normal gastroesophageal junction organoid images.

[0035] Figure 21 shows Barrett’s esophagus organoids H&E images.

[0036] Figure 22 shows Barrett’s esophagus immunofluorescence images.

[0037] The following figures show evidence that pre-cancer PDOs that harbor mutations which increase the risk of colorectal cancers (CRCs) can be used as dynamic platforms for screening for drugs that can reverse the morphological defects that are associated with sternness (failure to differentiate).

[0038] Figure 23 shows first using APC min mice derived organoids (which are animal models of the human disease, Familial Adenopatous Polyposis [FAP]) to test the concept that a PRKAB1 agonist [PF] was predicted to increase differentiation and crypt budding. The figure shows that is indeed possible and there is statistically increased crypt budding.

[0039] Figure 24 shows confirmation that this sort of differentiation into crypts in vitro was associated with reduced tumor burden in the mice: Left panel shows study design and right panels show the metrics used for assessing tumor burden.

[0040] Figure 25 shows confirmation that treatment with the PRKAB1 agonist [PF] causes bCat to move outside the nucleus and go to the junctions (lower left panel).

[0041] Figure 26 shows the table of patients and their PDOs enrolled into this study to conduct human preclinical studies.

[0042] Figure 27 show the data demonstrating that PRKAB1 agonist (PF) causes increased budding in FAP PDOs that were either collected from polyps or from noninvolved (NI) colon segments.

[0043] Figure 28 shows the evidence that PRKAB1 agonist (PF) causes increased budding in Lynch Syndrome PDOs that were either collected from non-involved colon segments.

[0044] Figure 29 shows the evidence that PRKAB1 agonist (PF) causes increased budding in PJS PDOs that were either collected from polyps or from non-involved (NI) colon segments.

[0045] Figure 30 shows the evidence that PRKAB1 agonist (PF) causes increased budding in JPS PDOs that were either collected from polyps or from non-involved (NI) colon segments.DETAILED DESCRIPTION

[0046] Various further aspects and embodiments of the disclosure are provided by the following description. Before further describing various embodiments of the presently disclosed inventive concepts in more detail by way of exemplary description, examples,and results, it is to be understood that the presently disclosed inventive concepts are not limited in application to the details of methods and compositions as set forth in the following description. The presently disclosed inventive concepts are capable of other embodiments or of being practiced or carried out in various ways. As such, the language used herein is intended to be given the broadest possible scope and meaning; and the embodiments are meant to be exemplary, not exhaustive. Also, it is to be understood that the phraseology and terminology employed herein is for the purpose of description and should not be regarded as limiting unless otherwise indicated as so. Moreover, in the following detailed description, numerous specific details are set forth in order to provide a more thorough understanding of the disclosure. However, it will be apparent to a person having ordinary skill in the art that the presently disclosed inventive concepts may be practiced without these specific details. In other instances, features which are well known to persons of ordinary skill in the art have not been described in detail to avoid unnecessary' complication of the description. All of the compositions and methods of production and application and use thereof disclosed herein can be made and executed without undue experimentation in light of the present disclosure.

[0047] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0048] Unless defined otherwise, all technical and scientific terms and any acronyms used herein have the same meanings as commonly understood by one of ordinary' skill in the art in the field of the invention. Although any methods and materials similar or equivalent to those described herein can be used in the practice of the present invention, the exemplary methods, devices, and materials are described herein.

[0049] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry and immunology, w'hich are within the skill of the art. Such techniques are explained fully in the literature, such as, Molecular Cloning: A Laboratory Manual, 2nded. (Sambrook et al., 1989); Oligonucleotide Synthesis (M. J. Gait, ed., 1984); Animal Cell Culture (R. 1. Freshney. ed., 1987); Methods in Enzymology (Academic Press, Inc ); Current Protocols in Molecular Biology (F. M. Ausubel et al., eds., 1987, and periodic updates); PCR: ThePolymerase Chain Reaction (Mullis et al., eds., 1994); Remington, The Science and Practice of Pharmacy, 20thed., (Lippincott, Williams & Wilkins 2003), and Remington, The Science and Practice of Pharmacy, 22thed., (Pharmaceutical Press and Philadelphia College of Pharmacy at University of the Sciences 2012).

[0050] As used herein, the terms ‘'comprises,’’ ‘'comprising,” '‘includes,” “including,” “has,” “having,” “contains”, “containing,” “characterized by,” or any other variation thereof, are intended to encompass a non-exclusive inclusion, subject to any limitation explicitly indicated otherwise, of the recited components. For example, a cell, a pharmaceutical composition, and / or a method that “comprises” a list of elements (e.g.. components, features, or steps) is not necessarily limited to only those elements (or components or steps), but may include other elements (or components or steps) not expressly listed or inherent to the cell, pharmaceutical composition and / or method.

[0051] As used herein, the transitional phrases “consists of’ and “consisting of’ exclude any element, step, or component not specified. For example, “consists of’ or “consisting of’ used in a claim would limit the claim to the components, materials or steps specifically recited in the claim except for impurities ordinarily associated therewith (i.e., impurities within a given component). When the phrase “consists of’ or “consisting of’ appears in a clause of the body of a claim, rather than immediately following the preamble, the phrase “consists of’ or “consisting of’ limits only the elements (or components or steps) set forth in that clause; other elements (or components) are not excluded from the claim as a whole.

[0052] As used herein, the transitional phrases “consists essentially of’ and “consisting essentially of’ are used to define a fusion protein, pharmaceutical composition, and / or method that includes materials, steps, features, components, or elements, in addition to those literally disclosed, provided that these additional materials, steps, features, components, or elements do not materially affect the basic and novel characteristic(s) of the claimed invention. The term “consisting essentially of’ occupies a middle ground between “comprising” and “consisting of’.

[0053] When introducing elements of the present invention or the preferred embodiment(s) thereof, the articles “a”, “an”, “the” and “said” are intended to mean that there are one or more of the elements. The terms “comprising”, “including” and “having” are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0054] The term “and / or” when used in a list of two or more items, means that any one of the listed items can be employed by itself or in combination with any one or more of the listed items. For example, the expression “A and / or B” is intended to mean either or both of A and B, i.e. A alone, B alone or A and B in combination. The expression “A, B and / or C” is intended to mean A alone, B alone, C alone, A and B in combination, A and C in combination, B and C in combination or A, B, and C in combination.

[0055] It is understood that aspects and embodiments of the invention described herein include “consisting” and / or “consisting essentially of’ aspects and embodiments.

[0056] It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range. Values or ranges may be also be expressed herein as “about,” from “about” one particular value, and / or to “about” another particular value. When such values or ranges are expressed, other embodiments disclosed include the specific value recited, from the one particular value, and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another embodiment. It will be further understood that there are a number of values disclosed therein, and that each value is also herein disclosed as “about” that particular value in addition to the value itself. In embodiments, “about” can be used to mean, for example, within 10% of the recited value, within 5% of the recited value, or within 2% of the recited value.

[0057] The present invention provides compositions and methods, useful as research tools for example, for identifying biological processes in and agents effective for treating colorectal cancer.

[0058] The invention provides methods of modeling a biological process in a human MASC organoid.

[0059] In embodiments, the invention further provides determining a resulting biological process or a treatment effective response by a candidate therapeutic agent administered to the composition.

[0060] In embodiments, the invention provides methods of identifying a therapeutic agent effective for treating a colorectal cancer, the method comprising administering a candidate therapeutic agent to the cell culture composition and assessing whether the therapeutic agent is effective in treating or preventing the colorectal cancer.

[0061] In further embodiments, the compositions and methods are scalable, propagable. and personalized. In embodiments, the composition and method comprise or consist essentially of the components described herein.

[0062] In one aspect, a method for modeling a biological process in a human colorectal cancer is provided.

[0063] In another aspect, a method of identifying a therapeutic agent effective for treating a colorectal cancer is provided.

[0064] The term “model” or “modelling” refers to a non-naturally occurring research tool or procedure for observing effects analogous to a biological cell, organ or process such that positive correlations may be made therebetween.

[0065] As used herein, a composition containing an enriched cell population means that at least 10%, 20%, 30%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, of the cells in the composition are of the identified ty pe.

[0066] “Culture” or “cell culture” refers to the maintenance, growth and / or differentiation of cells in an in vitro environment. “Cell culture media,” “culture media” (singular “medium” in each case), “supplement” and “media supplement” refer to nutritive compositions that cultivate cell cultures.

[0067] “Cultivate,” or “maintain.” refers to the sustaining, propagating (growing) and / or differentiating of cells outside of tissue or the body, for example in a sterile plastic (or coated plastic) cell culture dish or flask. “Cultivation,” or “maintaining,” may utilize a culture medium as a source of nutrients, hormones and / or other factors helpful to propagate and / or sustain the cells.

[0068] In some embodiments, one or more of the media of the culture platform is a feeder-free environment, and optionally is substantially free of cytokines and / or growth factors. In some embodiments, the cell culture media contains supplements such as serums, extracts, grow th factors, hormones, cytokines and the like. Generally, the cultureplatform comprises one or more of stage specific feeder-free, serum-free media, each of which further comprises one or more of the following: nutrients / extracts, grow th factors, hormones, cytokines and medium additives. Suitable nutrients / extracts may include, for example, DMEM / F-12 (Dulbecco's Modified Eagle Medium / Nutrient Mixture F-12), which is a widely used basal medium for supporting the growth of many different mammalian cells; KOSR (knockout serum replacement); L-glut; NEAA (Non-Essential Amino Acids). Other medium additives may include, but not limited to, MTG, ITS, (ME, anti-oxidants (for example, ascorbic acid). In some embodiments, a culture medium of the present invention comprises one or more of the following cytokines or growth factors: epidermal growth factor (EGF), acidic fibroblast growth factor (aFGF), basic fibroblast growth factor (bFGF), leukemia inhibitory' factor (LIF), hepatocyte growth factor (HGF), insulin-like growth factor 1 (IGF-1), insulin-like growth factor 2 (IGF-2), keratinocyte growth factor (KGF). nerve growth factor (NGF). platelet-derived growth factor (PDGF), transforming growth factor beta (TGF-0), bone morphogenetic protein (BMP4), vascular endothelial cell growth factor (VEGF) transferrin, various interleukins (such as IL-1 through IL- 18), various colony-stimulating factors (such as granulocyte / macrophage colony-stimulating factor (GM-CSF)), various interferons (such as IFN-y) and other cytokines having effects upon stem cells such as stem cell factor (SCF) and erythropoietin (EPO). These cytokines may be obtained commercially, for example from R&D Systems (Minneapolis, Minn.), and may be either natural or recombinant. In some other embodiments, the culture medium of the present invention comprises one or more of bone morphogenetic protein (BMP4). insulin-like growth factor-1 (IGF-1), basic fibroblast growth factor (bFGF), vascular endothelial growth factor (VEGF), hematopoietic growth factor (for example, SCF, GMCSF, GCSF, EPO, IL3, TPO, EPO), Fms-Related Tyrosine Kinase 3 Ligand (Flt3L); and one or more cytokines from Leukemia inhibitory factor (LIF), IL3, IL6. IL7, IL11. IL15. In some embodiments, the growth factors / mitogens and cytokines are stage and / or cell type specific in concentrations that are determined empirically or as guided by the established cytokine art.

[0069] As used herein, “patient” or “subject” means a human or animal subject to be treated.

[0070] As used herein the term “pharmaceutical composition” refers to pharmaceutically acceptable compositions, wherein the composition comprises a pharmaceutically active agent, and in some embodiments further comprises apharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition may be a combination of pharmaceutically active agents and carriers.

[0071] As used herein the term “pharmaceutically acceptable'’ means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopoeia, other generally recognized pharmacopoeia in addition to other formulations that are safe for use in animals, and more particularly in humans and / or nonhuman mammals.

[0072] As used herein, “therapeutically effective amount” refers to an amount of a pharmaceutically active compound(s) that is sufficient to treat or ameliorate, or in some manner reduce the symptoms associated with diseases and medical conditions. When used with reference to a method, the method is sufficiently effective to treat or ameliorate, or in some manner reduce the symptoms associated with diseases or conditions. For example, an effective amount in reference to diseases is that amount which is sufficient to block or prevent onset; or if disease pathology has begun, to palliate, ameliorate, stabilize, reverse or slow progression of the disease, or otherwise reduce pathological consequences of the disease. In any case, an effective amount may be given in single or divided doses.

[0073] As used herein, the terms “treat,” “treatment.” or “treating” embraces at least an amelioration of the symptoms associated with diseases in the patient, where amelioration is used in a broad sense to refer to at least a reduction in the magnitude of a parameter, e.g. a symptom associated with the disease or condition being treated. As such, “treatment” also includes situations where the disease, disorder, or pathological condition, or at least symptoms associated therewith, are completely inhibited (e.g. prevented from happening) or stopped (e.g. terminated) such that the patient no longer suffers from the condition, or at least the symptoms that characterize the condition.

[0074] As used herein, and unless otherwise specified, the terms "prevent," "preventing" and "prevention" refer to the prevention of the onset, recurrence or spread of a disease or disorder, or of one or more symptoms thereof. In certain embodiments, the terms refer to the treatment with or administration of a compound or dosage form provided herein, with or without one or more other additional active agent(s), prior to the onset of symptoms, particularly to subjects at risk of disease or disorders provided herein. The terms encompass the inhibition or reduction of a symptom of the particular disease. In certain embodiments, subjects with familial history of a disease are potential candidates for preventive regimens. In certain embodiments, subjects who have a history' of recurringsymptoms are also potential candidates for prevention. In this regard, the term "prevention" may be interchangeably used with the term "prophylactic treatment."

[0075] As used herein, and unless otherwise specified, a "prophylactically effective amount" of a compound is an amount sufficient to prevent a disease or disorder, or prevent its recurrence. A prophylactically effective amount of a compound means an amount of therapeutic agent, alone or in combination with one or more other agent(s), which provides a prophylactic benefit in the prevention of the disease. The term "prophylactically effective amount" can encompass an amount that improves overall prophylaxis or enhances the prophylactic efficacy of another prophylactic agent.

[0076] Figure 1 shows study design. A database including 1567 gene expression data from both 1406 human samples and 88 mouse samples was used to build up a computational model of colorectal carcinoma (CRC) via a Boolean Network Explorer (BoNE). Boolean implication network using machine learning identifies invariant gene signatures during CRC initiation and progression. These gene signatures were tested in patient-derived diseased organoid models (Normal, FAP, HNPCC, PJS, and JPS). Moreover, these transcriptome changes were tested in murine models of CRC initiation challenged with different gut microbes. The cancer-associated microbe Fusobacterium nucleatum (Fn) could dnve cancer initiation and / or progression where the presence of host genetic factors further triggers the cancer initiation and progression.

[0077] Figure 2 show s a Boolean Network map of CRC, and Establishment and Evaluation. Panel A: Boolean network analysis was performed on a pooled transcriptomics dataset of normal colon and adenoma dataset downloaded from Gene Expression Omnibus (see Materials and Methods) to identify global pathways and biological processes that are enriched in a continuum of cellular states from normal colon to adenoma. Genes with similar expression profiles were organized into clusters, and relationships between clusters were represented as color-coded edges connecting clusters. Panel B: Graph displaying the distribution of the sizes of clusters of genes that are equivalent to each other confirms scale-free architecture of the Boolean network in A. Panel C: Top: Schematic showing an example of how Boolean cluster relationships are used to chart disease paths, beginning with the largest cluster in healthy controls and following specific Boolean invariants: A low => B low; A opposite B, and A high => B high. Bottom: Schematic showing individual gene expression changes along a Boolean path within the normal to adenoma continuum, illustrating the changing levels ofexpression of the genes in the above example. Panel D: Boolean network contains the six possible Boolean relationships between genes as color-coded edges connecting clusters. Panel E: Reactome pathway analysis of each cluster along the top continuum paths was performed to identify the signaling pathways and cellular processes that are enriched during adenoma initiation and progression. Panel F: Selection of the Boolean path using machine learning. GSE76987 data set was used to select the best path that can separate normal (n=41) and adenoma (n=41) samples. The coefficient of each path score (at the center) with 95% confidence intervals (as error bars) and the p values expressed as significant (*) were illustrated in the bar plot. The p-value for each term tests the null hypothesis that the coefficient is equal to zero (no effect). Clusters C#l-2-3-4-5, and cluster C# 1-2-4 showed the best performance. Panel G: A heatmap of the expression profile of genes within Boolean clusters superimposed on sample type (top bar) shows the accuracy of Boolean analysis in sample segregation into normal and adenomas. Panel H: Comparison of Boolean (using the C#l-2-3-4-5 path), Bayesian, and Differential analysis in segregation Normal from adenoma samples in the dataset GSE76987. Panel I: Validation of Boolean analysis (the C#l-2-3-4-5 path) in testing other data sets (Pooled GEO, GSE77953. GSE117606 / 7. and phs001384.vl.pl). Panels J-K: Detailed view of two prominent disease paths identifying downregulation of PRKAA2 (Cluster 1; AMPKa2) and CHGA (Chromogranin A) as two early events in the process; both are accompanied by an up-regulation of Lgr5, CEMIP, IL-8 and Claudin 2 (CLDN2, a leaky TJ protein). Panels L-M: Overlap of differentially expressed genes (upregulated or downregulated) between two independent datasets (test cohort and Validation cohorts). The overlap of differently expressed genes is illustrated using two Venn diagrams. The overlap between the genes in different clusters (C#l-2-3-4-5) is presented in (M) Downregulated genes that overlap in clusters C# 1-2-3 (left), and upregulated genes that overlap in C#4-5 (right).

[0078] Figure 3 shows validation of the invariant gene signatures using human CRC tissues and organoids. Panel A: Disease pathways identify the clusters and gene changes during normal to adenoma initiation. Panel B: Heatmap of the qRT-PCR gene expression predicted from BoNE using colon tissues derived from healthy human (n=3), and genetically predisposed CRC patients (FAP, lynch, PJS. and JPS) where non-involved (NI. n= 11) and polyp (P=4) tissues were assessed. See Figure 10 for more analysis. Key genes in clusters 1,3, and 4 are presented on the left of the heatmap. Data presented as DCT (Ct of target gene- Ct of the endogenous control gene). The expression level rangedfrom -1 to +1 ; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Panel C: Schematic showing the risk of CRC in different hereditary' CRC syndromes, where FAP (red) represents the highest risk of CRC development and sporadic (blue greyscale) represents the lowest risk and other categories in between them. Data collected from the published literature and described in Table 2. Panel D: Bar and violin plots from non-involved tissues (left) and polyp tissues (right). Violin plots display the rank ordering of different CRC patient samples using the average gene expression patterns of the invariant genes analyzed by RNA sequencing. ROC-AUC statistics are measured to determine the classification strength of the samples. Panel E: Schematic showing the study design involving the development of 3D organoids from hereditary polyposis samples and assessment of the invariant gene signatures using RT-qPCR and RNA seq. Panel F: Major phenotypic features observed in organoid lines derived from hereditary CRC syndromes. Top: 4 morphological phenotypic organoids were presented (disintegrated (light blue greyscale), multi-lumen organoids (dotted violet), single lumen off-center (azure), and single lumen (dark blue greyscale)). Bottom: representatives of multi-lumen organoids from different hereditary CRC syndromes. Panel G: The percentage of each organoid phenotype (as color coded in Fig. 3 Panel F) in different CRC background were presented. Panel H: The percentage of multi-lumen organoids in organoids derived from the colon of healthy human (black), colon of non-involved region from genetic hereditary' CRC patients (red), and colon of polyp region from genetic hereditary CRC patients (blue greyscale). * represents the difference between CRC and healthy, and * represents the difference between polyp and NI from the same genetic background group. Panel I: Gene expression analysis of patient-derived organoids targeting the invariant gene signatures using qRT-PCR. Top: represent bar and violin plots show the rank ordering of organoid samples (using PRKAA2, chgA, CLDN-2. CEMIP, IL-8 and Lgr-5) derived from healthy individual’s vs organoid-derived from non-involved regions vs organoid-derived from polyp regions of hereditary CRC syndromes from different locations with technical repeats Bottom: heat map displays the target genes of analysis. Panel J: Bar and violin plot showing the expression of genes in the C# 1-2-3 -4-5 path (left) or C #4 only (right) in RNA seq data from non-involved (NI from 1 FAP and 2 Lynch) and polyp (P from 3 FAP and 1 PJS) regions. Panel K: The level of IL-8 cytokine was measured in the supernatant of 3D organoids derived from healthy and different genetic CRC patients by ELISA. *Represents the difference in the level IL-8 between CRC (either NI or P) and healthy, and * represents the difference between polyp and NI from the same genetic background group. The same color represents NI and polyp derived from the same genetic background. *, **, *** means p < 0.05, 0.01, and 0.001, respectively as determined by Student’s t-test.

[0079] Figure 4 shows validation of the BONE predicted targets Claudin2 and CEMIP by IF and IHC. Panel A: Expression of CLDN2 was tested in healthy (black, grey, n=4), FAP NI (red, n=2) and FAP polyp (blue greyscale, n=2) PDOs by IF. (left) representative images for each PDOs group, (right) represent log2 fold change expression of CLDN2 intensity signals of FAP divided by the intensity of CLDN2 in healthy organoids. Panel B: Expression of CEMIP was tested in healthy (black, grey, n=4), FAP NI (red, n=2) and FAP polyp (blue greyscale, n=2) PDOs by IHC. (left) representative images for each PDOs group, (right) represents fold change expression of CEMIP. Fold change = FAP [(No. CEMIP positive organoids / the total number of counted organoids) / Healthy (No. CEMIP positive organoids / the total number of counted organoids)].

[0080] Figure 5 shows that cancer-adjacent polyps and colon cancer-associated microbe recapitulate the gene expression changes predicted in BONE and observed during CRC initiation and progression. Panel A: Schematic showing the previously published (Druliner et al., 2018, PMID:29453410) time-lapse model for CRC initiation and progression. Tubular or villous adenomas that were adjacent to (Cancer-adjacent polyps; CAPs) or free of (Cancer-free polyps; CFPs) cancer and their matched normal colon from each was processed for RNA-Seq. Panel B: Comparison of Boolean (the C#l-2-3-4-5 path), Bayesian, and Differential analysis in separating CFP from CAP samples in PMID:29453410. Panel C: Bar plot and heatmap of genes in C#4 show the separation of CAP vs CFP samples. Panel D: Left: Schematic of study design involving patients with ulcerative colitis-associated CRCs. Right: Bar plot and heatmap of genes in C#4 separates patients with proven risk for CRCs (UC with remote neoplasia; nUC) from those with low risk (quiescent UC; qUC). Panel E: Bar plots and heatmaps of genes in the C#l-2-3-4-5 path show the accurate separation of samples into Fn-infected or uninfected Caco2 monolayers (left) and Fn-infected tumors vs. adjacent normal colon tissues (right). Panel F: The gene expression of PRKAA2, chgA, CLDN-2. CEMIP, IL-8 and Lgr-5 was compared in EDMs developed from WT mouse vs EDMs developed from genetically predisposed APC min mouse, both EDMs were challenged with Fn. qRT-PCR was done for PRKAA2, chgA, CLDN-2, CEMIP, IL-8 and Lgr-5 and presented as DCT. Top: Thebar plot and violin plot show the sample ranking in WT EDMs vs. APC min EDMs that are infected or not with Fn. Bottom: Heat map showing the expression of these genes. The expression below 0 indicated low expression and was represented by blue greyscale, and the expression above 0 indicated high expression and was represented by red greyscale. Panel G: The level of CXCL1 KC (IL-8 homologue) was measured in the supernatant of the EDMs from APC min mice (uninfected and Fn infected) as used in (F) by ELISA, and the level of cytokine was compared in infected EDMs vs uninfected EDMs. *, *** means p < 0.05 and 0.001 as determined by Student’s t-test. Panel H: Summary and proposed working model for the initiation of CRCs by Fn through induction of MACS genes. Microbial dysbiosis-mediated stress-induced TJ-collapse and loss of cell polarity which is accompanied by a gene expression signature that is permissive of key cellular processes such as EMT, sternness, leakiness of the gut barrier, and a distinct type of inflammation that is IL8-predominant. This signature is distinctly associated with dysplastic progression in the epithelium (as in adenomas) and colon cancer initiation within adenomas.

[0081] Figure 6 shows the Boolean Network Explorer (BoNE): A tool for clustering and visualization of the Boolean implication network. Panel A: Overview of the computational steps used in BoNE. Panel B: BoNE was applied to analyze normal and adenoma datasets to develop a model of polarization. A pooled dataset is used to build the Boolean implication network. Panel C: BooleanNet algorithm is applied to identify Boolean implication relationships. The BoNE uses Boolean equivalent relationships to cluster genes and identify relationships between clusters. Panel D: A graphical display of cluster size analysis shows a linear trend in log-log scatterplots between clusters sorted by size and the number of clusters of any size. Panel E: Sample ordering based on single Boolean Implication relationships. Panel F: Sample ordering based on a sequence of high => high, high => low, low => low Boolean relationships. Panel G: Like panel F. four Boolean implication relationships ‘A equivalent to B’, B opposite C’, ‘C low => D low’ and ‘D low => E’ low constitute a Boolean path that can be used to develop a computational model of the progression from normal to adenoma. A suitable path is selected by using machine learning that optimizes the strength of normal / adenoma classification.

[0082] Figure 7 shows that the C#I-2-3-4-5 path separates healthy vs adenoma samples. Machine learning identified the C#l-2-3-4-5 path segregates normal vs adenoma samples in two datasets: GEO Pooled and GSE76987. The C#l-2-3-4-5 path is applied tosixteen validation datasets (13 human datasets and 3 mouse dataset) to predict normal vs adenoma samples: Human datasets are GSE77953, GSE117607, phs001384.vl.pl, GSE4183, GSE8671, GSE24713, GSE41258, GSE74843, GSE79462, GSE111156, GSE94919, GSE102573, and SRP007584. Mice datasets are GSE784, GSE422, and GSE50794. The strength of the sample separation is determined by the number of samples, ROC AUC, Accuracy, and Fisher exact p-values.

[0083] Figure 8 shows the heatmap of genes present in the C#l-2-3-4-5 path in the test cohort. Key genes in the clusters are presented on the left of the heatmap. The expression level ranged from -1 to +1; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C 1-2-3-4-5 path is shown in the bar plot above the heatmap.

[0084] Figure 9 shows the heatmap of the gene expression values using genes present in the C#l-2-3-4-5 path in validation cohort #1 (accession: phs001384.vl .pl). Key genes in the clusters are presented on the left of the heatmap. The expression level ranged from -1 to +1; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C#l-2-3-4-5 path is shown in the bar plot above the heatmap.

[0085] Figure 10 shows the heatmap of the gene expression values using genes present in the C#l-2-3-4-5 path in validation cohort #2 (GSE117606, GSE117607). Key genes in the clusters are presented on the left of the heatmap. The expression level ranged from -1 to +1 ; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C#l-2-3-4-5 path is shown in the bar plot above the heatmap.

[0086] Figure 11 shows the heatmap of the gene expression values using genes present in the C 1-2-3-4-5 path in validation cohort #3 (GSE77953). Key genes in the clusters are presented on the left of the heatmap. The expression level ranged from -1 to +1; the expression below 0 indicated low expression and represented with blue greyscale and the expression above 0 indicated high expression and represented with red greyscale. Sample ranking using the C#l-2-3-4-5 path is shown in the bar plot above the heatmap.

[0087] Figure 12 shows the analysis of RT-qPCR on various samples based on DCT values. Genes included in the analysis are the core MACS genes: PRKAA2, ChgA, Lgr-5, CLDN-2, IL-8 and CEMIP. Top-. Bar plot showing the separation of different samples based on the composite score. Middle'. Violin plot showing the composite score for each of the samples by groups. The p-value from Student’s t-test is shown if there is a significant difference (p < 0.05) between the first and selected group. Bottom-. Heatmap of core MACS gene expression. Panel A: Core MACS genes can separate healthy colon samples (n=3), colon samples from non-involved (NI) regions collected from genetically predisposed patients (n=I l) and colon samples from polyp (P) regions collected from genetically predisposed patients (n=4). Panel B: Breakdown of colon samples from NI regions into specific diseases: PJS (n=3), JPS (n=2), Lynch (n=2) and FAP (n=4). Panel C: Comparison of NI vs P regions from patients with JPS. Panel D: Comparison of NI vs P regions from patients with FAP.

[0088] Figure 13 shows in the top the bar plot and violin plot using the C#1 -2-3-4- 5 path (left) or C#4 (right) can segregate pediatric (PED) colon samples collected from non-involved (NI) regions (n=3) from colon samples collected from polyp (P) region (n=4). Samples were harvested from genetically predisposed CRC pediatric patients. Analysis of the gene expression was done by RNA-seq. The bottom shows the heat map showed the expression of the invariant genes from non-involved and polyp areas. Core MACS genes are presented on the left of the heatmap.

[0089] Figure 14 shows the bar and violin plots using various gene signatures show separation of FAP non-involved (NI) from FAP polyp (P) samples that are obtained from patient derived organoids. Analysis of the samples was done by RNA-seq.

[0090] Figure 15 shows in the top row bar plots and violin plots showing the separation of samples from uninfected organoids and organoids infected with pks+ E. coli (GSE140929). In the middle row are shown bar plots and violin plots showing the segregation of genes of uninfected primary mouse epithelial cells and epithelial cells infected with pks+ E. coli. In the bottom row is shown the combined data of A & B. The sample ranks were based on genes in theC# 1-2-3 -4-5 path (left column) and C#4 genes (right column).

[0091] Figure 16 shows that the C#l-2-3-4-5 path cannot separate uninfected compared to infected using probiotics Lactobacillus (left) and enteric pathogens such as E. coli KI 2 and E. coli 0157 strains (middle), and Shigella and mutants (right).

[0092] Figure 17 shows Barret’s esophagus organoid images.

[0093] Figure 18 shows normal esophagus organoid images.

[0094] Figure 19 shows Barret’s gastroesophageal junction organoid images.

[0095] Figure 20 shows normal gastroesophageal junction organoid images.

[0096] Figure 21 shows Barret’s esophagus organoids H&E images.

[0097] Figure 22 shows Barret’s esophagus immunofluorescence images.

[0098] The following figures show evidence that pre-cancer PDOs that harbor mutations which increase the risk of colorectal cancers (CRCs) can be used as dynamic platforms for screening for drugs that can reverse the morphological defects that are associated with sternness (failure to differentiate).

[0099] Figure 23 shows first using APC min mice derived organoids (which are animal models of the human disease, Familial Adenopatous Polyposis [FAP]) to test the concept that a PRKAB1 agonist [PF] was predicted to increase differentiation and crypt budding. The figure shows that is indeed possible and there is statistically increased crypt budding.

[0100] Figure 24 shows confirmation that this sort of differentiation into crypts in vitro was associated with reduced tumor burden in the mice: Left panel shows study design and right panels show the metrics used for assessing tumor burden.

[0101] Figure 25 shows confirmation that treatment with the PRKAB1 agonist [PF] causes bCat to move outside the nucleus and go to the junctions (lower left panel).

[0102] Figure 26 shows the table of patients and their PDOs enrolled into this study to conduct human preclinical studies.

[0103] Figure 27 show the data demonstrating that PRKAB1 agonist (PF) causes increased budding in FAP PDOs that were either collected from polyps or from noninvolved (NI) colon segments.

[0104] Figure 28 shows the evidence that PRKAB1 agonist (PF) causes increased budding in Lynch Syndrome PDOs that were either collected from non-involved colon segments.

[0105] Figure 29 shows the evidence that PRKAB1 agonist (PF) causes increased budding in PJS PDOs that were either collected from polyps or from non-involved (NI) colon segments.

[0106] Figure 30 shows the evidence that PRKAB1 agonist (PF) causes increased budding in JPS PDOs that were either collected from polyps or from non-involved (NI) colon segments.EXAMPLES

[0107] It will be understood from the foregoing description that various modifications and changes may be made in the various embodiments of the present disclosure without departing from their true spirit. The description provided herein is intended for purposes of illustration only and is not intended to be construed in a limiting sense. Thus, while the presently disclosed inventive concepts have been described herein in connection with certain embodiments so that aspects thereof may be more fully understood and appreciated, it is not intended that the presently disclosed inventive concepts be limited to these particular embodiments. On the contrary, it is intended that all alternatives, modifications and equivalents are included within the scope of the presently disclosed inventive concepts as defined herein. Thus the examples described herein, which include particular embodiments, will serve to illustrate the practice of the presently disclosed inventive concepts, it being understood that the particulars shown are by way of example and for purposes of illustrative discussion of particular embodiments of the presently disclosed inventive concepts only and are presented in the cause of providing what is believed to be a useful and readily understood description of procedures as well as of the principles and conceptual aspects of the inventive concepts. Changes may be made in the construction and formulation of the various components and compositions described herein, the methods described herein or in the steps or the sequence of steps of the methods described herein without departing from the spirit and scope of the presently disclosed inventive concepts.Computational map of colon polyps identifies molecular signatures predicting the initiation of microbe-associated cancer

[0108] Colorectal carcinoma (CRC) represents the third most prevalent cancer worldwide1. The global incidence of CRCs is predicted to increase by 60%, accounting for ~1.1 million deaths by 20302. According to the American Cancer Society, about 5% of people who develop CRC have inherited mutations associated with CRC -related cancer predisposition syndromes. The most common inherited syndromes linked with colorectalcancers are Lynch syndrome (hereditary non-polyposis colorectal cancer, or HNPCC)3and familial adenomatous polyposis (FAP)4,5. The mismatch repair genes MLH1 (mutL homolog 1), MSH2 (mutS homolog 2), MSH6 (mutS homolog 6), and PMS2 (postmeiotic segregation 2), as well as the gene EPCAM (epithelial cellular adhesion molecule), are associated with the Lynch Syndrome6. On the other hand, in FAP, there is loss of function of the adenomatous polyposis coli (APC) gene which is a negative regulator of WNT / p>- Catenin, a proliferative signaling pathway whose upregulation is associated with cancer development. MUTYH-associated polyposis and certain hamartomatous polyposis conditions, including Peutz-Jeghers Syndrome (PJS) and Juvenile Polyposis Syndrome (JPS), have also been associated with the an increased risk of CRC7,8

[0109] Beyond genetic CRC predispositions and other non-modifiable risk factors,, approximately 70-75% of CRC cases are associated with modifiable risk factors such as lifestyle, diet, and microbial populations9. Omics-technology has been used to investigate the association of the gut microbiota with cancer initiation. Recently, Cao et al. have shown that a group of commensals together can cause CRC by generating DNA damage from the microbe-derived metabolites10. A large cohort of fecal metagenomic and metabolomic studies showed the shift in microbial population and their metabolite profile is linked to CRC11. It is also possible that a compromised gut barrier allows microbes and toxins to cross the gut epithelium and trigger inflammation, which may fuel the initiation and progression of CRCs12’15. Microbial biofilms have also been implicated in driving CRCs in both the major syndromes. Lynch and FAP, i.e., genetically engineered mouse models of these CRC syndromes show that specific bacteria could work together to induce colon inflammation and tumor formation"16,17. The transplantation of conventional microbiota increased microsatellite instability in untransformed intestinal epithelium of Msh2-Lynch mice, indicating that the microbial composition influences the rate of mutagenesis in MSH2-deficient crypts18. Pathogenic bacteria such as enterotoxigenic Bacteroides fragilis and pks+ E. coli have been linked to colitis-associated CRCs and Fusobacterium nucleatum (Fn) has been linked to sporadic CRCs19’23. Other microbes that are found in CRC include Clostridium difficile, Enterococcus faecalis, Helicobacter pylori, and Streptococcus gallolyticus (formerly known as S. bovis-type I)19. Despite increasing evidence of these associations, how microbes fuel CRCs remains poorlyunderstood24. Furthermore, the individual and combined contribution(s) of microbes and / or host genetics towards the risk of CRCs remains to be defined.

[0110] To address this gap, a computational map was generated to identify the gene expression changes during normal to adenoma progression (Fig. 1). Those network- derived signatures were validated in genetically predisposed human and mouse colon samples. To assess the gene signatures transcriptionally and translationally. patient- derived organoids (PDOs) were developed from the colonic tissues of subjects with hereditary polyposis syndromes. Finally, 3D organoid models and 2D enteroid-derived monolayers co-cultured with CRC-associated microbe were used to understand the impact of microbe on host gene signatures linked to cancer initiation.Boolean implication network identifies key pathways in the progression from normal colonic tissue to adenoma

[0111] Computational algorithms have enhanced the ability to analyze “big data'’, such as transcriptomics. with the goal of understanding complex human diseases and prioritizing diagnostic, prognostic, and therapeutic markers. For example, methods using networks to map relationships between genes have been widely utilized for understanding human diseases25'30. Pair-wise gene relationships are usually identified through methods implementing correlation31'36, mutual information28, or other linear-based computational techniques37dimension reduction38and clustering39’40. However, these results may not be reproducible in the real world. A method to build a network using transcriptomics data, in which gene clusters are connected by Boolean invariant relationships41,42, has been used to track the progression of cellular states along any disease continuum. Previously, this methodology was used to identify translationally relevant cellular states with high degrees of accuracy in diverse samples and tissues43'56; most recently, it was used to identify a therapeutic target to protect the gut barrier in inflammatory bowel disease42.

[0112] Using publicly available transcriptomic datasets of normal and adenoma colonic tissues (Fig. 2 Panel A, Fig. 6, Table 1), a Boolean implication network was built using BoNE42(Boolean Network Explorer), where few large clusters were formed, while smaller-sized clusters were the most common (Fig. 2 Panel B). The Boolean implication network showed charting of the Boolean paths (Fig. 2 Panel C) where clusters are linked to each other through the six possible Boolean relationships (Fig. 2 Panel C, Fig. 2 PanelD). Reactome pathway analysis of these clusters along the path continuum revealed the most important biological processes involved in normal to adenoma progression (Fig. 2 Panel E). Each cluster was put on the normal or adenoma side depending on the average gene expression value of a cluster, and then the clusters were organized from normal (left) to adenoma (right), showing the biological processes during the initiation and progression of adenoma (Fig. 2 Panel E, Table 2), as described previously in IBD network42There is a clear initiation cluster (C#l) and an equally clear termination cluster (C#4 and C#5); multiple paths converged on the latter. C#L which was the farthest cluster in the normal colon, was enriched primarily in pathways that are initiated by AMPK (AMP-activated protein kinase) (Table 2). The termination cluster C#4 was enriched in pCatenin / TCF / LEF Wnt signaling programs, their target genes and cell-cell junction organization and C#5 include TLRs, MyD88 involved in microbial sensing (Table 2). These initiation and termination clusters also represented elements of either genetic predisposition, injury and / or immune activation (Chemokine receptors, MAPK, IL-6 and T-cell response) or oncogene activation (PI3K and KRAS). Common genetic and epigenetic aberrations known to propel adenoma progression were seen: KRAS (C#7), BRAF (C#4 and C#14), NOTCH and Eph / Ephrin signaling (C#6), cell-cycle and DNA damage (C#13, C#15). Interestingly, an IL8-specific signature in adenomas was also found (C#4) (Fig. 2 PanelE).

[0113] Table 1 : List of datasets used for analysisMACS Path: Cl-2-3-4-5Cl weight:-5 C2 weight:-0.3 C3 weight:0.1 C4 weight:2.9 C5 weight:-4PRKAA2 FERMT2 MALL RAD18 ECM1TMEM200B RAB34 LOC400960 TACSTD2 TPD52L1LAYN R0B01 ENDOD1 Clorf59 RPP25CD109 ATP8B2 CPNE8 ASCL2 TLR4CFL2 PTPRM CCDC68 LOC730101 APCDD1DCLK1 SDC2 GUCA2A REX02 AGTTCEAL7 ABI3BP VIPR1 HOXB9 ALDH1B1CLGN FYB TMCC3 PTPN13 BRIP1PDE1A WIPF1 PDXP ARHGEF1O TTYH3FLJ25076 MAGEH1 LOC646627 TRIP6 KIF2CS0X7 JAM2 ACAA1 PAFAH1B3 FUT8NAP1L3 ZNF304 CHP2 RAB15 KLK11BEX1 ADH1B SULT1A2 SLC36A4 TNS4RASSF8 IL10RA ITPKA DPP7 CCL24SDPR ASPA GCNT3 CCNB1IP1 OXGR1PEG3 ASPN HSD17B2 LOC728568 ODAMC2orfl2 SLC26A3 CDK4 KLK12RECK LRRC19 MMP7 KCNN4DLC1 ATP2B1 ZNF473 MGC11082GIMAP6 MXD1 ZNF703 GRIN2DVCAM1 PTPRR FOXQ1 CADPSSPG20 GUCA2B RCC1 LGR6CNRIP1 SECTM1 PSMG4 ETV4HEG1 MGC4172 RUVBL1 SNTB1CADM1 PTPRH OSBPL3 TCN1CD2 C7orfl0 KIAA1549 FDXRCRISPLD2 STAP2 LOC652993 FAIM2PKD2 KIF16B SORD RPESPAP1S2 IL6R TRIB3 GPC4FA 126A GDPD3 QPCTABCA8 KIAA1211 CYP4X1NLRC3 HIST1H1C GLMNARHGAP25 CCL14 GEMIN5FLU MEP1A 7A5KLRB1 SLC4A4 AZGP1C7orf58 RAPGEFL1 TRIM16MAN1C1 CDKN2B LGR5ZEB1 TUBAL3 CLCC1GNB4 FLVCR2 TBC1D16STMN2 EDN3 PHKA10LFML1 PRDX6 B9D1SAMSN1 SPPL2A ZNF259SETBP1 SLC17A4 CLDN1RASSF5 LAMA1 MTERFD3EMCN C17orf76 RIPK2RCAN2 IGSF9 METTMEM204 SLC25A20 CD44RDX RND3 WDR77PCDH7 ACSS2 MYCMY05A RHOF SLC7A5CLIP4 LOC400573 CADLDB2 Clorfl06 LOC645166ARMCX1 PIGZ SLC6A6GMFG GCNT2 SLC29A1ARHGEF6 SULT1A1 PDCD2LLIFR TP53INP2 KLHL29GPC6 CNNM4 ABCC1P2RY14 GGT6 CDC25BGALNAC4S-6ST SRI ALS2CR4CYYR1 OAF RNASEH2AIL18BP MGC13057 CEP78PTPLAD2 KCNK5 KIAA1199GIMAP7 SCIN ARNTL2HERC5 SAMD9 Clorf67FRMD6 PDLIM2 PROXIRCSD1 PCK1 XPO5CYBRD1 NHSL1 CBR3SSBP2 TMEM45B EPHB2COL14A1 ENTPD5 SP5OLFML3 TMEM120A JUBITPR1 TRPM6 OTUB2MAMDC2 TMEM37 FAM152BSLC2A5 C14orfl39 C6orfl25LBH CPM VSNL1RGL1 RP11-285G1.3 COPG2WWTR1 SCNN1B ZNRF3TMEM47 LOC100133660 NFE2L3PALM2-AKAP2 MMP28 DACH1GIMAP4 SLC36A1 TBX3C10orfl28 CA4 GALNT6DOCKIO TSPAN1 TEAD4DP YD CA2 ATP 11ASCPEP1 PLA2G10 SLC39A10AKT3 PADI2 S100A2MITF ABCG2 IFITM2CSRP2 SLC22A18AS IL8PTPRC CAI ICA1MEF2C RELL1 CLDN2GIMAP8 CMBL RNF43LOC283666 CA7 RGNEFPPP1R16B STB DI GTF2IRD1ITGA4 PRSS8 LOC100129762SGCE CLDN23 LRRC6SYNE1 LPAR1 LPCAT1KIAA1946 CNNM2 DGAT2FLRT2 ATP1B3 KRT80C14orfl32 SLC25A34 TESCMAP1B TST PPM1HTSPYL5 SDCBP2 TDGF1HCLS1 LOC100130886 SLC35E4GIMAP1 CLIC5 REPS2 PDE7B C10orf54 FXN HLA-DMB AGPAT9 LOC652900 CXCL12 SMPDL3A TRAP1 LY86 CLCN2 RAD54B TEK GRAMD3 FAM92A1 MAFB MEIS1 TGFBI CLEC2B CASP7 CDH3 MSRB3 KRT20 AXIN2 CCDC88A TSP AN 7 LOC254057 PMP22 RHOU EPHB3 RBMS1 ITM2C LOC100134295 FBN1 DDX60 FGGY AKR1B1 PEX26 CCDC113FILIP1L ALPI FAM148A AKAP12 TMEM171 KIF18A TRAC OSTbeta HIST3H2A DSE SPIB SLC38A5 LAMA4 HSD11B2 MRE11A DCN HIGD1A CYB5R2 EVI2A DHRS9 NOB1 JAZF1 CES2 ASNS FZD1 PKIB CFH AHCYL2 ZEB2 PPAP2A RASSF2 MS4A12CD48 SLC22A5 FAM129A SGK2 CLIC2 RUNDC3B EVI2B SEMA6D MAF RHOJ RFTN1 LM02 NKX2-3 ARHGAP15 APBB1IP ANK2 QKI Clorf54PGCP DOCK2 EFEMP1ARHGAP30MCCDARCSCARA5LIX1LCD27GNG2RBMS3SRPX

[0114] Table 2: Genes identified in each cluster (Excel)QPCT 20243 l_s_at MYC -14.064346 6.77E-28 -1.8081099CYP4X1 203510_at MET -17.34115 1.14E-38 -1.7694252GLMN 1438_at EPHB3 -11.90169 4.98E-20 -1.7109562GEMIN5 203798_s_at VSNL1 -12.830632 5.04E-21 -1.69768837A5 228656_at PROXI -10.571876 2.76E-17 -1.6739625AZGP1 212806_at LOC100129762 -12.336786 1.61E-21 -1.6514641TRIM16 218412_s_at GTF2IRD1 -23.557046 5.07E-45 -1.6057514LGR5 219494_at RAD54B -15.39964 8.67E-28 -1.5935512CLCC1 219911_s_at LOC100134295 -14.314115 1.71E-24 -1.591636TBC1D16 232151_at 7A5 -17.000877 4.38E-31 -1.5820053PHKA1 230875_s_at ATP11A -10.120108 8.55E-17 -1.5780662B9D1 209627_s_at 0SBPL3 -17.23808 3.58E-34 -1.5563918ZNF259 201801_s_at SLC29A1 -12.568184 1.58E-21 -1.5401055CLDN1 201853_s_at CDC25B -13.647922 1.14E-23 -1.500063MTERFD3 204268_at S100A2 -10.893168 4.77E-19 -1.468869RIPK2 225295_at SLC39A10 -10.620771 2.66E-19 -1.4500583MET 222890_at CCDC113 -13.389942 2.60E-23 -1.4426685CD44 231849_at KRT80 -13.054023 2.50E-21 -1.4416742WDR77 232370_at LOC254057 -13.007802 1.94E-21 -1.4336496MYC 41037_at TEAD4 -14.739829 6.42E-26 -1.4326031SLC7A5 218872_at TESC -11.155278 7.54E-18 -1.3825364CAD 218145_at TRIB3 -11.96436 1.41E-19 -1.3580462LOC645166 235391_at FAM92A1 -10.088051 2.67E-16 -1.3515291SLC6A6 223018_at NOB1 -18.701444 3.01E-38 -1.3425966SLC29A1 204341_at TRIM16 -11.668811 9.30E-20 -1.3194549PDCD2L 204201_s_at PTPN13 -11.652342 1.30E-18 -1.314285KLHL29 230002_at CLCC1 -13.025254 5.16E-24 -1.2905482ABCC1 222116_s_at TBC1D16 -17.678532 9.03E-36 -1.2815504CDC25B 229310_at KLHL29 -10.926444 9.49E-18 -1.2807452ALS2CR4 209589_s_at EPHB2 -12.217729 2.84E-22 -1.2784727RNASEH2A 224467_s_at PDCD2L -19.434752 7.26E-40 -1.270457CEP78 201614_s_at RUVBL1 -13.921267 3.06E-27 -1.2513633KIAA1199 227425_at REPS2 -10.645279 9.95E-19 -1.2425804ARNTL2 202804_at ABCC1 -16.991591 1.07E-32 -1.2355875Clorf67 205395_s_at MRE11A -11.101854 1.13E-19 -1.2336981PROXI 235845_at SP5 -18.670008 7.39E-35 -1.2240719XPO5 226064_s_at DGAT2 -15.631253 2.89E-28 -1.2240605CBR3 213248_at LOC730101 -12.4545 6.41E-22 -1.2025325EPHB2 221258_s_at KIF18A -9.4748778 3.55E-15 -1.2004011SP5 220658_s_at ARNTL2 -12.578639 8.95E-22 -1.1918989JUB 201420_s_at WDR77 -18.050658 3.62E-35 -1.1918168OTUB2 217988_at CCNB1IP1 -12.116074 4.64E-24 -1.1601354FAM152B 201818_at LPCAT1 -12.390396 2.36E-22 -1.1580693C6orfl25 218194_at REX02 -10.011365 6.13E-17 -1.1535549VSNL1 222760_at ZNF703 -17.049527 4.27E-32 -1.1484539C0PG2 232994_s_at RGNEF -14.963685 1.15E-27 -1.1475388ZNRF3 209129_at TRIP6 -10.7301 1.09E-17 -1.1473144NFE2L3 216417_x_at H0XB9 -11.223996 3.97E-19 -1.1380056DACH1 207153_s_at GLMN -14.720986 1.38E-29 -1.1280265TBX3 1568623_a_at SLC35E4 -13.350951 5.17E-23 -1.1265455GALNT6 206483_at LRRC6 -14.90807 3.11E-27 -1.122303TEAD4 203022_at RNASEH2A -9.9105475 4.26E-16 -1.1070781ATP11A 202246_s_at CDK4 -10.703367 3.03E-19 -1.0927069SLC39A10 59697_at RAB15 -11.886495 2.82E-21 -1.0791183S100A2 225346_at MTERFD3 -13.138633 7.05E-26 -1.0723551IFITM2 201315_x_at IFITM2 -8.6774737 1.12E-13 -1.0569559IL8 229876_at PHKA1 -12.273401 9.54E-21 -1.0449705ICA1 224200_s_at RAD18 -10.872165 6.65E-19 -1.0418898CLDN2 225841_at Clorf59 -11.640112 5.51E-20 -1.0390865RNF43 210534_s_at B9D1 -11.076798 4.20E-19 -1.036791RGNEF 205379_at CBR3 -12.105533 1.98E-21 -1.0339169GTF2IRD1 205565_s_at FXN -11.169712 1.80E-19 -1.0280467LOC100129762 228217_s_at PSMG4 -11.041245 1.82E-20 -1.0151108LRRC6 201391_at TRAP1 -12.93236 8.26E-23 -0.9998286LPCAT1 209545_s_at RIPK2 -13.139982 2.87E-24 -0.9911643DGAT2 218720_x_at LOC652900 -11.981233 2.99E-21 -0.9804084KRT80 223457_at C0PG2 -10.894497 2.24E-18 -0.9797153TESC 219718_at FGGY -12.895598 1.54E-22 -0.9794381PPM1H 223575_at KIAA1549 -13.148502 7.30E-23 -0.9782701TDGF1 205047_s_at ASNS -7.4831278 4.22E-11 -0.9720524SLC35E4 221582_at HIST3H2A -12.753363 5.87E-22 -0.957259REPS2 217998_at LOC652993 -12.043865 8.05E-21 -0.9565272FXN 206499_s_at RCC1 -11.132859 1.93E-19 -0.9457039LOC652900 1553956_at ALS2CR4 -11.012251 4.79E-19 -0.9408626TRAP1 202715_at CAD -12.472893 1.49E-22 -0.9372061RAD54B 234973_at SLC38A5 -8.8496589 8.17E-14 -0.9356674FAM92A1 203228_at PAFAH1B3 -11.351648 1.26E-20 -0.9287938TGFBI 210547_x_at ICA1 -11.677055 3.37E-21 -0.8910058CDH3 219369_s_at 0TUB2 -10.894578 8.75E-18 -0.8904202AXIN2 223057_s_at XP05 -10.329456 3.60E-17 -0.8761233LOC254057 238012_at DPP7 -10.067592 1.48E-16 -0.8754043EPHB3 225712_at GEMIN5 -10.592461 8.66E-19 -0.8596632LOC100134295 220230_s_at CYB5R2 -9.8008351 7.81E-16 -0.84947FGGY 200054_at ZNF259 -9.9673076 1.70E-16 -0.8334939CCDC113 242283_at Clorf67 -9.9891688 3.98E-16 -0.8330357FAM148A 228774 at CEP78 -9.7946624 3.40E-16 -0.7905884KIF18A 234978_at SLC36A4 -8.8812669 1.63E-14 -0.7769029HIST3H2A 228158_at LOC645166 -6.2701862 1.45E-08 -0.7688007SLC38A5 212527_at FAM152B -8.4841106 2.37E-13 -0.7588751MRE11A 226943_at LOC728568 -7.653796 7.95E-12 -0.7471652CYB5R2 224448_s_at C6orfl25 -7.7290764 2.41E-12 -0.7386832N0B1 21662O_s_at ARHGEF10 -7.4570279 5.56E-11 -0.7259855ASNS 213124 at ZNF473 -9.1772854 1.99E-14 -0.7001635Core MACS GenesWeight:-1 Weight: 1 ID Name T P Log2FCPRKAA2 CXCL8 / IL-8 238441_at PRKAA2 12.722135 4.132351e-28 1.109298CHGA LGR5 204697_s_at CHGA 15.084556 5.13E-32 2.130831CEMIP 202859_x_at IL8 -9.860601 1.961075e-17 -2.348483CLDN2 213880_at LGR5 -18.862805 3.39E-36 -3.105867212942 s at KIAA1199 -50.029244 9.94E-73 -4.489338223509_at CLDN2 -19.912716 2.750723e-31 -2.823662

[0115] Next, machine learning was used to identify the best gene clusters (nodes) connected by Boolean implication relationships (edges) in distinguishing normal from adenoma samples (Fig. 2 Panel F). Using the training dataset (normal samples: n = 41; adenoma samples: n = 41) and different possible cluster combinations, the machine learning identified clusters 1-2-3-4-5 (C#l -2-3-4-5) as the best in separating normal samples from adenoma samples with the highest accuracy (Fig. 2 Panel F, Fig. 2 PanelG). Then the C#l-2-3-4-5 path was evaluated in different cohorts, test cohorts (n=2) and validation cohorts (n=16; 13 human cohorts and 3 mouse cohorts), revealing that the C#l-2-3-4-5 path performed consistently well across all the independent cohorts (Fig. 7-11). To assess the accuracy of predicting normal vs adenoma samples using the C#l-2-3-4-5 path vs. other conventional bioinformatics tools, the Boolean analysis was compared directly with differential and Bayesian approaches that were used by others for analysis of gene expression profiles in adenoma and CRC samples57‘59. Using the training dataset, it was found that the C#l-2-3-4-5 path w as more accurate than the other two approaches and the only significant approach in segregation normal from adenoma samples (Fig. 2 PanelH). The C#l-2-3-4-5 path could classify normal from adenoma in other validation cohorts with high accuracy (ROC-AUC 0.98-1.00) (Fig. 2 Panel I). All the previous findingsshow the power of Boolean networks in accurately modeling gene expression changes that occur during normal to adenoma progression.

[0116] Using the disease map in Fig. 2 Panel E, it was found that the catalytic a- 2 subunit of AMPK (encoded by the PRKAA2 gene) was one of the first genes to be downregulated within the most prominent disease path identified (C#l), which was associated with upregulation of 4 key genes in C#4; CldvlIP a.k.a KIAA1199, which encodes the Cell migration-inducing and hyaluronan-binding protein; Cldn2, which encodes the cation-selective channels in the paracellular space; Lgr5, which encodes the sternness-reporter Leucine-rich repeat-containing G-protein coupled receptor, and interleukin-8 (IL8)-predominant inflammation (Fig. 2 Panel J). Another path with the same endpoint was the downregulation of CHGA (Fig. 2 Panel K), which was followed by the downregulation of 3 other genes that are often used as markers of differentiation in the colon: Carbonic anhydrase-1 (CAI), Membrane Spanning 4-Domains, Subfamily A, Member 12 (MS4A12), and Solute Carrier Family 26 Member 3 (SLC26A3, a.k.a down- regulated in adenoma, DRA) (Fig. 2 Panel K). The downregulation of CHGA was also associated with the upregulation of C#4 genes mentioned in Fig. 2 Panel J and Fig. 2 Panel K. Differential expression analysis revealed limited overlap between the upregulated and downregulated genes in the test pooled cohort and the three validation cohorts (Fig. 2 Panel L). The C#l-2-3-4-5 path exhibited minimal overlap in its commonly differentially expressed genes (upregulated and downregulated) (Fig. 2 Panel M).Validation of the invariant gene signatures using human CRC tissues and organoids

[0117] Looking at the major genes that are known to be mutated in hereditary CRC, it was noticed that many of these genes are not present in C 1-2-3-4-5 but are present in C#6 (APC, STK1), and C#7 (MSH2 and SMAD), which are connected to C#l-2-3-4-5 (Fig. 3 Panel A, Table 2). These predictions were validated by qRT-PCR on colon tissues collected from healthy controls and from non-involved and polyp tissues collected from the colons of genetically predisposed subjects (Table 3). It was found that the genes along the C#l-2-3-4 path were up / dowmregulated exactly as predicted: PRKAA2 and CHGA were downregulated concomitantly with the upregulation of CEMIP, CLDN2, LGR5 and IL8 in the predisposed mucosa and in polyps (Fig. 3 Panel B). Next, it was asked if theC#l-2-3-4-5 path can classify the CRC syndromes based on the established lifetime CRC risk profiles associated with these syndromes, which were gathered from publicly available information related to disease epidemiology, and it is higher in FAP > Lynch > JPS ~ PJS (Fig. 3 Panel C, Table 3). Next, RNA-seq was performed from genetically predisposed human colon samples, and it was found that the C#l-2-3-4-5 path, used as a gene signature, could separate different samples either derived from non-involved (Fig. 3 Panel D, left) or polyp region (Fig. 3 Panel D, right) that are associated with low vs high risk of CRC.

[0118] Table 3: Host genetics and the associated risk.RISK (AVG.Cluster PredominantDisease Gene Incidence AGE AT CRC number CancerDIAGNOSIS)FamilialColorectal, adenomatous 1 in 7,000 to 1 in Lifetime riskCluster 6 small bowel, polyposis 22,000 live births. 100% (39 years) gastric, etc. [FAP] AttenuatedLifetime risk FAP APC Unknown. Cluster 670% (56 years) [AFAP]MLH1, MLH1,MultipleLynch MSH2, MSH21 :370 to 1 :2,000 in Lifetime risk 60- (including syndrome MSH6, and (Cluster 7) Western populations 80% (45 years) colorectal and (HNPCC) PMS2. MSH6 endometrial)EPCAM (Cluster 8)Lifetime risk 80% for biallelic;MUTYH-1 per 10 000 and 5-7% for associatedCluster 6 Colorectal 40.000 newborns. monoallelic. polyposis(50 years; 26-98)MultipleLifetime risk (includingPeutz-Jeghers STK] 1 1 in 8300 to 1 in 280 39% (fifth Cluster 6 colorectal, small syndrome (LKB1) 000 individuals decade of life) bowel, pancreas)Lifetime risk 40-JuvenileBMPR1A, 1 in 100,000 to 50%. (third SMAD 4 Colorectal and polyposisSMAD4 160,000 individuals decade of life) (Cluster 7) gastric cancer syndromeLifetime risk MultipleCowden PTEN 1 case per 200,000 16% (sixth Cluster 6 (including syndrome (TEP1) population decade) colorectal) MultipleLi-Fraumeni p53 (TP53) Unknown. Cluster 8 (including colorectal)APC 40 per 100,000 [68 Cluster 6Sporadic colon KRAS Lifetime risk 5- (men) and 72 Singleton Colon cancer 6%NRAS (women)] Cluster 7RRAS2 Cluster 7 BRAF Singleton PIK3CA Cluster 7 P53 Cluster 8 SMAD4 Cluster 7 MLII1 Cluster 7

[0119] To determine if the colonic epithelium shows network-predicted gene expression changes, stem cells were isolated from the colonic tissues of genetically predisposed CRC patients and healthy individuals, then they were grown as 3D organoid models, and then transcriptome analysis was performed on the culturing organoids (Fig. 3 Panel E). Four distinct phenotypical features were observed during the process of their derivation: a) disintegrated (loosely packed) organoids, b) multi-lumen organoids, c) organoids with a single, but center-off lumen, and d) organoids with a centrally located lumen (Fig. 3 Panel F top row, samples listed in Table 4). The multi-lumen organoids are collections of several organoids that appeared to be connected (Fig. 3 Panel F bottom row). It was hypothesized the multi-lumen structures could be related to epithelial-mesenchymal transition (EMT). Multi-lumen structures were commonly encountered and were seen at a higher frequency when organoids were derived from polyp (P) region compared to organoids derived from non-involved (NI) region of the same patient in all the CRC categories polyp. The abundance of these multi-lumen structures correlated with the lifetime risk of CRC development, i.e., the percentage of multi-lumen organoids is in the following direction: FAP > Lynch > PJS ~ JPS (Fig. 3 Panel H). Analysis of the transcriptome of the PDOs by qRT-PCR showed the predicted gene expression profile from Boolean analysis (Fig. 3 Panel I). RNA-seq data derived from these organoids revealed that the C#l-2-3-4-5 path could segregate organoids derived from the NI region from organoids derived from the polyp region with high accuracy in CRC syndromes (Fig.3 Panel J, Figs. 12-14).

[0120] Table 4: Characteristics of patients used in this studyGenotype Patient information Age (mutation) Gender (years)H4 Healthy M 60-64H14 Healthy M 45-49H19 Healthy F 45-49FamilialAdenomatous Polyposis, OneFAP1 copy of the C 646OT8.5 (p.Arg216Ter) pathogenic variant in the APC gene FAP (familial adenomatousFAP4 polyposis) M 16 c.1370C>A APC(p.Ser457Ter)FAP (familial adenomatousFAP5 polyposis) F 13.5 c 643C>T APC (p Gln215Ter FAP (familial adenomatousFAP6 polyposis) F 9C 643OT APC(p.Gln215TerFAP (familial adenomatousFAP7 polyposis) D12.6 F 141145 G>A MYH; 494 A>G MYHLynch c 2228C>G 1 (p.S743X) in MSH2.M12Lynch C 2228OGF2 (P.S743X) in MSH2. 14 Peutz-Jeghers syndrome, Chromosome 22q13PJS1 microdeletion F 16 syndrome; SGS (short gut syndrome), Dumping syndrome Peutz-Jeghers syndrome,PJS2 Duodenal mass . . 13STK11 , c.385 dupA,Maa alteration pMetl 29fs Juvenile polyposis syndrome, the patient isJPS heterozygous for M 14 the EX4-5 del gross deletion in the BMPR1 A gene

[0121] Next, the level of C#4 genes in the protein level was assessed, especially IL-8, CEMIP. and CLDN2. Measuring the level of IL-8 cytokines in the culture of PDOs, it was found that the level of IL-8 was significantly higher in the supernatant of cultured organoids derived from polyp regions compared to the NI regions of the same mice in all CRC categories (Fig. 3 Panel K), and the level of IL-8 was matched with the lifetime CRC development (Fig. 3 Panel K, Table 3). In addition, immunofluorescence (IF) analyses showed the expression of Claudin-2 was higher in organoids derived from FAP patients compared to organoids derived from healthy individuals, and the expression was much higher in organoids derived from polyp and compared to organoids from NI regions (Fig. 4 Panel A). In a complementary approach, IHC staining was used, and it revealed that CEMIP staining was significantly higher in FAP-derived organoids than organoids derived from healthy subjects (Fig. 4 Panel B).Cancer-adjacent polyps and colon cancer-associated microbe recapitulate the gene expression changes observed during CRC initiation and progression

[0122] It was asked if the changes in gene expression along the C#l-2-3-4-5 path are associated with the risk of polyp to CRC progression. To this end, a publicly available dataset was leveraged that represents a time-lapse model for CRC initiation and progression in humans60. In that model, cancer adjacent polyps (CAPs) were used as a model to study cancer progression temporally because the precursor polyp of origin remains in direct contiguity' with its related polyps61’63. Cancer-Free Polyp (CFP) cases, on the other hand, are polyps that have remained cancer free, despite sharing similar size, histologic features and degrees of dysplasia as CAPs (Fig. 5 Panel A). Thus, in this model the laser-dissected pre-neoplastic tissues from the CAPs represent polyps with a proven high risk of CRCs, CFPs represent polyps at low' risk and paired normal colons sampled ~8 cm away from the polyps served as controls. Performance of the C#l-2-3-4-5 path (derived from the Boolean implication network) was better than other methods of analysis, such as Bayesian and differential analysis, as C#4 alone was sufficient to distinguish normal colon samples from adenomas with high accuracy (Fig. 5, Panels B and C). Furthermore, Boolean analysis using C#4 can separate quiescent UC samples (low risk of CRC) from UC with remote neoplasia (high risk of CRC) (Fig. 5 Panel D).

[0123] As the main goal of this application is to understand the impact of microbes in the progression of CRC, the publicly available dataset was searcehd with microbes involved in colorectal carcinogenesis. An analysis was conducted of the RNA seq data from organoids and polarized monolayers derived from primary murine colon epithelial cells infected with genotoxic colibactin-producing pks+ Escherichia coli strains (Fig. 15). Using the C#l-2- 3-4-5 path, EDMs infected with colon cancer microbes can be segregated from noninfected EDMs with high accuracy (Fig. 15). In contrast, the C#l-2-3-4-5 path cannot separate uninfected compared to infected using probiotics and enteric pathogens (Fig. 16).

[0124] Next, it was asked if colorectal cancer-associated microbes can increase the risk factor-associated with CRC progression. The CRC-associated microbe, Fn22was used as a model bacteria as it is the most abundant microbe enriched in colonic adenomas and CRCs compared to adjacent normal tissues22,64. Fn can trigger tumorigenesis in the mice22that resembles human polyposis and is associated with poor survival and chemoresistance among patients with CRCs65. The choice of Fn as a model carcinopathogen was further rationalized by the analysis of publicly available transcriptomic datasets with Fu-infected colonic epithelial cell lines and / ’’^-associated tumors and adjacent normal tissues66. The Boolean prediction was validated using these two publicly available transcriptomic datasets: (a) Caco2 monolayers infected or not with Fusobacterium nucleatum Fn) (Fig. 5 Panel E, left) and (b) Fra-infected tumors and adjacent normal tissues (Fig. 5 Panel E, right). In both cases, the invariant gene expression signature predicted by the Boolean analysis and that is associated with normal- to-adenoma conversion in the colon could distinguish between Fw-infected Caco2 monolayers and tumors from their respective controls (Fig. 5 Panel E). Specifically, the 6 genes (PRKAA2, CHGA. CEMIP, CLDN2, LGR5 and IL8) that are differentially expressed in the NI and polyp area also distinguish between uninfected and Fw-infected Caco2 monolayers. These 6 genes were classified as microbe-associated cancer signature (MACS). To understand the impact of CRC-associated Fn in genetically predisposed conditions, Fn was used as a model bacterium in the genetically CRC-predisposed murine model. APCMm+ / _mice strains and their WT littermates were used. APCKIin / _mice have a point mutation in the Ape gene and are a model for human familial adenomatous polyposis. The CPC- APCMin+ / " mice67were chosen, which carry a CDX2P-NLS Cre recombinase transgene and a flox-targeted Ape allele to achieve a colon-specific genedepletion. Much like human CRCs, these mice develop tumors predominantly in the distal colon and display biallelic Ape inactivation, P-catenin dysregulation, global DNA hypomethylation, and aneuploidy.

[0125] To assess the effect of genetic background on the initiation and / or progression of CRC in the presence of microbes, WT murine EDMs and CPC- APCMin+' derived) EDMs were infected with Fn and then evaluated, and the impact of infection on the previous genes was evaluated. It was found that Fn infection caused upregulation of CLDN2 and IL8 in both the normal (WT) (and the genetically predisposed (CPC- APCMin+ / '-derived) EDMs (Fig. 5 Panel F), but Fw-infection, induced CEMIP and LGR5 exclusively in the CPC- APCMin+ / ’-derived EDMs, indicating that infection may fuel EMT and sternness only in the genetically predisposed epithelium. (Fig. 5 Panel F).

[0126] Since IL-8 was part of the identified Boolean network path (C#4), the level of IL- 8 was assessed using ELISA (Fig. 5 Panel G) and it was found that Fn induces IL-8 production. Collectively, these findings confirm that recapitulating the distinct IL8- dominant nature of inflammation was identified earlier by Boolean analysis (Fig. 2). It was concluded that infection with colon cancer-associated microbes (such as Fn) triggered the distinct gene expression signature in EDMs that mirrors the changes observed along the normal-to-adenoma disease paths in the human colon. It was also concluded that the presence of genetic predisposition could fuel the initiation and / or progression of CRC, especially in the presence of CRC-associated microbe Fn (Fig. 5 Panel H).

[0127] Microbial infections contribute to 20% of all cancers68'70. Chronic infections can generate abnormal immune responses that fuel inflammation, tissue destruction and DNA damage. Infection-mediated inflammation is predicted to be a predisposing factor for CRCs, but the mechanism is not well-detailed. The available therapies for advanced CRCs are inadequate to cure most cancers. Also, post-polypectomy guidelines recommend a repeat colonoscopy in 3 to 10 years based on the size, number, and histology of polyps; however, there is no molecular basis for risk assessment for CRC progression in patients who develop these polyps. Additionally, the guidelines for the appropriate interval between surveillance colonoscopies are poorly supported by evidence71'73. This study is based on the urgent need for knowledge needed to predict the risk for CRC progression. The computational approaches described herein, with the predicted gene signature andvalidation in patient-derived organoids together with colonoscopy screening, could help detect and prevent CRC at an earlier stage. Tissues and patient-derived organoids from genetically predisposed human CRCs and animal models (CPC APC models that mimic human FAP samples) were used to identify the transcriptome changes during CRC initiation and / or progression. In addition, differentiated organoid-derived colonic cells were challenged with CRC-associated microbe- / ' / ? to assess the impact of the infection on CRC development in the genetically predisposed model (APC EDM).

[0128] Artificial intelligence (Al) has been used in the diagnosis and treatment of CRC74. Al is also used for the detection of microsatellite instability in CRC7s, accurate recognition of CRC using semi-supervised deep learning on pathological images76, development of IncRNA signature for improving outcomes in CRC77', and determination of T1 CRC metastasis risk to lymph node78. No study has utilized computational approaches to predict the microbial signatures associated with CRC risk that could be useful in preventing cancer initiation. Here, the Boolean implication network was used to identify the gene signature during adenoma initiation and progression regardless of the phenotypic (histologic) and genotypic heterogeneity within the analyzed dataset. The predictive capacity of Boolean analysis in the time-lapse model was assessed, where CAPs, CFP and their matched normal colons60were analyzed60. In that model, CAPs were used to study cancer progression temporally since the precursor polyp of origin remains in direct contiguity with its related CRC61-63. CFPs are polyps that have remained cancer-free despite the same size, histologic features, and degrees of dysplasia as CAPs. The Boolean analysis described herein successfully segregated normal colon from adenomas with 100% accuracy, indicating that the predicted gene clusters along the major path (C#l-2-3-4-5) identified previously in the NCBI GEO discovery cohort is also valid in this time-lapse cohort. This analysis showed that C#4 alone was best at segregating CAPs. CAP and colon cancer-associated microbe Fn recapitulate the gene expression changes (by C#4 with 122 genes) in cancer tissue and colon cancer cells. RNA-based panels like the 70-gene MammaPrinf in breast cancer79,80and 18-gene ‘ColoPrint’ in CRC81-83were previously reported. This application describes further validation of RNA- based 6 MACS genes from the 122 gene list (PRKAA2, CHGA, CEMIP, CLDN2, LGR5 and IL8) to predict the risk of progression to CRC.

[0129] The MACS genes showed that the catalytic a-2 subunit of AMPK (PRKAA2) was one of the first genes to be downregulated and the downregulation of PRKAA2 is predicted to have four associated changes: 1) trigger the upregulation of key genes in C#4 that are known to trigger epithelial-mesenchymal transition (EMT) during CRC initiation and progression [CEMIP; a.k.a KIAA 1 199. which encodes the Cell migration-inducing and hyaluronan-binding protein84'86]; 2) promote leakiness in an inflamed gut barrier [CLDN2, which encodes the cation-selective channels in the paracellular space, Claudin-2 which is associated with barrier dysfunction87’90]; 3) induce sternness [LGR5. which encodes the sternness-reporter Leucine-rich repeat-containing G-protein coupled receptor 591]; 4) induce interleukin-8 (IL8)-predominant inflammation (CC#4). While the expression of CLDN2 has been shown to increase in CRCs92,93and promote sternness94, the pro-inflammatory cytokine IL8 has previously been implicated in the induction of proliferation, migration, angiogenesis, sternness and metastatic potential in CRCs95‘98.

[0130] To assess the impact of microbes on the initiation and / or progression of CRC, a stem-cell-based gut-in-a-dish model was co-cultured with CRC-associated microbe Fn. Interestingly, it was found that Fn increases the risk in the genetically CRC -predisposed conditions. It was found that F -infection led to the downregulation of PRKAA2 and CHGA and upregulation of CLDN2 and IL8 in both the normal (WT) and the genetically predisposed (CPC-APC or CDX2Cre-APCMin-derived) EDMs; however, Fn infection induced CEMIP and LGR5 exclusively in the CDX2Cre - APC '"-denved EDMs, indicating that infection may fuel EMT and sternness only in the genetically predisposed epithelium. Fn also recapitulates the distinct IL8-dominant nature of inflammation identified earlier by Boolean analysis. Similarly, previous studies showed the cooccurrence of IL8 overexpression with ^-enriched tumors14, and that Fn invasion of CRCs is associated with an IL8-predominant inflammatory signature99.

[0131] This genetic predisposition to CRCs arises from the germline mutations or epimutations in the DNA mismatch repair (MMR) genes MLH1, MSH2, MSH6 and PMS2 for hereditary non-polyposis colorectal cancer (HNPCC) [Lynch syndrome], and in APC and MUTYH (recessive inheritance) for familial adenomatous polyposis (FAP) syndromes. Mutations in SMAD4, BMPR1A, STK11 and PTEN predispose to hamartomatous polyps in less common syndromes such as, the Peutz-Jeghers syndrome (PJS) and the juvenile polyposis syndrome (JPS). Patient samples from these 4 groups(FAP, Lynch, PJS and JPS) were added, and organoids were isolated to validate the expression of MACS genes. This study will be the basis of future studies to incorporate more tissue samples and PDOs with fibroblast and immune cells to recapitulate the microenvironment. A recent study showed the presence of fungi Candida albicans in 35 different cancer tissues after relocation from the gut to other organs100’101. The expression status of MACS after infection with other microbes, including viruses and fungi is another potent line of inquiry. It should be understood that the changes in genomic integrity, oncogenic signaling, cellular migration, inflammatory states, and epigenetic changes created by other microbes are important and need further mechanistic studies.

[0132] In conclusion, these findings help to understand and predict infectioninflammation-driven oncogenesis. Identifying host gene signatures will help identify markers that can predict the risk of CRC patients in an early stage. As microbes are known to fuel other cancers (pancreas, breast), the relevance of these findings goes beyond CRCs and can help identify new microbes as triggers for cancer initiation.

[0133] Table 5: Key Resources TableMETHODSData collection and annotation

[0134] Publicly available microarray and RNASeq databases were downloaded from the National Center for Biotechnology Information (NCBI) Gene Expression Omnibus (GEO) website104-106. Gene expression summarization was performed by normalizing Affymetrix platforms by RMA (Robust Multichip Average)107,108and RNAseq platforms by computing TPM (Transcripts Per Millions)109values whenever normalized data were not available in GEO. Log2(TPM) if TPM > 1 and (TPM — 1) if TPM < 1 was used as the final gene expression value for analyses. Log2(TPM + 1) was also used in some datasets. Publicly available datasets that were normalized using RPKM110, FPKM111,112, TPM113,114, and CPM115 116were also used.Adenoma datasets used for network analysis

[0135] Previously published normal colon and adenoma datasets from GEO were used to perform network analysis (Table 1). Samples of normal colon (N) and adenoma (A) were used for this analysis. Three validation datasets are used to test the gene signature: GSE77953 (13 N. 17 A), GSE117606 / GSE117607 (65 N. 204 A) and phs001384.vl.pl (20 N, 41 A). See Table 1 for all datasets analysed in this work.Computational approachesStepMiner analysis

[0136] StepMiner is a computational tool that identifies step-wise transitions in a time-series data117. StepMiner performs an adaptive regression scheme to identify the best possible step up or down based on sum-of-square errors. The steps are placed between time points at the sharpest change between low expression and high expression levels, which gives insight into the timing of the gene expression-switching event. To fit a step function, the algorithm evaluates all possible step positions, and for each position, itcomputes the average of the values on both sides of the step for the constant segments. An adaptive regression scheme is used that chooses the step positions that minimize the square error with the fitted data. Finally, a regression test statistic is computed as follows:Where .¥ffor ; - to « are the values, A?tfor i - i to « are fitted values, m is the degrees of freedom used for the adaptive regression analysis. ,¥ is the average of all the values:For a step position at k, the fitted values A?, are computed by using

[0137] Boolean logic is a simple mathematical relationship of two values, i.e., high / low, 1 / 0. or positive / negative. The Boolean analysis of gene expression data requires the conversion of expression levels into two possible values. The StepMiner algorithm is reused to perform Boolean analysis of gene expression data118. The Boolean analysis is a statistical approach which creates binary logical inferences that explain the relationships between phenomena. Boolean analysis is performed to determine the relationship between the expression levels of pairs of genes. The StepMiner algorithm is applied to gene expression levels to convert them into Boolean values (high and low). In this algorithm, first the expression values are sorted from low to high and a rising step function is fitted to the series to identify the threshold. Middle of the step is used as the StepMiner threshold. This threshold is used to convert gene expression values into Boolean values. A noise margin of 2-fold change is applied around the threshold to determine intermediate values, and these values are ignored during Boolean analysis. In a scatter plot, there are four possible quadrants based on Boolean values: (low, low), (low, high), (high, low), (high, high). A Boolean implication relationship is observed if any one of the four possible quadrants or two diagonally opposite quadrants are sparsely populated. Based on this rule, there are six kinds of Boolean implication relationships. Two of them are symmetric: equivalent (corresponding to the positively correlated genes), opposite (corresponding to the highly negatively correlated genes). Four of the Boolean relationships are asymmetric, and each corresponds to one sparse quadrant: (low => low), (high => low), (low => high), (high => high). BooleanNet statistics is used to assess the sparsity of a quadrant and thesignificance of the Boolean implication relationships43,118. Given a pair of genes A and B, four quadrants are identified by using the StepMiner thresholds on A and B by ignoring the Intermediate values defined by the noise margin of 2-fold change (+ / - 0.5 around StepMiner threshold). Number of samples in each quadrant are defined as aoo, aoi, aio, and an. which is different from X in the previous equation of F stat. Total number of samples where gene expression values for A and B are low is computed using the following equations.

[0138] Total number of samples considered is computed using the following

[0139] Expected number of samples in each quadrant is computed by assuming independence between A and B. For example, expected number of samples in the bottom left quadrant■ is computed as probability of A lowmultiplied bymultiplied by total number of samples. Following equation is used to compute the expected number of samples.

[0140] To check whether a quadrant is sparse, a statistical test for (eoo > aoo) or is performed by computing Soo and poo using following equations. A quadrant is considered sparse if Soo is highand poo is small.

[0143] A suitable threshold is chosen for Soo > sThr and poo < pThr to check sparse quadrant. A Boolean implication relationship is identified when a sparse quadrant is discovered using following equation.

[0144] Boolean Implication = (Sij > sThr, pij < pThr)A relationship is called Boolean equivalent if top-left and bottom-right quadrants are sparse.Boolean opposite relationships have sparse top-right (an) and bottom-left (aoo) quadrants,

[0145] Boolean equivalent and opposite are symmetric relationships because the relationship from A to B is the same as from B to A. Asymmetric relationship forms when there is only one quadrant sparse (A low => B low: top-left; A low => B high: bottom-left; A high=> B high: bottom-right; A high => B low: top-right). These relationships are asymmetric because the relationship from A to B is different from B to A. For example, A low => B low and B low => A low are two different relationships.A low => B high is discovered if the bottom-left (aoo) quadrant is sparse and this relationship satisfies following conditions.Similarly, A low => B low is identified if the top-left (aoi) quadrant is sparse.A high => B high Boolean implication is established if the bottom-right (aio) quadrant is sparse as described below.Boolean implication A high => B low is found if the top-right (an) quadrant is sparse using following equation.For each quadrant, a statistic Sij and an error rate pij is computed. Sij > sThr and pij < pThr are the thresholds used on the BooleanNet statistics to identify Boolean implication relationships.Boolean network explorer (BoNE)

[0146] Boolean network explorer (BoNE) provides an integrated platform for the construction, visualization and querying of a network of progressive changes underlying a disease or a biological process in three steps (Fig. 6 Panel A): First, the expression levels of all genes in these datasets were converted to binary values (high or low) using the StepMiner algorithm. Second, gene expression relationships between pairs of genes were classified into one-of-six possible Boolean Implication Relationships (BIRs), two symmetric and four asymmetric, and expressed as Boolean implication statements. This offers a distinct advantage from conventional computational methods (Bayesian, Differential, etc.) that rely exclusively on symmetric linear relationships in networks. The other advantage of using BIRs is that they are robust to the noise of sample heterogeneity (i.e., healthy, diseased, genotypic, phenotypic, ethnic, interventions, disease severity) andevery sample follows the same mathematical equation, and hence is likely to be reproducible in independent validation datasets. Third, genes with similar expression architectures, determined by sharing at least half of the equivalences among gene pairs, were grouped into clusters and organized into a network by determining the overwhelming Boolean relationships observed between any two clusters. In the resultant Boolean implication network, clusters of genes are the nodes, and the BIR between the clusters are the directed edges; BoNE enables their discover}' in an unsupervised way while remaining agnostic to the sample type.Statistical analyses

[0147] Gene signature is used to classify sample categories and the performance of the multi-class classification is measured by ROC-AUC (Receiver Operating Characteristics Area Under The Curve) values. A color-coded bar plot is combined with a density or violin+swarm plot to visualize the gene signature-based classification. All statistical tests were performed using R version 3.2.3 (2015-12-10). Standard t-tests were performed using python scipy. stats. ttesfyind package (version 0.19.0) with Welch’s Two Sample t-test (unpaired, unequal variance (equal_var=False), and unequal sample size) parameters. Multiple hypothesis corrections were performed by adjusting p values with statsmodels. stats. multitest. multipletests (fdr_bh: Benj amini / Hochberg principles). The results were independently validated with R statistical software (R version 3.6.1; 2019-07- 05). Pathway analysis of gene lists were carried out via the Reactome database and algorithm119. Reactome identifies signaling and metabolic molecules and organizes their relations into biological pathways and processes. Kaplan-Meier analysis is performed using lifelines python package version 0.14.6.Boolean implication network construction

[0148] A Boolean implication network (BIN) is created by identifying all significant pairwise Boolean implication relationships (BIRs) on a pooled dataset that included 160 normal colon and 68 adenoma samples from 11 datasets available in the human NCBI GEO Global Database for colon (Fig. 6 Panel A). The Boolean implication network contains the six possible Boolean relationships between genes in the form of a directed graph with nodes as genes and edges as the Boolean relationship between the genes. The nodes in the BIN are genes and the edges correspond to BIRs. Equivalent andOpposite relationships are denoted by undirected edges and the other four types (low => low; high => low; low => high; high => high) of BIRs are denoted by having a directed edge between them. The network of equivalences seems to follow a scale-free trend; however, other asymmetric relations in the network do not follow scale-free properties. BIR is strong and robust when the sample sizes are usually more than 200. The normal colon and adenoma datasets were pooled for Boolean analysis by filtering genes that had a reasonable dynamic range of expression values. When the dynamic range of expression values was small, it was difficult to distinguish if the values were all low or all high or if there were some high and some low values. Thus, it was determined that it would be best to ignore them during the Boolean analysis. The filtering step was performed by analyzing the fraction of high and low values identified by the StepMiner algorithm117. Any probe set or genes that contained less than 5% of high or low values were dropped from the analysis.Clusters Boolean implication network

[0149] Clustering was performed in the Boolean implication network to dramatically reduce the complexity of the network (Fig. 6 Panel C). A clustered Boolean implication network (CBIN) was created by clustering nodes in the original BIN by following the equivalent BIRs. One approach is to build connected components in an undirected graph of Boolean equivalences. However, because of noise the connected components become internally inconsistent e.g. two genes opposite to each other becomes part of the same connected component. To avoid such a situation, one needs to break the component by removing the weak links. To identify the weakest links, first was computed a minimum spanning tree for the graph and computed Jaccard similarity coefficient for every edge in this tree. Ideally, if two members are part of the same cluster, they should share as many connections as possible. If they share less than half of their total individual connections (Jaccard similarity coefficient less than 0.5) the edges are dropped from further analysis. Thus, many weak equivalences were dropped using the above algorithm, leaving the clusters internally consistent. All edges that have Jaccard similarity coefficient of less than 0.5 were removed, and the connected components were built with the rest. The connected components were used to cluster the BIN which is converted to the nodes of the CBIN. Increasing the Jaccard similarity cut-off will result in more compact and correlated clusters in CBIN. The distribution of cluster sizes was plotted in a log-log scale to observethe characteristics of the Boolean network (Fig. 6 Panel D). To ensure that the cluster sizes exhibit scale-free properties, the Jaccard similarity cut-off is modified such that they are evenly distributed along a straight line on a log-log plot (Fig. 6 Panel D). A new graph was built that connected the individual clusters to each other using Boolean relationships. Genes in each cluster is ranked based on the number of equivalences within the cluster. Link between two clusters (A, B) was established by using the top representative node from A that was connected to most of the members of A and, sampling 6 nodes from cluster B and identifying the overwhelming majority of BIRs (Fig. 6 Panel C) between the nodes from each cluster. The 6 nodes include the top representative gene (first rank), the gene next to top (second rank), middle (floor(n / 2)thrank where n is the cluster size), gene next to middle (floor(n / 2) - 1 rank), middle from top half (floor(n / 4)thranked gene), and middle from the top l / 4th(floor(n / 8)thranked gene) representative nodes from cluster B if size of the cluster is greater than 10. If the size of the cluster is between 2 to 10, top two and middle one are picked to test the relationship with cluster A. If the size of the cluster is 1, then it is used to test the relationship with cluster A. Testing multiple nodes provides the most common type of relationships found between clusters A and B. Refer to the codebase released for additional details.

[0150] A CBIN was created using the pooled normal colon and adenoma dataset. Each cluster was associated with normal colon or adenoma samples based on where these gene clusters were highly expressed - either on the healthy or disease side. The edges between the clusters represented the Boolean relationships that are color-coded in greyscale as follows: orange greyscale for low => high, dark blue greyscale for low => low, green greyscale for high => high, red greyscale for high => low, light blue greyscale for equivalent and black for opposite.Boolean paths

[0151] The asymmetric BIRs provide a unique dimension to the network that is fundamentally different from any other gene expression network in the literature. Traversing a set of nodes in a directed graph of the Boolean network constitutes a Boolean path that can be interpreted as follows. A simple Boolean path involves two nodes and the directed edge between them. This simple Boolean path can be interpreted as shown in the Fig. 6 Panel E. For the nodes X and Y with X low => Y low only quadrant #1 is sparse; the other quadrants #0, #2, and #3 are filled with samples (Fig. 6 Panel E). Assumingmonotonicity in X and Y, the quadrants can be ordered in two possible ways: 0-2-3 and 3- 2-0. The path corresponds to 0-2-3 begins with X low and Y low. This is interpreted as X turns on first and then Y turns on along a hypothetical biological path defined by the sample order. Similarly, Y turns off first and then X turns off on path 3-2-0. A complex path in the Boolean network involves more than one Boolean implication relationship (Fig. 6 Panel F). Three Boolean implication relationships can be used to group samples into five bins and the bins can be ordered in two possible ways (Fig. 6 Panel F, forward, reverse). Another example of a path is illustrated in Fig. 6 Panel G.Discover^’ of paths in clustered Boolean implication network

[0152] Paths that are transitive (such as Fig. 6, Panels F and G) are the focus because they represent a simple change in gene regulation, i.e., going from low-to-high or high-to-low once along a path (See Boolean paths above). By contrast, complex change refers to changes of gene regulation multiple times along a path, such as a gene going from high-to-low and then back to high. The discover}7of paths starts with a node that represents the biggest cluster in the CBIN. Since a path of high=>high, high=>low, and low=>low can be used to order samples, as shown in Fig. 6 Panel F, paths of this type that intersect the big clusters (top 5, based on size) were identified in the network. To maintain the transitivity this path can be expanded as the chain of high => high, followed by high => low, followed by another chain of low => low. It is desirable to keep one high => low in a path because that will cover genes that are both up- and down-regulated. Since the path A high => B high can also be written as B low => A low. the chain of high => high can be reduced to the chain of low => low in the reverse direction. Therefore, one must focus only on the high => low and chain of low => low. A simple, intuitive algorithm was developed that traverses the nodes of the CBIN, starting with the biggest cluster and greedily chooses the next big cluster connected to the nodes visited in sequence. The emphasis on cluster sizes comes from the fundamental assumption that size determines importance and relevance. Therefore, one can start from a big cluster (Al from the top 5) and identify other clusters that form a chain of low7=> low7. Further, one can identify other clusters that are either opposite to Al or they have high=>low relationship with Al, and the biggest cluster (A2) among these clusters was chosen. In addition, a chain of low=> low relationship from A2 is identified. In each subsequent step, again the biggest cluster among the different choices was greedily chosen. Finally, equivalence relationship fromeach cluster is used to gather more genes in each cluster and the whole path is clustered based on equivalence relationships. Depth-first traversal (DFS) was used to follow the path of low => low where bigger clusters are visited first. The search was performed until a cluster was reached for which there is no low => low relationships. For example, starting with cluster S, the search will return S low => Al low, Al low => A2 low, and A2 low => A3 low if A3 doesn’t have any low => low relationships. Similarly, a new starting point is considered S2 such that S2 is the biggest cluster X that has either S high => X low or S Opposite X. From cluster S2 another DFS was performed to retrieve the longest possible path of low => low. The search may return S2 low => Bl low, Bl low => B2 low if B2 doesn’t have any low => low relationships. In summary, the most prominent Boolean path was discovered by starting with the largest cluster and then exploring edges that connected to the next largest cluster in a greedy manner. This process was repeated to explore paths that connect the big clusters in the network.Scoring Boolean path for sample order

[0153] A composite score was computed for a specified Boolean path that can be used to order the sample which was consistent with the logical order. To compute the score, first the genes present in each cluster were normalized and averaged. Gene expression values were normalized according to a modified Z-score approach centered around StepMiner threshold (formula = (expr - (SThr+0.5)) / 3*stddev). Weighted linear combination of the averages from the clusters of a Boolean path was used to create a score for each sample. The weights along the path either monotonically increased or decreased to make the sample order consistent with the logical order based on BIR. The samples were ordered based on the final weighted (RNA-seq C#l-2-3-4-5: -5 for C#l, -0.3 for C#2, 0.1 for C#3, 2.9 for C#4 and -4 for C#5; qRT-PCR: -1 for PRKAA2 and CHGA, and + 1 for IL-8, LGR5, CEMIP and CLDN2; RNA-seq C#4: +1 for C#4) and linearly combined score. The direction of the path is derived from the connection from a normal colon cluster to an adenoma cluster. The sample order is visualized by a color-coded bar plot and a violin+swarm plot. A noise margin is computed for this composite score which follows the same linearly weighted combined score on 2-fold change (+ / - 0.5 around StepMiner threshold).Summary of genes in the clusters

[0154] Reactome pathway analysis of each cluster along the top continuum paths was performed to identify the enriched pathways119. The pathway description was used to summarize at a high-level what kind of biological processes are enriched in a cluster. List of genes and the pathways enriched in them are provided in Table 2.Cross-species gene name conversion

[0155] Orthologous human and mouse genes were identified using ensemble GRCh38.pl3-100 gene annotations. Human to mouse gene name conversion and vice- versa used this database.Charting the disease path from the normal colon to adenomas

[0156] BoNE uses Boolean implication network on macrophage dataset to build a signature for normal colon to adenoma disease progression. Selected clusters by size connected by high => high (green greyscale arrow), high => low (red greyscale arrows) and low => low (blue greyscale arrows) Boolean implication relationships. Reactome analysis of each clusters shows the biological processes the genes are involved in. A path is selected in the network that is used to test normal colon vs adenoma classification.Boolean analysis to chart the disease path from the normal colon to adenomas

[0157] Supervised learning was implemented, in which labelled training data of normal and adenoma states were used to train a model that can recognize a continuum of disease states. Briefly, to identify gene regulatory changes during normal colon to adenoma, the MiDReG (Mining Developmentally Regulated Genes) algorithm was employed, which utilizes statistical learning techniques.4344By applying statistical model checking to Boolean invariant rules within a static cross-sectional dataset, MiDReG infers the underlying temporal events. It identifies temporal logical changes in gene regulation by exploring transitive Boolean paths. The MiDReG algorithm was applied to analyze large and diverse normal colon and adenoma datasets (Pooled GEO; 160 normal colon, 68 adenoma), discovering Boolean invariant rules and constructing a clustered Boolean Implication Network. The model was trained by using labels from GSE76987 with normal colon (n = 41), adenoma (n = 41). The algorithm takes the adenoma network (selected graph) and this labeled dataset GSE76987 as inputs and identifies the best model torecognize the labels. The algorithm searches for transitive Boolean paths that contain three nodes with one high => low relationship. The high => low relationships cover both up / down-regulated genes and the additional Boolean path of high => high or low => low provides features to improve predictions. An unbiased search for these patterns results in 12 different Boolean paths [6, 10,11], [10,11,12], [11, 12,2], [1,2,3], [2,3,4], [6,7,8], [14,9,5], [12,2,4], [7,8,9], [1,2,4], [12,13,4], [14,15,4], The nodes on the high => high side were assigned negative weights (-2, -1, etc.) and the nodes on the low => low side assigned positive weights (1, 2 etc.) to compute an optimal composite score. Two additional paths [1,2, 3, 4, 5] and [1,2, 3, 4] were considered with weights from linear regression. Multivariate analysis was used to identify the best model. The performance of the Boolean path 1-2-3-4-5 was better than that of all the other paths.Generation of heat maps using gene clusters identified by Boolean analysis

[0158] To generate a heatmap, a Boolean path was first constructed by following the largest clusters in the Boolean Network. Genes along this path were selected to generate a heatmap that shows the gene expression values in different samples. Gene expression values were normalized according to a modified Z-score approach centered around StepMiner threshold (formula = (expr - (SThr+0.5) ) / 3*stddev). The samples are ordered according to an average of the normalized gene expression values in the largest cluster along the Boolean path. The heatmap uses red greyscales for the high values, white colors for the intermediate values and blue greyscales for low values. The selected genes are displayed on the left of the map.Human subjects

[0159] To assess the invariant gene signature and develop CRC patient-derived organoids, hereditary predisposed polyposis patients were enrolled at pediatric department at Rady Children Hospital and collection of colon biopsies from non-involved (NI) and polyp region for each patient was done following a research protocol compliant with the Human Research Protection Program (HRPP) and approved by the Institutional Review7Board (Project ID#190105). Part of the patient colon biopsies was sent to the pathology lab to assess the presence / absence of adenoma and degree CRC. In addition, the genetic mutation was assessed in each samples. Other data was gathered from each patient, such as age, gender, and previous history of the disease, following the rules of HIPAA. Eachpatient signed informed consent as an approval of the collection of colonic tissue biopsies for research purposes to generate 3D organoids. Another part of the colonic sample was processed for isolation and biobanking of organoids at the UC San Diego HUMANOID Center of Research Excellence (CoRE) (IRB: Project ID # 190105: PI Ghosh and Das). The third part of the colon samples was lysed directly and RNA from the tissues was used to assess the gene expression signatures. The study design and the use of human study participants was conducted in accordance with the criteria set by the Declaration of Helsinki.

[0160] For immunohistochemical (1HC) analysis of human tissue specimens, archived formaldehyde-fixed paraffin-embedded (FFPE) human colonic biopsies from healthy controls or patients with adenomas and / or carcinomas were obtained from the Gastroenterology Division, VA San Diego (IRB# 1132632).Animal studies

[0161] Colon tissues were collected from NI and polyp regions of WT C57BL / 6 and CPC-APCMmmice. Part of the colon samples were lysed directly to isolate RNA for assessment of invariant gene signatures. The other part of colon samples was used to isolate intestinal crypts for enteroid isolation. The previously validated CPC-APC (APCMir' / +CDX2-Cre) and control littermate mice were a kind gift from Michael Karin (UCSD); these mice are known to develop overt polyposis in the colon around the age of~3-4 months67 120Animals were bred, housed and euthanized according to all University’ of California San Diego Institutional Animal Care and Use Committee (IACUC) policies under the animal protocol number S I 8086.Bacterial culture

[0162] Fusobacterium nucleatum (ATCC-25586) was cultured in the chopped cooked meat media in an anaerobic chamber, including an anaerobic gas kit generating system (MGC, AnaeroPACK System, Japan).Development of 3D colon organoids from mouse and human colons

[0163] Stem cells were isolated from the colonic crypts of mouse and human tissue specimens by digesting with Collagenase type I [2 mg / ml; Life Technologies Corporation,NY) and cultured in stem-cell enriched conditioned media (CM) containing WNT 3a, R- spondin and Noggin as described previously102>121-123Preparation of enteroid-derived monolayers (EDMs) and infection with microbes

[0164] Enteroid-derived monolayers (EDMs) were prepared from 3D colon spheroids isolated from non-involved regions of CPC-APC^"^ mice as previously described102,124Briefly, the organoids were digested with trypsin and the isolated filtered cells were resuspended in 5% CM and plated with matrigel in 0.4 m polyester transwell membrane (Coming. Cat #3470) at the density of 2x105cell / well. EDMs were differentiated for 2 days, and then were challenged with live microbes Fusobacterium nucleatum (Fn) at a multiplicity of infection 1: 100 as described before124The supernatant was collected from the basolateral and apical part of the transwell for cytokine analysis, and the cells were lysed for RNA extraction, followed by an analysis of target gene expression by qPCR.Immunohistochemistry for CEMIP and Lgr-5 in organoids

[0165] Thick sections of 4 pm formalin-fixed, paraffin-embedded (FFPE) tissues were cut and placed on glass slides coated with poly-L-lysine, followed by deparaffinization and hydration. Citrate buffer (pH-6) and a pressure cooker were used to perform the heat-induced epitope step. Tissue / organoid sections were incubated with 3% hydrogen peroxidase for 10-15 min to block endogenous peroxidase activity, followed by incubation with primary antibodies for 30-90 min in a humidified chamber at room temperature. Rabbit polyclonal anti-KIAAl 199 antibody [1 :50], and mouse monoclonal anti -LGR-5 antibody (1:250, SANTA CRUZ BIOTECHNOLOGY, INC) was used for immunostaining of the organoids, and anti-AMPK 2 (1:50, anti-rabbit, Abeam) was used for immunostaining of tissue section. A labeled streptavidin-biotin using 3,3'- diaminobenzidine was used as a chromogen and hematoxylin for visualization. Samples were quantitatively analyzed and scored based on the intensity of staining using the following scale: 0 to 3, where 0 = no staining, 1 = light brown, 2 = brown, and 3 = dark brown. Data is expressed as the frequency of staining scores, and a Chi-square test was used to determine significance.Immunofluorescence of FFPE-embedded patient-derived organoids for Claudin 2 and Lgr- 5

[0166] FFPE organoids were placed on glass slides, followed by deparaffinization and hydration. The antigen retrieval step was done as described in IHC. then the slides were incubated with a blocking buffer at room temperature for 1 hr. Rabbit polyclonal anti-Claudin 2 antibody (1:500, Abeam) and mouse monoclonal anti-LGR-5 antibody (1:250, SANTA CRUZ BIOTECHNOLOGY, INC) were incubated with the slides at 4°C for overnight. DAPI (1: 1000) was used for Dapi staining. Goat anti-rabbit Alexa 594 (1:500 dilution) and goat anti-mouse Alexa 488 (1:500 dilution) were used as secondary antibodies for Claudin 2 and Lgr-5, respectively. Leica CTR4000 Confocal Microscope was used for image capture, Image J software for image processing, and Illustrator software (Adobe) for image assemble and presentation.RNA extraction and quantitative-(q)RT-PCR

[0167] Total RNA was extracted from NI and polyp colon biopsies of genetically predisposed CRC patients and CPC-APC mouse model RNA MiniPrep Kit (Zymo Research, USA) according to the manufacturer's instructions. RNA was isolated from patient-derived organoids, APC mouse organoids, 2D EDMs challenged with different microbes using the Quick-RNA MicroPrep Kit (Zymo Research, USA) according to the manufacturer's instruction. cDNA was prepared by the cDNA mastermix (qScript™, Quantabio). Quantitative RT-PCR (qPCR) was carried out using 2x SYBR Green qPCR Master Mix (Biotool™. USA), and the cycle threshold (Ct) of target genes was normalized to 18s rRNA gene. The fold change in the mRNA expression was determined using the AACt method. Primers used in qPCR reactions were designed using NCBI Primer Blast software and Roche Universal Probe Library Assay Design software.Quantification of mouse and human IL-8 by ELISA

[0168] Supernatants collected from organoid culture media of genetically predisposed patients and from the basolateral part of uninfected and infected monolayer were tested for IL-8 using Human IL-8 ELISA kit (Biolegend, San Diego. CA)) and CXCL1 / KC DuoSet ELISA kit (R&D Systems, Minneapolis, MN), respectively according to the manufacturer’s instructions. IL-8 levels were compared between healthy organoids and PDO and betw een uninfected EDMs and infected EDMs.

[0169] Statistical analyses

[0170] Data are expressed as the mean + / - S.E.M, unless stated otherwise. Statistical significance is determined using Mann-Whitney, Student’s t-test and one-way ANOVA test, p value is < 0.05 considers significant. The graphs were generated using GraphPad version 8. Statistical tests used in each data is indicated in each figure. Statistical tests were performed using R version 3.2.3 (2015-12-10) and skleam python packages. Heatmaps are generated using python matplotlib packages.

[0171] Barrett’s Esophagus and Gastroesophageal Junction [GEJ] PrecancerousOrganoid Models

[0172] Currently, there are no patient derived organoids (PDOs, expandable lines) for these pre-cancer conditions. The next best option, which is used by others is primary epithelial cells from barrett’s esophagus or GEJs. What follows are detailed protocols for creating Barrett’s Esophageal PDOs and PDOs from GEJs.

[0173] Table 6: Barrett's Esophagus / GEJ Organoid Media

[0174] Table 7: Normal Esophagus / GEJ Organoid MediaNormal / Barretts Esophagus / GEJ Adult Stem Cell Isolation ProtocolBiopsy Acquisition

[0175] Due to the structure of Barrett’s and Normal Esophagus tissues, it is crucial to take a biopsy using jumbo forceps to get a tissue sample deep enough to reach the basal layer where the stem cells are found.Reagents Needed

[0176] PBS - 2 mL; Penicillin / Streptomycin - 80 pL; Collagenase Type I - 750 pL (Add 150 pL collagenase Type I (12500 U / mL) to 3.6mL GI wash. Filtrate with a 0.22 pM syringe filter. Aliquot into 750 pL, store at -20°C. Working concentration of 4mg / mL); Dispase - 750 pL (Add 180pL dispase (from 5U / mL stock) to 570pL PBS, per Eppendorf for a 750pL aliquot of 1.2U / mL. Aliquots are stored at -20°C.); Primocin - 7 pL; Wash Media - 19 mL (Adv. DMEM / F12 + 10% FBS and 1 : 100 Glutamax and 1: 100 PenStrep); Grow th Media (BE Media or ESO Media) - 500 pLOther Consumables to Have Ready

[0177] 1, dry 60-mm Petri dish (This dish is where you will cut and digest the tissue); 1, 70 pm cell strainer; 1, 50 mL Conical Tube (For 70 pm cell strainer); 1, 15 mL Conical Tube (For pelleting strained cells); 1, 1.5 mL Eppendorf Tube; 1, 24 Well Tissue Culture Treated Plate.Workstation Preparation

[0178] 1. Turn on Germinator 500 (Takes -30 minutes to get up to temperature).

[0179] 2. Prepare 2 15 mL conical tubes labeled Wash 1 and Wash 2, each containing 3 mL of GI wash. Prepare 1 15 mL conical tube with 12 mL of GI wash. Prepare a 1.5 mL Eppendorf with 1 mL of GI Wash.

[0180] 3. Prepare 1 1.5 mL Eppendorf Tube by adding 750pL of collagenase I,750pL of dispase I, 3 pL of Primocin (1 :500) (keep on ice until use).

[0181] 4. Prepare 1 60-mm Petri dish by adding 2 mL of cold PBS, 80 uL ofPenicillin / Streptomycin to make 4X, and 4 uL of Primocin to the dish.

[0182] 5. Once Germinator 500 is preheated, spray surgical scissors with ethanol and place in Germinator 500 for 30 seconds. Remove and place in hood ensuring they remain sterile and do not touch anything (can be placed in a sterile autoclaved beaker or upside down in a 1.5 mL tube rack)Isolation Procedure

[0183] 1. Thaw cryovial containing biopsy in water bath until biopsy is no longer surrounded by ice (~1.5 minutes).

[0184] 2. Use a P1000 Pipette to suction to the Biopsy and remove from the vial try ing to transfer as little media as possible, and transfer to tube containing 3 mL of GI wash labeled Wash 1. Discard tip.

[0185] 3. With new tip. pipette media at the biopsy 10 to 15 times being careful to not suck up the biopsy.

[0186] 4. Once the biopsy is thoroughly washed, transfer it to the tube containing 3 mL of GI wash labeled Wash 2 and discard tip. Repeat wash procedure from step 3.

[0187] 5. Transfer biopsy to petti dish containing 2 mL of cold PBS, 80 uL ofPenicillin / Streptomycin, and 4 uL of Primocin and discard tip. Repeat wash procedure from step 3.

[0188] 6. After biopsy is sufficiently washed, transfer to dry760-mm Petri dish.Mince tissue with fine surgical scissors. Tissue fragments should be small enough to pipette through a 1000 pL pipette tip. (You only need to use the forceps if tissue gets stuck to the surgical scissors).

[0189] 7. Add Collagenase / Dispase mixture to the petri dish containing tissue fragments (Wash any tissue off the surgical scissors with Pl 000 while adding mixture to dish). Pipette up and down and incubate at 37°C (Save P1000 Tip).

[0190] 8. During incubation pipette up and down every 10 minutes, attempting to break dow n larger tissue fragments against the surface of the dish.

[0191] 9. Check progress of digestion under microscope. Repeat incubation for 5-10 min and mechanical breakage until a majority of the single epithelial cells have separated from larger tissue fragments.

[0192] 10. When 80-90% of cells have been released (—3-6 rounds of incubation) prepare to fdter tissue. Set a 70-pm cell strainer in a 50 mL conical tube and wet with 1 mL of GI wash media.

[0193] 11. Add 2 mL of GI wash media to tissue solution to neutralize collagenase and pipette onto filter.

[0194] 12. Wash dish 2-3 more times with GI wash media and pipette onto filter until all 12 mL of GI Wash is used.

[0195] 13. Place 70-pm cell strainer in the cap of the 50 mL conical tube and set aside (Used for Myofibroblast Isolation)

[0196] 14. With a new P1000 tip, transfer cells and media from 50 mL tube into empty 15 mL conical tube.

[0197] 15. Centrifuge for 4 min at 300 X g.

[0198] 16. Carefully aspirate the supernatant using a 10 pL tip for gentle suction.

[0199] 17. Resuspend the pellet in 1 mL of GI wash media (from 1.5 mLEppendorf tube), pipette up and down 4-5 times, and transfer to an Eppendorf tube.

[0200] 18. Centrifuge for 4 min at 300 X g.

[0201] 19. Carefully aspirate the supernatant and place the tube on ice.

[0202] 20. Resuspend the cells in 15 pL of Matrigel per well of a 24 well plate(Generally, 1 fresh biopsy can be seeded into 1 well of a 24 well plate).

[0203] 21. Keep tube on ice and pipette up and down gently to break up cell pellet and evenly disperse cells in matrigel. Avoid generating air bubbles.

[0204] 22. Add 15 pl of cell-Matrigel suspension to the center of each well using a20-pl pipette. Spread using the pipette tip.

[0205] 23. Flip the plate upside down and incubate at 37°C for 10 min.

[0206] 24. Add 500 pl of Growth Media to each well.Organoid Maintenance

[0207] Aspirate media from wells and replace with fresh growth media every 2-3 days. Split organoids every 7 days (See below for protocol).Normal / Barretts Esophagus / GEJ Organoid Splitting ProtocolReagents Needed

[0208] PBS-EDTA: 1 : 1000 0.5M EDTA in IX PBS; TrypLE; Wash Media: (DMEM / F12 ham nutrient mix + 10% FBS and 1: 100 Glutamax and 1 : 100 Pen Strep); Trypan Blue; Matrigel; BE or ESO Growth Media.

[0209] Workstation Preparation

[0210] 1. Label 4 properly sized tubes with the names and volumes needed PBS-EDTA, TrypLE, Gi Wash, and one tube with the cell line name.

[0211] 2. Add 10 uL of Trypan Blue to a 1 .5 mL Eppendorf Tube for cell count.

[0212] 3. Label an additional 1.5 mL Eppendorf Tube with cell line name for final cell pellet.

[0213] 4. Aliquot appropriate volumes of reagents into respective labeled tubes.

[0214] Organoid Splitting Procedure

[0215] 1. Aspirate media from edge of wells. Use a 10 pl tip throughout for gentle aspiration.

[0216] 2. Add 1 mL of PBS-EDTA to each well. Scratch the Matrigel from the surface of the well with the same pl 000 tip.

[0217] 3. Collect cell suspension from each well and add it to the cell collection tube labeled with the cell line name.

[0218] 4. Once all wells have been collected, Use the extra PBS-EDTA to wash the wells one more time than add to cell collection tube.

[0219] 5. Centrifuge @ 300g for 4 min.

[0220] 6. Aspirate supernatant and add TrypLE and gently pipette a few- times to break up pellet. Put in 37°C water bath for 4-5 min. (Allow longer incubation for larger, healthier spheroids)

[0221] 7. After incubation, pipette up and down vigorously (around 10-15 times) to further break apart spheroids.

[0222] 8. Add wash media to the tube to inactivate TrypI.E (Add 1 :2 TrypLE: wash media ratio). Pipette up and dow n to mix.

[0223] 9. Centrifuge @ 300g for 4 min.

[0224] 10. Aspirate the supernatant and resuspend in 1 mL of wash media. Quickly after resuspension collect 10 uL of the cell suspension and add it to the tube containing Trypan Blue.

[0225] 11. Move the suspension to the sterile Eppendorf tube with the cell line name.

[0226] 12. Centrifuge @ 300g for 4 min.

[0227] 13. Count the cells by taking 10 uL from the cell / trypan blue mixture and using automated cell counter or hemocytometer. Record cell count.

[0228] 14. Aspirate the supernatant and resuspend with the appropriate amount ofMatrigel (see table). Be sure no bubbles are introduced.

[0229] 15. Place the plate on ice.

[0230] 16. Add equal volume of Matrigel / cell mixture to the center of each well and spread into circle with the pipette tip, avoiding bubbles. The Matrigel should not touch the edge of the well.

[0231] 17. Flip the plate upside down and incubate for 10 min @ 37°C

[0232] 18. Add the appropriate amount of growth media to each well.Results

[0233] Figure 17 shows Barrett's Esophagus Organoid Images.

[0234] Figure 18 shows Normal Esophagus Organoid Images.

[0235] Figure 19 shows Barrett’s Gastroesophageal Junction Organoid Images.

[0236] Figure 20 shows Normal Gastroesophageal Junction Organoid Images.

[0237] Figure 21 shows Barrett’s Esophagus Organoids H&E Images.

[0238] Figure 22 shows Barrett’s Esophagus Immunofluorescence ImagesREFERENCES

[0239] 1. Bray, F. et al. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 68, 394-424, doi: 10.3322 / caac.21492 (2018).

[0240] 2. Arnold, M. et al. Global patterns and trends in colorectal cancer incidence and mortality. Gut 66, 683-691, doi: 10.1136 / gutjnl-2015-310912 (2017).

[0241] 3. Bonadona, V. et al. Cancer risks associated with germline mutations inMLH1, MSH2, and MSH6 genes in Lynch syndrome. JAMA 305, 2304-2310, doi : 10. 1001 / j ama.2011.743 (2011 ).

[0242] 4. Bulow, S., Berk, T. & Neale, K. The history of familial adenomatous polyposis. Fam Cancer 5, 213-220, doi: 10.1007 / sl0689-005-5854-0 (2006).

[0243] 5. Fodde, R. The APC gene in colorectal cancer. Eur J Cancer 38, 867-871, doi: 10. 1016 / s0959-8049(02)00040-0 (2002).

[0244] 6. Pearlman, R. et al. Prevalence and Spectrum of Germline CancerSusceptibility Gene Mutations Among Patients With Early-Onset Colorectal Cancer. JAMA Oncol 3, 464-471, doi:10.1001 / jamaoncol.2016.5194 (2017).

[0245] 7. Cobum. M. C., Pricolo, V. E.. DeLuca, F. G. & Bland, K. I. Malignant potential in intestinal juvenile polyposis syndromes. Ann Surg Oncol 2, 386-391, doi: 10.1007 / BF02306370 (1995).

[0246] 8. Hearle, N. et al. Frequency and spectrum of cancers in the Peutz-Jeghers syndrome. Clin Cancer Res 12, 3209-3215, doi: 10.1158 / 1078-0432.CCR-06-0083 (2006).

[0247] 9. Collaborators, G. B. D. C. C. Global, regional, and national burden of colorectal cancer and its risk factors, 1990-2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet Gastroenterol Hepatol 7, 627-647, doi: 10. 1016 / S2468-1253(22)00044-9 (2022).

[0248] 10. Cao. Y. et al. Commensal microbiota from patients with inflammatory bowel disease produce genotoxic metabolites. Science 378, eabm3233, doi: 10. 1126 / science.abm3233 (2022).

[0249] 11. Yachida, S. et al. Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer. Nature medicine 25, 968-976, doi: 10. 1038 / s41591-019-0458-7 (2019).

[0250] 12. Bischoff, S. C. et al. Intestinal permeability— a new target for disease prevention and therapy. BMC Gastroenterol 14, 189, doi: 10.1186 / sl2876-014-0189-7 (2014).

[0251] 13. McCoy. A. N. et al. Fusobacterium is associated with colorectal adenomas. PLoS One 8, e53653, doi: 10. 1371 / joumal. pone.0053653 (2013).

[0252] 14. Warren, R. L. et al. Co-occurrence of anaerobic bacteria in colorectal carcinomas. Microbiome 1, 16, doi: 10.1186 / 2049-2618-l-16 (2013).

[0253] 15. Genua, F.. Raghunathan, V., Jenab, M., Gallagher, W. M. & Hughes,D. J. The Role of Gut Barrier Dysfunction and Microbiome Dysbiosis in Colorectal Cancer Development. Front Oncol 11, 626349, doi: 10.3389 / fonc.2021.626349 (2021).

[0254] 16. Dejea, C. M. et al. Patients with familial adenomatous polyposis harbor colonic biofdms containing tumorigenic bacteria. Science 359, 592-597, doi: 10. 1126 / science.aah3648 (2018).

[0255] 17. Tomkovich, S. & Jobin, C. Microbial networking in cancer: when two toxins collide. Br J Cancer 118, 1407-1409, doi: 10.1038 / s41416-018-0101-2 (2018).

[0256] 18. Pieters, W. et al. Pro-mutagenic effects of the gut microbiota in a Lynch syndrome mouse model. Gut Microbes 14, 2035660, doi: 10. 1080 / 19490976.2022.2035660 (2022).

[0257] 19. Sayed, I. M., Ramadan, H. K.-A., El-Mokhtar, M. A. & Abdel-Wahid,L. Microbiome and gastrointestinal malignancies. Current Opinion in Physiology 22, 100451 , doi : https : / / doi . org / 10. 1016 / j . cophy s.2021.06.005 (2021 ) .

[0258] 20. Arthur, J. C. et al. Intestinal inflammation targets cancer-inducing activity of the microbiota. Science 338, 120-123, doi: 10. 1 126 / science. 1224820 (2012).

[0259] 21. Wu, S. et al. A human colonic commensal promotes colon tumorigenesis via activation of T helper type 17 T cell responses. Nature medicine 15, 1016-1022, doi: 10.1038 / nm.2015 (2009).

[0260] 22. Kostic, A. D. et al. Fusobacterium nucleatum potentiates intestinal tumorigenesis and modulates the tumor-immune microenvironment. Cell host & microbe 14, 207-215, doi: 10. 1016 / j.chom.2013.07.007 (2013).

[0261] 23. Sayed IM et al. The DNA Glycosylase NEIL2 SuppressesFusobacterium-Infection-Induced Inflammation and DNA Damage in Colonic Epithelial Cells. Cells 9, doi: 10.3390 / cells9091980 (2020).

[0262] 24. Gagnaire, A., Nadel, B., Raoult, D., Neefjes, J. & Gorvel, J. P.Collateral damage: insights into bacterial mechanisms that predispose host cells to cancer. Nat Rev Microbiol 15, 109-128, doi: 10.1038 / nrmicro.2016.171 (2017).

[0263] 25. Butte, A. J. & Kohane, I. S. Unsupervised knowledge discovery in medical databases using relevance networks. ProcAMIA Symp, 711-715 (1999).

[0264] 26. Shameer, K., Readhead, B. & Dudley, J. T. Computational and experimental advances in drug repositioning for accelerated therapeutic stratification. Curr TopMedChem 15, 5-20, doi: 10.2174 / 1568026615666150112103510 (2015).

[0265] 27. Shen. Y. et al. Systematic, network-based characterization of therapeutic target inhibitors. PLoS Comput Biol 13, el 005599, doi: 10. 1371 / joumal.pcbi. 1005599 (2017).

[0266] 28. Margolin, A. A. et al. Reverse engineering cellular networks. NatProtoc 1, 662-671, doi: 10.1038 / nprot.2006.106 (2006).

[0267] 29. van Someren, E. P., Wessels, L. F., Backer, E. & Reinders, M. J.Genetic network modeling. Pharmacogenomics 3, 507-525, doi : 10. 1517 / 14622416.3.4.507 (2002).

[0268] 30. Kusonmano, K. Gene Expression Analysis Through Network Biology:Bioinformatics Approaches. Adv Biochem Eng Biotechnol 160, 15-32, doi: 10.1007 / 10_2016_44 (2017).

[0269] 31. Allocco, D. J., Kohane, I. S. & Butte, A. J. Quantifying the relationship between co-expression, co-regulation and gene function. BMC Bioinformatics 5, 18, doi: 10. 1186 / 1471-2105-5-18 (2004).

[0270] 32. Arkin, A. & Ross, J. Statistical construction of chemical reaction mechanisms from measured time-series. The Journal of Physical Chemistry 99, 970-979 (1995).

[0271] 33. Jordan, I. K., Marino-Ramirez. L., Wolf, Y. I. & Koonin. E. V.Conservation and coevolution in the scale-free human gene coexpression network. Mol Biol Evol 21, 2058-2070, doi:10.1093 / molbev / msh222 (2004).

[0272] 34. Lee, H. K., Hsu, A. K., Sajdak. J., Qin, J. & Pavlidis, P. Coexpression analysis of human genes across many microarray data sets. Genome Res 14, 1085-1094, doi: 10.1101 / gr.1910904 (2004).

[0273] 35. Tavazoie, S., Hughes, J. D., Campbell, M. J., Cho, R. J. & Church, G.M. Systematic determination of genetic network architecture. Nat Genet 22, 281-285, doi: 10. 1038 / 10343 (1999).

[0274] 36. Butte, A. J., Tamayo, P., Slonim, D._ Golub. T. R. & Kohane, I. S.Discovering functional relationships between RNA expression and chemotherapeutic susceptibility using relevance networks. Proc Natl Acad Set U S A 97, 12182-12186, doi: 10. 1073 / pnas.220392197 (2000).

[0275] 37. Paik, S. et al. A multigene assay to predict recurrence of tamoxifen- treated, node-negative breast cancer. N Engl J Med 351, 2817-2826. doi:10.1056 / NEJMoa041588 (2004).

[0276] 38. Witten, D. M. & Tibshirani, R. Survival analysis with high-dimensional covariates. Stat Methods Med Res 19, 29-51, doi: 10. 1177 / 0962280209105024 (2010).

[0277] 39. Alizadeh, A. A. et al. Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling. Nature 403, 503-511, doi: 10. 1038 / 35000501 (2000).

[0278] 40. Zhao, H. et al. Gene expression profiling predicts survival in conventional renal cell carcinoma. PLoS Med 3, el3. doi: 10. 1371 / joumal.pmed.0030013 (2006).

[0279] 41. Ghosh, P. et al. Machine learning identifies signatures of macrophage reactivity and tolerance that predict disease outcomes. EBioMedicine 94, 104719, doi: 10.1016 / j.ebiom.2023. 104719 (2023).

[0280] 42. Sahoo, D. et al. Artificial intelligence guided discovery of a barrier- protective therapy in inflammatory bowel disease. Nat Commun 12, 4246, doi: 10. 1038 / s41467-021-24470-5 (2021).

[0281] 43. Sahoo, D. et al. MiDReG: a method of mining developmentally regulated genes using Boolean implications. Proc Natl Acad Set U S A 107, 5732-5737, doi: 10. 1073 / pnas.0913635107 (2010).

[0282] 44. Inlay, M. A. et al. Ly6d marks the earliest stage of B-cell specification and identifies the branchpoint between B-cell and T-cell development. Genes Dev 23,

[0284] 46. Chan, C. K. et al. Identification and specification of the mouse skeletal stem cell. Cell 160, 285-298, doi:10.1016 / j.cell.2014.12.002 (2015).

[0285] 47. Dimov, I. K. et al. Discriminating cellular heterogeneity using microwell-based RNA cytometry. Nat Commun 5, 3451, doi: 10.1038 / ncomms4451 (2014).

[0286] 48. Chan, C. K. et al. Clonal precursor of bone, cartilage, and hematopoietic niche stromal cells. Proc Natl Acad Set U S A 110, 12643-12648, doi: 10. 1073 / pnas. 1310212110 (2013).

[0287] 49. Seita, J. et al. Gene Expression Commons: an open platform for absolute gene expression profiling. PLoS One 7, e40321, doi: 10.1371 / joumal.pone.0040321 (2012).

[0288] 50. Cheah, M. T. et al. CD14-expressing cancer cells establish the inflammatory and proliferative tumor microenvironment in bladder cancer. Proc Natl Acad Sci USA 112, 4725-4730, doi: 10. 1073 / pnas. 1424795112 (2015).

[0289] 51. Volkmer, J. P. et al. Three differentiation states risk-stratify bladder cancer into distinct subtypes. Proc Natl Acad Sci U S A 109, 2078-2083, doi: 10. 1073 / pnas. 1120605109 (2012).

[0290] 52. Shin, K. et al. Hedgehog signaling restrains bladder cancer progression by eliciting stromal production of urothelial differentiation factors. Cancer Cell 26, 521- 533, doi: 10.1016 / j.ccell.2014.09.001 (2014).

[0291] 53. Sin, M. L. Y. el al. Deep Sequencing of Urinary RNAs for BladderCancer Molecular Diagnostics. Clin Cancer Res 23, 3700-3710, doi: 10. 1158 / 1078- 0432. CCR- 16-2610 (2017).

[0292] 54. Sahoo, D. et al. Boolean analysis identifies CD38 as a biomarker of aggressive localized prostate cancer. Oncotarget 9. 6550-6561, doi: 10.18632 / oncotarget.23973 (2018).

[0293] 55. Bhamre, S., Sahoo, D., Tibshirani, R., Dill, D. L. & Brooks, J. D.Temporal changes in gene expression induced by sulforaphane in human prostate cancer cells. Prostate 69, 181-190, doi: 10.1002 / pros.20869 (2009).

[0294] 56. Ghosh, P. et al. Al-assisted discovery’ of an ethnicity-influenced driver of cell transformation in esophageal and gastroesophageal junction adenocarcinomas. JCl Insight 7, doi: 10. 1172 / jci. insight. 161334 (2022).

[0295] 57. Sinha, S. Reproducibility of parameter learning with missing observations in naive Wnt Bayesian network trained on colorectal cancer samples and doxycycline-treated cell lines. Mol Biosyst 11, 1802-1819, doi: 10.1039 / c5mb00117j (2015).

[0296] 58. Du, G. et al. Comparative gene expression profiling of normal and human colorectal adenomatous tissues. Oncol Lett 8, 2081-2085, doi: 10.3892 / ol.2014.2485 (2014).

[0297] 59. Lee, S., Bang, S., Song, K. & Lee, I. Differential expression in normal- adenoma-carcinoma sequence suggests complex molecular carcinogenesis in colon. Oncol Rep 16, 747-754 (2006).

[0298] 60. Druliner, B. R. et al. Molecular characterization of colorectal adenomas with and without malignancy reveals distinguishing genome, transcriptome and methylome alterations. Sci Rep 8, 3161, doi: 10. 1038 / s41598-018-21525-4 (2018).

[0299] 61. Druliner, B. R. et al. Colorectal Cancer with Residual Polyp of Origin:A Model of Malignant Transformation. Transl Oncol 9, 280-286, doi : 10. 1016 / j . tranon.2016.06.002 (2016).

[0300] 62. Druliner, B. R. et al. Time Lapse to Colorectal Cancer: TelomereDynamics Define the Malignant Potential of Polyps. Clin Transl Gastroenterol 7, el 88, doi : 10. 1038 / ctg.2016.48 (2016).

[0301] 63. Kim, T. M. et al. Clonal origins and parallel evolution of regionally synchronous colorectal adenoma and carcinoma. Oncotarget 6, 27725-27735, doi: 10. 18632 / oncotarget.4834 (2015).

[0302] 64. Kostic, A. D. et al. Genomic analysis identifies association ofFusobacterium with colorectal carcinoma. Genome research 22, 292-298, doi: 10. 1101 / gr. 126573. 111 (2012).

[0303] 65. Yu, T. et al. Fusobacterium nucleatum Promotes Chemoresistance toColorectal Cancer by Modulating Autophagy. Cell 170, 548-563. e516, doi: 10. 1016 / j. cell.2017.07.008 (2017).

[0304] 66. Castellarin, M. et al. Fusobacterium nucleatum infection is prevalent in human colorectal carcinoma. Genome Res 22, 299-306, doi: 10.1101 / gr. 126516.111 (2012).

[0305] 67. Hinoi, T. et al. Mouse model of colonic adenoma-carcinoma progression based on somatic Ape inactivation. Cancer Res 67, 9721-9730, doi: 10. 1158 / 0008-5472.Can-07-2735 (2007).

[0306] 68. Bosch, F. X. et al. Prevalence of human papillomavirus in cervical cancer: a worldwide perspective. International biological study on cervical cancer (IBSCC) Study Group. J Natl Cancer Inst 87, 796-802, doi: 10. 1093 / jnci / 87.11.796 (1995).

[0307] 69. El-Serag, H. B. Epidemiology of viral hepatitis and hepatocellular carcinoma. Gastroenterology 142, 1264-1273. el261. doi: 10.1053 / j.gastro.2011.12.061 (2012).

[0308] 70. van Elsland, D. & Neefjes, J. Bacterial infections and cancer. EMBORep 19, doi:10.15252 / embr.201846632 (2018).

[0309] 71. Gupta. S. et al. Recommendations for Follow-Up After Colonoscopy and Polypectomy: A Consensus Update by the US Multi-Society Task Force on Colorectal Cancer. Gastroenterology 158, 1131-1153. el l35, doi: 10.1053 / j.gastro.2019.10.026 (2020).

[0310] 72. Lieberman, D. A. et al. Guidelines for colonoscopy surveillance after screening and polypectomy: a consensus update by the US Multi-Society Task Force on Colorectal Cancer. Gastroenterology! 143, 844-857, doi: 10.1053 / j.gastro.2012.06.001 (2012).

[0311] 73. Short, M. W„ Layton, M. C„ Teer, B. N. & Domagalski, J. E.Colorectal cancer screening and surveillance. Am Fam Physician 91, 93-100 (2015).

[0312] 74. Wang, Y. et al. Application of artificial intelligence to the diagnosis and therapy of colorectal cancer. Am J Cancer Res 10, 3575-3598 (2020).

[0313] 75. Echle, A. et al. Artificial intelligence for detection of microsatellite instability in colorectal cancer-a multicentric analysis of a pre-screening tool for clinical application. ESMO Open 7, 100400, doi: 10. 1016 / j.esmoop.2022. 100400 (2022).

[0314] 76. Yu, G. et al. Accurate recognition of colorectal cancer with semisupervised deep learning on pathological images. Nat Commun 12, 6311,

[0316] 78. Kudo, S. E. et al. Artificial Intelligence System to Determine Risk ofT1 Colorectal Cancer Metastasis to Lymph Node. Gastroenterology 160, 1075- 1084.el072, doi: 10.1053 / j.gastro.2020.09.027 (2021).

[0317] 79. Slodkowska, E. A. & Ross, J. S. MammaPrint 70-gene signature: another milestone in personalized medical care for breast cancer patients. Expert Rev Mol Diagn 9, 417-422, doi: 10.1586 / erm.09.32 (2009).

[0318] 80. Bedard, P. L. et al. MammaPrint 70-gene profde quantifies the likelihood of recurrence for early breast cancer. Expert Opin Med Diagn 3, 193-205, doi: 10.1517 / 17530050902751618 (2009).

[0319] 81. Maak, M. et al. Independent validation of a prognostic genomic signature (ColoPrint) for patients with stage II colon cancer. Ann Surg 257, 1053-1058, doi: 10. 1097 / SLA.0b013e31827cl 180 (2013).

[0320] 82. Kopetz, S. et al. Genomic classifier ColoPrint predicts recurrence in stage II colorectal cancer patients more accurately than clinical factors. Oncologist 20, 127-133, doi: 10.1634 / theoncologist.2014-0325 (2015).

[0321] 83. Tan, I. B. & Tan, P. Genetics: an 18-gene signature (ColoPrint®) for colon cancer prognosis. Nat Rev Clin Oncol 8, 131-133, doi: 10.1038 / nrclinonc.2010.229 (2011).

[0322] 84. Hartmans, E. et al. Functional Genomic mRNA Profiling of ColorectalAdenomas: Identification and in vivo Validation of CD44 and Splice Variant CD44v6 as Molecular Imaging Targets. Theranoslics 7, 482-492, doi: 10.7150 / thno. 16816 (2017).

[0323] 85. Fink, S. P. et al. Induction of KIAA1199 / CEMIP is associated with colon cancer phenotype and poor patient survival. Oncotarget 6, 30500-30515, doi: 10. 18632 / oncotarget.5921 (2015).

[0324] 86. Birkenkamp-Demtroder, K. et al. Repression of KIAA1199 attenuatesWnt-signalling and decreases the proliferation of colon cancer cells. Br J Cancer 105, 552-561, doi: 10.1038 / bjc.2011.268 (2011).

[0325] 87. Rosenthal, R. et al. Claudin-2, a component of the tight junction, forms a paracellular water channel. J Cell Sci 123, 1913-1921, doi:10.1242 / jcs.060665 (2010).

[0326] 88. Amasheh, S. et al. Claudin-2 expression induces cation-selective channels in tight junctions of epithelial cells. J Cell Sci 115, 4969-4976. doi: 10.1242 / jcs.00165 (2002).

[0327] 89. Luettig, J., Rosenthal, R., Barmeyer, C. & Schulzke, J. D. Claudin-2 as a mediator of leaky gut barrier during intestinal inflammation. Tissue Barriers 3, e977176, doi: 10.4161 / 21688370.2014.977176 (2015).

[0328] 90. Zeissig, S. et al. Changes in expression and distribution of claudin 2, 5 and 8 lead to discontinuous tight junctions and barrier dysfunction in active Crohn's disease. Gut 56, 61-72, doi: 10. 1136 / gut.2006.094375 (2007).

[0329] 91. de Sousa e Melo, F. et al. A distinct role for Lgr5(+) stem cells in primary and metastatic colon cancer. Nature 543, 676-680. doi: 10.1038 / nature21713 (2017).

[0330] 92. Dhawan, P. et al. Claudin-2 expression increases tumorigemcity of colon cancer cells: role of epidermal growth factor receptor activation. Oncogene 30, 3234-3247, doi:10.1038 / onc.2011.43 (2011).

[0331] 93. Kinugasa, T. et al. Selective up-regulation of claudin-1 and claudin-2 in colorectal cancer. Anticancer Res 27, 3729-3734 (2007).

[0332] 94. Paquet-Fifield, S. et al. Tight Junction Protein Claudin-2 PromotesSelf-Renewal of Human Colorectal Cancer Stem-like Cells. Cancer Res 78, 2925-2938, doi: 10. 1158 / 0008-5472.Can-17-1869 (2018).

[0333] 95. Xia, W. et al. Prognostic value, clinicopathologic features and diagnostic accuracy of interleukin-8 in colorectal cancer: a meta-analysis. PLoS One 10, e0123484, doi: 10.1371 / joumal.pone.0123484 (2015).

[0334] 96. Ning, Y. et al. Interleukin-8 is associated with proliferation, migration, angiogenesis and chemosensitivity in vitro and in vivo in colon cancer cell line models. Int J Cancer 128, 2038-2049, doi: 10.1 02 / ijc.25562 (201 1).

[0335] 97. Jin, W. J., Xu, J. M., Xu, W. L., Gu, D. H. & Li, P. W. Diagnostic value of interleukin-8 in colorectal cancer: a case-control study and meta-analysis. World J Gastroenterol 20. 16334-16342, doi: 10.3748 / wjg.v20.i43. 16334 (2014).

[0336] 98. Rubie, C. et al. Correlation of IL-8 with induction, progression and metastatic potential of colorectal cancer. World J Gastroenterol 13, 4996-5002, doi : 10.3748 / wj g. v 13.137.4996 (2007).

[0337] 99. Quah, S. Y., Bergenholtz, G. & Tan, K. S. Fusobacterium nucleatum induces cytokine production through Toll-like-receptor-independent mechanism. Int EndodJ , 550-559, doi: 10.1111 / iej.12185 (2014).

[0338] 100. Dohlman, A. B. et al. A pan-cancer mycobiome analysis reveals fungal involvement in gastrointestinal and lung tumors. Cell 185, 3807-3822 e3812, doi : 10. 1016 / j . cell.2022.09.015 (2022).

[0339] 101. Narunsky-Haziza, L. et al. Pan-cancer analyses reveal cancer-type- specific fungal ecologies and bacteriome interactions. Cell 185, 3789-3806 e3717, doi : 10. 1016 / j . cell .2022.09.005 (2022).

[0340] 102. Sayed, I. M., Tindle, C., Fonseca, A. G., Ghosh, P. & Das, S.Functional assays with human patient-derived enteroid monolayers to assess the human gut barrier. STARProtoc 2. 100680, doi: 10. 1016 / j. xpro.2021.100680 (2021).

[0341] 103. Schneider, C. A., Rasband, W. S. & Eliceiri, K. W. NIH Image toImageJ: 25 years of image analysis. Nat Methods 9, 671-675, doi:10.1038 / nmeth.2089 (2012).

[0342] 104. Barrett, T. et al. NCBI GEO: mining tens of millions of expression profiles-database and tools update. Nucleic Acids Res 35, D760-765, doi: 10. 1093 / nar / gkl887 (2007).

[0343] 105. Barrett, T. et al. NCBI GEO: archive for functional genomics data sets--update. Nucleic Acids Res 41, D991-995, doi: 10.1093 / nar / gksl l93 (2013).

[0344] 106. Edgar, R., Domrachev, M. & Lash, A. E. Gene Expression Omnibus:NCBI gene expression and hybridization array data repository. Nucleic Acids Res 30, 207- 210, doi: 10.1093 / nar / 30.1.207 (2002).

[0345] 107. Irizarry, R. A. et al. Summaries of Affymetrix GeneChip probe level data. Nucleic Acids Res 31. el5, doi: 10. 1093 / nar / gng015 (2003).

[0346] 108. Irizarry, R. A. et al. Exploration, normalization, and summaries of high density oligonucleotide array probe level data. Biostatistics 4, 249-264, doi: 10.1093 / biostatistics / 4.2.249 (2003).

[0347] 109. Li, B. & Dewey. C. N. RSEM: accurate transcript quantification fromRNA-Seq data with or without a reference genome. BMC Bioinformatics 12, 323, doi: 10. 1186 / 1471-2105-12-323 (2011).

[0348] 110. Mortazavi, A., Williams, B. A., McCue, K., Schaeffer, L. & Wold, B.Mapping and quantifying mammalian transcriptomes by RNA-Seq. Nat Methods 5, 621- 628. doi : 10. 1038 / nmeth.1226 (2008).

[0349] 1 1 1. Trapnell, C. et al. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat Protoc 7, 562-578, doi : 10. 1038 / nprot.2012.016 (2012).

[0350] 112. Trapnell, C., Pachter, L. & Salzberg, S. L. TopHat: discovering splice junctions with RNA-Seq. Bioinformatics 25, 1105-1111, doi: 10. 1093 / bioinformatics / btpl20 (2009).

[0351] 113. Li, B., Ruotti, V., Stewart, R. M., Thomson, J. A. & Dewey, C. N.RNA-Seq gene expression estimation with read mapping uncertainty. Bioinformatics 26, 493-500, doi: 10.1093 / bioinformatics / btp692 (2010).

[0352] 114. Wagner, G. P., Kin, K. & Lynch, V. J. Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples. Theory Biosci 131, 281-285, doi: 10.1007 / sl2064-012-0162-3 (2012).

[0353] 115. Law, C. W. et al. RNA-seq analysis is easy as 1-2-3 with limma,Glimma and edgeR. FlOOORes 5, doi: 10.12688 / fl000research.9005.3 (2016).

[0354] 116. Robinson, M. D., McCarthy, D. J. & Smyth, G. K. edgeR: aBioconductor package for differential expression analysis of digital gene expression data. Bioinformatics 26, 139-140, doi: 10.1093 / bioinformatics / btp616 (2010).

[0355] 117. Sahoo, D., Dill, D. L., Tibshirani, R. & Plevritis, S. K. Extracting binary signals from microarray time-course data. Nucleic Acids Res 35, 3705-3712, doi: 10. 1093 / nar / gkm284 (2007).

[0356] 118. Sahoo, D., Dill, D. L., Gentles, A. J., Tibshirani, R. & Plevritis, S. K.Boolean implication networks derived from large scale, whole genome microarray datasets. Genome Biol 9, R157, doi: 10. 1 186 / gb-2008-9-10-rl57 (2008).

[0357] 119. Fabregat, A. et al. The Reactome Pathway Knowledgebase. NucleicAcids Res 46, D649-D655, doi: 10.1093 / nar / gkxl l32 (2018).

[0358] 120. Sanchez-Lopez. E. et al. Targeting colorectal cancer via its microenvironment by inhibiting IGF-1 receptor-insulin receptor substrate and STAT3 signaling. Oncogene 35, 2634-2644, doi: 10.1038 / onc.2015.326 (2016).

[0359] 121. Mahe, M. M., Sundaram, N., Watson, C. L., Shroyer, N. F. &Helmrath. M. A. Establishment of human epithelial enteroids and colonoids from whole tissue and biopsy. J Pis Exp, doi: 10.3791 / 52483 (2015).

[0360] 122. Sato, T. et al. Single Lgr5 stem cells build crypt-villus structures in vitro without a mesenchymal niche. Nature 459, 262-265, doi: 10.1038 / nature07935 (2009).

[0361] 123. Miyoshi. H. & Stappenbeck. T. S. In vitro expansion and genetic modification of gastrointestinal stem cells in spheroid culture. Nat Protoc 8, 2471-2482, doi : 10. 1038 / nprot.2013.153 (2013).

[0362] 124. Sayed, I. M. et al. The DNA Glycosylase NEIL2 SuppressesFusobacterium-Infection-Induced Inflammation and DNA Damage in Colonic Epithelial Cells. Cells 9. doi: 10.3390 / cells9091980 (2020).

Claims

1. CLAIMSWhat is claimed is:

1. A method of treating colorectal cancer comprising: isolating a sample from a patient, screening the sample for microbe-associated colorectal signatures (MACS), identifying an effective treatment based on the MACS, and treating the patient with the identified effective treatment.

2. The method of claim 1. wherein the MACS comprises the increased expression of at least one gene.

3. The method of claim 1, wherein the MACS comprises the increased expression of a gene selected from the group consisting of Claudin-2, Lgr-5, CEMIP, and IL-8.

4. The method of claim 3. wherein the MACS comprises the increased expression of all of Claudin-2, Lgr-5, CEMIP, and IL-8.

5. The method of claim 1, wherein the sample is comprised of stem cells isolated from colonic crypts in a human patient.

6. The method of claim 1, wherein the sample is obtained from the blood of a human patient.

7. The method of claim 1, wherein the human patient has a CRC -related cancer predisposition syndrome.

8. The method of claim 1, wherein the treatment comprises the administration of an effective amount of a therapeutic compound.

9. The method of claim 1, additionally comprising the isolation of colonic tissue from a human subject and preparation of an organoid from said tissue.

10. The method of claim 9, additionally comprising administering a therapeutic compound to the organoid.

11. The method of claim 10, additionally comprising monitoring the impact of the therapeutic compound on the MACS.

12. A composition comprising a patient-derived organoid, wherein the organoid exhibits one or more MACS.

13. The composition of claim 12, wherein the MACS comprises increased expression of at least one gene selected from the group consisting of Claudin-2, Lgr-5, CEMIP, and IL-8.

14. The composition of claim 12, wherein the organoid is derived from the colonic tissue of a patient.

15. The composition of claim 14, wherein the patient has a CRC-related cancer predisposition syndrome.

16. The composition of claim 12, additionally comprising a therapeutic compound.

17. A method of preparation of an organoid, comprising: isolation of a colonic tissue sample from a human patient, digestion of the tissue sample, and culturing the digested tissue sample in stem-cell enriched conditioned media.

18. The method of claim 17, wherein the digestion is with collagenase type I, and wherein the stem-cell enriched conditioned media contains WNT 3a, R-spondin and Noggin.

19. The method of claim 17, additionally comprising the administration of a therapeutic compound.

20. The method of claim 17, wherein the colonic tissue sample comprises stem cells isolated from colonic crypts.

Citation Information

Patent Citations

  • Immune cell organoid co-cultures

    WO2019122388A1