Chromatin landscapes and cell-free DNA fragmentation for colorectal cancer detection
Genome-wide chromatin accessibility profiling and cfDNA fragmentation analysis provide a precise diagnostic and therapeutic approach for colorectal cancer by identifying tissue-specific signatures and distinct fragmentation profiles, addressing limitations of previous methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- JOHNS HOPKINS UNIVERSITY
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
AI Technical Summary
Existing methods for analyzing chromatin accessibility in colorectal cancer development are limited by complex cell type composition and low tumor purity, hampering the identification of regions of accessibility and progression to malignancy.
A method involving genome-wide tissue-specific chromatin accessibility profiling through ATAC-seq, followed by cfDNA fragmentation analysis, to diagnose and treat colorectal cancer by identifying tissue-specific chromatin accessibility signatures and correlating changes in cfDNA fragmentation profiles.
Enables accurate diagnosis and targeted treatment of colorectal cancer by identifying distinct chromatin accessibility and cfDNA fragmentation patterns, overcoming limitations of previous methods.
Smart Images

Figure US2025055637_21052026_PF_FP_ABST
Abstract
Description
DOCKET: 348358.19102CHROMATIN LANDSCAPES AND CELL-FREE DNA FRAGMENTATION FOR COLORECTAL CANCER DETECTIONCROSS-REFERENCE TO RELATED APPLICATIONThe present application claims the benefit of U.S. provisional application no. 63 / 720,509 filed November 14, 2024, which is incorporated by reference herein in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0001] This invention was made with Government support under grant nos. CA121113, CA006973, CA233259, CA062924 and CA271896 awarded by the National Institutes of Health. The Government has certain rights in the invention.BACKGROUND
[0002] Colorectal cancer is a leading cause of cancer death in men and women, with a nearly 1 million deaths and more than 1.9 million new cases diagnosed worldwide in 20221. It has been well established that colorectal tumors originate from precursor cells at the base of crypts in the colonic epithelium, which progress to CRC precursor lesions, namely, adenomas and eventually to cancer4. The somatic genetic, epigenetic alterations and genomic instability that lead to this progression have been well characterized, and such findings have led to actionable diagnostic and therapeutic intervention in patient care3 1 1. These driver alterations dysregulate the function of molecular pathways in the affected cells, often by disrupting genomic regulation encoded in the chromatin accessibility landscape8 12. However, the nature and extent of this chromatin dysregulation has yet to be fully characterized.
[0003] Chromatin accessibility is a key regulator of gene transcription, and is essential for the maintenance of cell state and cell identity13,14. Previous efforts to analyze the accessibility landscape of colorectal cancer development have used the assay for transposase accessible chromatin by sequencing (ATAC-seq) or this approach adapted to single cell analyses12 15 lx. Direct comparison of accessibility profiles from colorectal adenomas and cancers have indicated that genome-wide changes in accessibility occur at the onset of malignancy17, but these studies have been hampered by the complex cell type composition of colonic tissue and low tumor purity of samples analyzed. Conversely, single cell ATAC-seq179217191.1DOCKET: 348358.19102analyses of colon tissue has demonstrated genome-wide shifts in chromatin accessibility and transcription over the course of development to malignancy18, but this approach analyzed a small number of samples and therefore may be limited in the identification of the compendium of regions of accessibility using single-cell analyses19.SUMMARY
[0004] In one aspect, we now demonstrate that changes in chromatin remodeling is a feature of colorectal tumorigenesis and identify chromatin accessibility through evaluation of genome-wide cfDNA fragmentation.
[0005] Accordingly, in certain aspects, a method of diagnosing and treating a subject diagnosed with cancer is provided, comprising:a) producing a genome-wide tissue specific chromatin accessibility profile from a sample obtained from a subject;b) comparing the genome-wide tissue specific chromatin accessibility profile of the subject to a normal healthy control sample; andc) identifying changes in the genome-wide tissue specific chromatin accessibility profile of the subject and correlating these changes to a genome- wide tissue specific cancer chromatin accessibility profiles, thereby diagnosing the subject with cancer.
[0006] The diagnosed subject is then suitably treated with a cancer-specific therapy.
[0007] In certain embodiments, the genome-wide tissue specific chromatin accessibility profile comprises identifying genomic loci comprising tissue-specific signatures of chromatin accessibility. In certain embodiments, the genetic loci are identified by conducting assays for transposase accessible chromatin by sequencing (ATAC-seq) thereby generating libraries comprising peaks across the sample. In certain embodiments, the ATAC-seq generated libraries exclude libraries with low complexity, sequencing quality, signal-to-noise ratio, or replicate concordance.
[0008] In certain embodiments, overlap of each location of peaks across samples identify accessible consensus peaks. In certain embodiments, the accessible consensus peaks obtained from the subject’s samples are compared to accessible consensus peaks obtained from cancer samples and normal samples. In certain embodiments, the accessible consensus peaks are diagnostic of the type of cancer. In certain embodiments, the accessible consensus peaks among2179217191.1DOCKET: 348358.19102samples within the same tissue type or disease state comprise a higher correlation or similarity. In certain embodiments, the accessible consensus peaks among normal samples within the same tissue type comprise a higher correlation or similarity. In certain embodiments, the accessible consensus peaks among normal samples are inaccessible in healthy samples as compared to the accessible consensus peaks among cancer samples.
[0009] In certain embodiments, identifying genes and pathways regulated by chromatin changes comprises mapping accessible consensus peaks to nearby genes and analyzed these for enrichment of gene sets or pathways. In certain embodiments, the subject is diagnosed with colorectal cancer.
[0010] In certain embodiments, the method further comprises producing a cell-free DNA (cfDNA) fragmentation profde at tissue specific chromatin accessibility regions in patients with and without cancer.
[0011] In certain embodiments, the cfDNA fragmentation profiles comprise cfDNA fragmentation profiles of consensus peaks. In certain embodiments, evaluation of cfDNA features to produce a cfDNA fragmentation profile comprises evaluating sequence coverage, fragment size, and frequency of fragments ending at specific positions normalized by sequence coverage.
[0012] In certain embodiments, higher chromatin accessibility at a center of consensus peaks comprises decreased cfDNA sequence coverage, increased fragment end frequency, a reduction in size of cfDNA fragments or combinations thereof. In certain embodiments, cfDNA fragmentation profiles of healthy subjects are similar. In certain embodiments, cfDNA fragmentation profiles of subjects with cancer are similar based on cancer type. In certain embodiments, a subject with colorectal cancer comprises a cfDNA fragmentation profile distinct to other cancer types.
[0013] In another aspect, a method of diagnosing and treating colorectal cancer in a subject is provided, comprising:a) producing a cell-free DNA (cfDNA) fragmentation profile at tissue specific chromatin accessibility regions in patients with and without cancer, wherein the subject with colorectal cancer comprises a cfDNA fragmentation profile distinct to other cancer types; and b) treating the subject with colorectal specific therapeutics.3179217191.1DOCKET: 348358.19102
[0014] In certain embodiments, the cfDNA fragmentation profiles comprise cfDNA fragmentation profiles of consensus peaks. In certain embodiments, evaluation of cfDNA features to produce a cfDNA fragmentation profile comprises evaluating sequence coverage, fragment size, and frequency of fragments ending at specific positions normalized by sequence coverage. In certain embodiments, higher chromatin accessibility at a center of consensus peaks comprises decreased cfDNA sequence coverage, increased fragment end frequency, a reduction in size of cfDNA fragments or combinations thereof. In certain embodiments, cfDNA fragmentation profiles of healthy subjects are similar. In certain embodiments, cfDNA fragmentation profiles of subjects with cancer are similar based on cancer type.
[0015] DEFINITIONS
[0016] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0017] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”
[0018] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value or range. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude within 5-fold, and also within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the4179217191.1DOCKET: 348358.19102term “about” meaning within an acceptable error range for the particular value should be assumed.
[0019] The terms “aligned”, “alignment”, “mapped” or “aligning”, “mapping” refer to one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Such alignment can be done manually or by a computer algorithm, examples including the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysts pipeline. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).
[0020] The term “cancer” as used herein is meant, a disease, condition, trait, genotype or phenotype characterized by unregulated cell growth or replication as is known in the art; including colorectal cancer (CRC), liver cancer (including hepatocellular carcinoma (HCC)), lung cancer (including non-small cell lung carcinoma), gastric cancer, colorectal cancer, as well as, for example, leukemias, e.g., acute myelogenous leukemia (AML), chronic myelogenous leukemia (CML), acute lymphocytic leukemia (ALL), and chronic lymphocytic leukemia, AIDS related cancers such as Kaposi's sarcoma; breast cancers; bone cancers such as Osteosarcoma, Chondrosarcomas, Ewing's sarcoma, Fibrosarcomas, Giant cell tumors, Adamantinomas, and Chordomas; Brain cancers such as Meningiomas, Glioblastomas, Lower- Grade Astrocytomas, Oligodendrocytomas, Pituitary Tumors, Schwannomas, and Metastatic brain cancers; cancers of the head and neck including various lymphomas such as mantle cell lymphoma, non-Hodgkins lymphoma, adenoma, squamous cell carcinoma, laryngeal carcinoma, gallbladder and bile duct cancers, cancers of the retina such as retinoblastoma, cancers of the esophagus, gastric cancers, multiple myeloma, ovarian cancer, uterine cancer, thyroid cancer, testicular cancer, endometrial cancer, melanoma, bladder cancer, prostate cancer, pancreatic cancer, sarcomas, Wilms' tumor, cervical cancer, head and neck cancer, skin cancers, nasopharyngeal carcinoma, liposarcoma, epithelial carcinoma, renal cell carcinoma, gallbladder adeno carcinoma, parotid adenocarcinoma, endometrial sarcoma, multidrug resistant cancers; and proliferative diseases and conditions, such as neovascularization associated with tumor angiogenesis.
[0021] The term “cell free nucleic acid,” “cell free DNA,” or “cfDNA” refers to nucleic acid fragments that circulate in an individual's body (e.g., bloodstream) and originate from one or5179217191.1DOCKET: 348358.19102more healthy cells and / or from one or more cancer cells. Additionally, cfDNA may come from other sources such as viruses, fetuses, etc.
[0022] The term “cfDNA sequence coverage” refers to the average number of cfDNA molecules overlapping a specific position.
[0023] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments that originate from tumor cells or other types of cancer cells, which may be released into an individual's bloodstream as result of biological processes such as apoptosis or necrosis of dying cells or actively released by viable tumor cells.
[0024] As used herein, the terms “comprising,” “comprise” or “comprised,” and variations thereof, in reference to defined or described elements of an item, composition, apparatus, method, process, system, etc. are meant to be inclusive or open ended, permitting additional elements, thereby indicating that the defined or described item, composition, apparatus, method, process, system, etc. includes those specified elements— or, as appropriate, equivalents thereof-and that other elements can be included and still fall within the scope / definition of the defined item, composition, apparatus, method, process, system, etc.
[0025] “Diagnostic” or “diagnosed” means identifying the presence or nature of a pathologic condition. Diagnostic methods differ in their sensitivity and specificity. The “sensitivity” of a diagnostic assay is the percentage of diseased individuals who test positive (percent of “true positives”). Diseased individuals not detected by the assay are “false negatives.” Subjects who are not diseased and who test negative in the assay, are termed “true negatives.” The “specificity” of a diagnostic assay is 1 minus the false positive rate, where the “false positive” rate is defined as the proportion of those without the disease who test positive. While a particular diagnostic method may not provide a definitive diagnosis of a condition, it suffices if the method provides a positive indication that aids in diagnosis.
[0026] An “effective amount” as used herein, means an amount which provides a therapeutic or prophylactic benefit.
[0027] As used herein, the terms “fragmentation profile,” “fragmentome profile”, are equivalent and can be used interchangeably. In some embodiments, determining a cfDNA fragmentation profile in a mammal can be used for identifying a mammal as having cancer. For example, cfDNA fragments obtained from a mammal (e.g., from a sample obtained from a6179217191.1DOCKET: 348358.19102mammal) can be subjected to low coverage whole- genome sequencing, and the sequenced fragments can be mapped to the genome (e.g., in non- overlapping windows) and assessed to determine a cfDNA fragmentation profile. As described herein, a cfDNA fragmentation profile of a mammal having cancer is more heterogeneous (e.g., in chromatin accessibility) than a cfDNA fragmentation profile of a healthy mammal (e.g., a mammal not having cancer). As such, this disclosure also provides methods and materials for assessing, monitoring, and / or treating mammals (e.g., humans) having, or suspected of having, cancer. In some embodiments, this document provides methods and materials for identifying a mammal as having cancer. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine the presence and, optionally, the tissue of origin of the cancer in the mammal based, at least in part, on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials for monitoring a mammal as having cancer are provided. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine the presence of the cancer in the mammal based, at least in part, on the cfDNA fragmentation profile of the mammal. In some embodiments, methods and materials for identifying a mammal as having cancer and administering one or more cancer treatments to the mammal to treat the mammal are provided. For example, a sample (e.g., a blood sample) obtained from a mammal can be assessed to determine if the mammal has cancer based, at least in part, on the cfDNA fragmentation profile of the mammal, and one or more cancer treatments can be administered to the mammal.
[0028] The term “genomic nucleic acid,” or “genomic DNA,” refers to nucleic acid including chromosomal DNA that originates from one or more healthy (e.g., non-tumor) cells or tumor cells. In various embodiments, genomic DNA can be extracted from a cell derived from a blood cell lineage, such as a white blood cell (WBC).
[0029] “Optional” or “optionally” means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0030] As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.7179217191.1DOCKET: 348358.19102
[0031] ‘Parenteral” administration of an immunogenic composition includes, e g., subcutaneous (s.c.), intravenous (i.v.), intramuscular (i.m.), or intrastemal injection, or infusion techniques.
[0032] The terms “patient” or “individual” or “subject” are used interchangeably herein, and refers to a mammalian subject to be treated, with human patients being preferred. In some embodiments, the methods of the invention find use in experimental animals, in veterinary application, and in the development of animal models for disease, including, but not limited to, rodents including mice, rats, and hamsters, and primates.
[0033] The term “reference genome” as used herein may refer to a digital or previously identified nucleic acid sequence database, assembled as a representative example of a species or subject. Reference genomes may be assembled from the nucleic acid sequences from multiple subjects, sample or organisms and does not necessarily represent the nucleic acid makeup of a single person. Reference genomes may be used to for mapping of sequencing reads from a sample to chromosomal positions. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov.
[0034] The term “read segment” or “read” refers to any nucleotide sequences including sequence reads obtained from an individual and / or nucleotide sequences derived from the initial sequence read from a sample obtained from an individual.
[0035] The terms “sample,” “patient sample,” “biological sample,” and the like, encompass a variety of sample types obtained from a patient, individual, or subject and can be used in a diagnostic, prognostic and / or monitoring assay. The patient sample may be obtained from a healthy subject, a diseased patient, or a patient with lung cancer. In certain embodiments, a sample that is “provided” can be obtained by the person (or machine) conducting the assay, or it can have been obtained by another, and transferred to the person (or machine) carrying out the assay. Moreover, a sample obtained from a patient can be divided and only a portion may be used for diagnosis. Further, the sample, or a portion thereof, can be stored under conditions to maintain sample for later analysis. The definition specifically encompasses blood and other liquid samples of biological origin (including, but not limited to, peripheral blood, serum, plasma, cord blood, amniotic fluid, cerebrospinal fluid, urine, saliva, stool and synovial fluid),8179217191.1DOCKET: 348358.19102solid tissue samples such as a biopsy specimen or tissue cultures or cells derived therefrom and the progeny thereof. In certain embodiment, a sample comprises cerebrospinal fluid. In a specific embodiment, a sample comprises a blood sample. In another embodiment, a sample comprises a plasma sample. In yet another embodiment, a serum sample is used. The definition of “sample” also includes samples that have been manipulated in any way after their procurement, such as by centrifugation, filtration, precipitation, dialysis, chromatography, treatment with reagents, washed, or enriched for certain cell populations. The terms further encompass a clinical sample, and also include cells in culture, cell supernatants, tissue samples, organs, and the like. Samples may also comprise fresh-frozen and / or formalin-fixed, paraffin-embedded tissue blocks, such as blocks prepared from clinical or pathological biopsies, prepared for pathological analysis or study by immunohistochemistry.
[0036] The term “sequence reads” refers to nucleotide sequences read from a sample obtained from an individual. Sequence reads can be obtained through various methods known in the art.
[0037] As defined herein, a “therapeutically effective” amount of a compound or agent (i.e., an effective dosage) means an amount sufficient to produce a therapeutically (e.g., clinically) desirable result. The compositions can be administered from one or more times per day to one or more times per week, including once every other day. The skilled artisan will appreciate that certain factors can influence the dosage and timing required to effectively treat a subject, including but not limited to the severity of the disease or disorder, previous treatments, the general health and / or age of the subject, and other diseases present. Moreover, treatment of a subject with a therapeutically effective amount of the compounds of the invention can include a single treatment or a series of treatments.
[0038] As used herein, the terms “treat,” treating,” “treatment,” and the like refer to reducing or ameliorating a disorder and / or symptoms associated therewith. It will be appreciated that, although not precluded, treating a disorder or condition does not require that the disorder, condition or symptoms associated therewith be completely eliminated.
[0039] Genes: All genes, gene names, and gene products disclosed herein are intended to correspond to homologs from any species for which the compositions and methods disclosed herein are applicable. It is understood that when a gene or gene product from a particular species9179217191.1DOCKET: 348358.19102is disclosed, this disclosure is intended to be exemplary only, and is not to be interpreted as a limitation unless the context in which it appears clearly indicates. Thus, for example, for the genes or gene products disclosed herein, are intended to encompass homologous and / or orthologous genes and gene products from other species.
[0040] Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.
[0041] Any compositions or methods provided herein can be combined with one or more of any of the other compositions and methods provided herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee.
[0043] FIG. 1 is an overview of organoid and cell-free DNA analyses. Chromatin accessibility analyses of were performed using patient-derived colorectal healthy, adenoma and carcinoma tissue samples as well as healthy peripheral blood mononuclear cells (PBMCs). Chromatin accessibility changes were assessed over the course of colorectal cancer development, identifying regions that were dynamically regulated. Analysis of accessibility peaks were used to identify transcription factor drivers regulating colorectal cancer development and features underlying cell-free fragmentation in healthy individuals as well as those with colorectal cancer.
[0044] FIGS. 2A, 2B are a series of plots and a heatmap demonstrating that chromatin accessibility features of colorectal cancer organoids reveal tissue specific peaks. FIG.2A: Median tissue-specific ATAC-seq tracks (top) generated from PBMCs (n=17) and colonic organoids from10179217191.1DOCKET: 348358.19102healthy (n=l 1), adenoma (n=41), and cancer (n=26) samples (bottom) at the CDK14 gene. FIG.2B: Heatmap of normalized ATAC-seq signal intensity at consensus peaks formed from resolving peak overlaps in all organoid samples reveals differences between healthy, adenoma and cancer samples.
[0045] FIGS. 3A-3F are a series of plots and heatmaps demonstrating that chromatin accessibility analyses identified signatures of colorectal cancer progression. FIG. 3A: Principal component analyses (PCA) of accessibility data from 305 ATAC-seq analyses at 90,028 consensus peaks. FIG. 3B: PCA of accessibility data from 78 colorectal organoid ATAC-seq analyses at 62,890 consensus peaks. FIGS. 3C, 3D: Heatmap of the ATAC-seq peaks cumulatively contributing to 90% of the total variation of PCI (FIG. 3C) and PC2 (FIG. 3D) defined in FIG. 3B identifies transcription factor binding elements that are enriched in peaks that are adenoma or cancer accessible or inaccessible. Samples are ordered by position in the relevant PC coordinate, and peaks are ordered by the magnitude of their contribution to that PC coordinate. FIGS.3E, 3F: Hypergeometric enrichments and 95% confidence intervals for genes within 2 Kbp of selected peaks from PCI (FIG.3E) and PC2 (FIG.3F), separated by sign (+ or -) of contribution to the principal component reveal gene pathways involved in different aspects of colorectal adenoma or cancer development.
[0046] FIGS. 4A-4C are a series of plots and heatmaps demonstrating that chromatin accessibility signatures identify regions of altered fragmentation in cfDNA. FIG. 4A: Median signal of 17 PBMC ATAC-seq analyses in 71,270 consensus peaks (±2.5 Kbp) from PBMC and colon organoid samples. All peaks are centered at genomic position 0 and ranked from top to bottom by ATAC-seq signal intensity in PBMC samples. FIG. 4B: Cell-free DNA fragmentation metrics, including overlapping fragment coverage, frequency of fragment ends, and size of cell-free DNA fragments for pooled cfDNA data from 251 healthy individuals ordered at the same loci shown in FIG. 3 A. FIG. 4C: Boxplot of fragmentation metrics for 251 healthy cfDNA samples. Comparing mean values of cfDNA fragmentation metrics at the center (position 0 - 25bp) of the 10,000 loci which are most accessible in PBMCs compared to those of the 10,000 loci which are the least accessible revealed decrease in fragment coverage and fragment ends, and an increase in fragment length at the most accessible loci. The center line in the boxplots represents the median, the upper limit of the boxplots represents the third quantile (75th percentile), the lower limit of the11179217191.1DOCKET: 348358.19102boxplots represents the first quantile (25th percentile), the upper whiskers is the maximum value of the data that is within 1.5 times the interquartile range over the 75th percentile, and the lower whisker is the minimum value of the data that is within 1.5 times the interquartile range under the 25th percentile.
[0047] FIGS. 5A-5D are a series of plots showing altered cfDNA fragmentation at regions of colon-specific chromatin accessibility in patients with colorectal cancer. FIG. 5A: Volcano plot showing selection of 23,911 peaks with enriched accessibility in colon organoids compared to PBMC. FIG. 5B: Cohort-level summary of cfDNA metrics (mean ± 1 SD) at colon-enriched peaks comparing healthy and cancer cohorts, faceted by genomic context reveal altered fragment coverage, ends and length characteristics in patients with cancer compared to healthy individuals.FIG. 5C: Receiver operating characteristic curves using fragment coverage at distal elements position 0 (±10bp) distinguishes healthy individuals from those with colon cancer with high performance. FIG. 5D: Normalized fragment ends at distal elements position 0 (±10bp) predicts cfDNA mutant allele fraction obtained independently using digital droplet PCR.
[0048] FIGS. 6A-6E show a series of plots, schematics and a heatmap demonstrating the chromatin accessibility changes during colorectal cancer development which reveal adenoma and cancer specific chromatin accessibility states. FIG. 6A: Heatmap of normalized accessibility values for all peaks in selected comparisons. FIG. 6B: Genome browser visualization of median aggregated ATAC-seq profiles at the adenoma-enriched APCDD1 / VAPA locus, highlighted with six adenoma-specific peaks and one adenoma / cancer shared peak. FIG.6C: Scatterplot of healthy, adenoma, and cancer-specific peaks showing the relationship between distance from transcription start sites and average change in accessibility, with loess regression trendlines and 95% confidence intervals. FIG. 6D: Cohort-level summary of cfDNA coverage (mean ± 1 SD) at differentially accessible distal elements in 37 healthy individuals (MAF = 0%) and 7 individuals with cancer (MAF > 50%) identifying different fragment coverage for individuals with cancer at colorectal cancer specific peaks compared to adenoma-specific peaks. FIG. 6E: Proposed model of chromatin accessibility regulation throughout colon cancer development, revealing specific chromatin accessibility states for adenomas and cancers.
[0049] FIGS. 7A-7C show colorectal organoid histopathology slides revealing features of normal, adenoma, and colon cancer tissues. FIGS. 7A-7C: Hematoxylin and eosin stain of12179217191.1DOCKET: 348358.19102representative healthy (FIG. 7 A), adenoma (FIG. 7B), and cancer (FIG. 7C) colorectal organoid samples used in this study. Organoids recapitulated histologic features of their primary tissue counterparts, with normal structures and goblet cells observed for normal organoids, disappearance of goblet cells and decreased differentiation for adenoma organoids, and a lack of differentiation and abnormal nucleus localization for colorectal cancer organoids.
[0050] FIGS. 8A-8F are a series of heatmaps and plots demonstrating the quality assessment of ATAC-seq analyses. FIG. 8A: Total sequencing reads for all ATAC-seq samples categorized by alignment to hgl9 reference genome, human repetitive sequence elements, and duplication status. FIG. 8B: Top-line output of FastQC for all ATAC-seq samples. FIG 8C: Quality control metrics for all ATAC-seq samples summarized by warning level (see Methods for details). FIGS. 8D-8E: Representative density plots showing the Pearson correlation of ATAC-seq peak intensity between corresponding transposition (FIG. 8D) and sequencing (FIG.8E) replicates. FIG 8F: Histogram showing the distribution of all Person correlations between corresponding replicates. FRiP, fraction of reads in peaks, NRF, nonredundant fraction of reads, PBC, PCR bottlenecking coefficient, TSS, transcription start site.
[0051] FIGS. 9A, 9B depict the correlation analyses of ATAC-seq samples which show high within-group agreement and low correlation between tissue or disease states. FIG. 9A:Heatmap showing Pearson correlation values of read counts at 61,754 colon consensus peaks between all combinations of samples passing quality control. FIG. 9B: Box and whisker plots of Pearson correlations from FIG. 9A (excluding self-comparisons), comparing ATAC-seq profiles within and between sample groups. PBMC samples are highly correlated with each either as a group, and are negatively correlated with colon samples. While colon cancers, colon adenomas or healthy colon organoids are positively correlated within each group, each disease state is negatively correlated to PBMCs or other disease states. The center line in the boxplots represents the median, the upper limit of the boxplots represents the third quantile (75th percentile), the lower limit of the boxplots represents the first quantile (25th percentile), the upper whiskers is the maximum value of the data that is within 1.5 times the interquartile range over the 75th percentile, and the lower whisker is the minimum value of the data that is within 1.5 times the interquartile range under the 25th percentile.13179217191.1DOCKET: 348358.19102
[0052] FTGS. 10A-10C show the genome-wide distribution of colon consensus peaks show enrichment at promoters and distal elements. FIG. 10A: Genomic annotation of 62,890 colon consensus peaks shown in FIG. 2B. FIG. 10B: Histogram showing the distribution of distances between each colon consensus peak to the nearest transcription start site. FIG. 10C:Circos plot showing the correspondence between gene density and consensus peak density in 5 Mbp bins tiled across the genome.
[0053] FIGS. 11A-11E show a genome-wide distribution of multi-tissue consensus peaks demonstrating enrichment at promoters and distal elements. FIG. 11A: Genomic annotation of 90,028 multi -tissue consensus peaks used to generate FIG. 3 A. FIG. 11B:Histogram showing the distribution of distances between each consensus peaks from the multitissue consensus peak set to the nearest transcription start site. FIG. 11C: Circos plot showing the correspondence between gene density and consensus peak density in 5 Mbp bins tiled across the genome. FIG 11D: Histogram showing total variance explained by each principal component. FIG HE: Principal components 2 and 3 of multi -tissue PC A from FIG. 3 A shows continued separation between tissues and disease states
[0054] FIGS. 12A, 12B demonstrate that unsupervised clustering validates chromatin accessibility signatures. FIG. 12A: PC A of accessibility data from 305 ATAC-seq analyses at 36,494 consensus peaks constructed using randomized group assignments for each sample. FIG.12B: PCA of accessibility data from 78 colorectal organoid ATAC-seq analyses at 52,011 consensus peaks constructed using randomized group assignments for each sample.
[0055] FIGS. 13A-13D demonstrate that chromatin accessibility signatures identify genes involved in colon cancer. FIG. 13A: Bar plot of the ten consensus peaks which have the strongest contribution to the definition of the first two principal components shown in FIG. 3B. FIG. 13B:Genome browser visualization of NF1 locus, with the consensus peak highlighted in pink. FIGS.13C, 13D: Bar plot of the DNA binding elements with the highest fold enrichment for their recognition sequences in PC 1 (FIG. 13C) and PC2 (FIG. 13D), as determined by Homer. Low frequency motifs are those whose recognition sequence was found in <1% of the background sequence set.
[0056] FIGS. 14A, 14B demonstrate that gene sets show similarity of gene enrichment across colon and other tissue or molecular processes. FIGS. 14A, 14B: Heatmaps of Jaccard14179217191.1DOCKET: 348358.19102similarity coefficients for all combinations of the top 25 gene sets enriched within 2Kbp of the peak sets selected for colon PCI (FIG. 14A) and PC2 (FIG. 14B), clustered hierarchically.
[0057] FIGS. 15A-15F demonstrate RNA-seq validation of chromatin signatures. FIGS.15A, 15B: Volcano plots of all genes harboring a PCI (FIG. 15A) or PC2 (FIG. 15B) chromatin signature peak. Genes are separated into those targeted by peaks with a positive or negative PC coordinate. A positive log2(Fold change) indicates the gene was more highly expressed in the cancer samples than in the adenoma samples. FIGS. 15C, 15E: Stacked bar plots corresponding to the enriched gene sets displayed in FIGS. 3G, 3H. Each bar represents the subset of genes targeted by chromatin signature peaks from PCI (FIG. 15C) or PC2 (FIG. 15E), and the color shows the proportion of those genes which are differentially expressed between colon adenoma and cancer samples. FIGS. 15D, 15F: Box and whisker plots showing the fold change difference between adenoma and cancer samples for the genes identified as differentially expressed in FIG.15C (1 D) and FIG. 15E (15G). The center line in the boxplots represents the median, the upper limit of the boxplots represents the third quantile (75th percentile), the lower limit of the boxplots represents the first quantile (25th percentile), the upper whiskers is the maximum value of the data that is within 1.5 times the interquartile range over the 75th percentile, and the lower whisker is the minimum value of the data that is within 1.5 times the interquartile range under the 25th percentile.
[0058] FIGS. 16A-16C demonstrate that genome-wide distribution of colon / PBMC consensus peaks shows similar enrichment at promoters and distal elements. FIG. 16A: Genomic annotation of 71,270 colon / PBMC consensus peaks shown in FIGS. 4A, 4B. FIG. 16B: Histogram showing the distribution of distances between each consensus peak to the nearest transcription start site. FIG. 16C: Circos plot showing the correspondence between gene density and consensus peak density in 5 Mbp bins tiled across the genome.
[0059] FIGS. 17A-17C demonstrate that randomly selected genomic loci show no changes in cfDNA fragment characteristics. FIG. 17A: Mean ATAC-seq of PBMC samples at 60,000 randomly selected regions, ordered by strength of accessibility at position 0. FIG 17B: cfDNA fragmentation metrics at randomly selected regions shown in FIG. 17A. FIG. 17C: Boxplot of fragmentation metrics for 251 healthy cfDNA samples. Comparing mean values of cfDNA fragmentation metrics at the center (position 0 - 25bp) of the 10,000 loci which are most accessible15179217191.1DOCKET: 348358.19102in PBMCs compared to those of the 10,000 loci which are the least accessible. The center line in the boxplots represents the median, the upper limit of the boxplots represents the third quantile (75th percentile), the lower limit of the boxplots represents the first quantile (25th percentile), the upper whiskers is the maximum value of the data that is within 1.5 times the interquartile range over the 75th percentile, and the lower whisker is the minimum value of the data that is within 1.5 times the interquartile range under the 25th percentile.
[0060] FIG. 18 demonstrates the differences in fragmentation features at distal element accessibility peaks for healthy individuals and cancer patients. Scatterplot comparing fragmentation features at colon-specific peak centers (±25 bp) to locus flanks (1-2.5 Kbp from peak center) for 251 healthy individuals, and 51 individuals with cancer.
[0061] FIGS. 19A-19E demonstrate that the genomic context influences cfDNA fragmentation feature differences at colon-specific peaks in healthy individuals and patients with cancer. FIG. 19A: Receiver operating characteristic curves for identifying individuals with cancer based on each cfDNA fragmentation metric, at each genomic context. FIG. 19B: Heatmap of normalized cfDNA fragmentation metric values for all cfDNA samples, faceted by genomic context. FIGS. 19C, 19D: Heatmap of Pearson correlations of cfDNA fragmentation metrics at each genomic context for all plasma samples from healthy individuals (FIG. 19C) and all those from individuals with cancer (FIG. 19E). FIG. 19E: Regression plots comparing fragmentation metrics at colon-specific peak centers (-10bp) to mutant allele fraction of 51 individuals with cancer, faceted by genomic context. The shaded area is the 95% confidence interval of the regression line.
[0062] FIGS. 20A-20D demonstrate that chromatin accessibility changes during colorectal cancer development reveal colon healthy, adenoma, and cancer specific chromatin accessibility states. FIG. 20A: Heatmap summarizing 12,718 differentially accessible loci (fold change > 2 and FDR < 0.05) between all combinations of colon healthy, colon adenoma, and colon cancer ATAC-seq sample groups. 521 loci were differentially accessible in two comparisons, resulting in 13,239 distinct significant differences between sample groups. FIG. 20B: Heatmap summarizing number of differential peaks (fold change > 2 and FDR < 0.05) between all combinations of colon healthy, colon adenoma, colon cancer and PBMC ATAC-seq sample groups. Marked comparisons are investigated in subsequent analyses. FIG. 20C: Bar plot of all16179217191.1DOCKET: 348358.19102values in top left quadrant of b (outlined in red). Based on the direction of accessibility change in the colon-colon comparisons, a peak was labeled as either Healthy specific, Heal thy / Adenoma shared, Adenoma specific, Adenoma / Cancer shared, Cancer specific, or Healthy / Cancer shared.FIG. 20D: Heatmap of normalized accessibility values for all peaks in selected comparisons.
[0063] FIGS. 21A-21C demonstrate that accessibility peaks cluster near putative regulators of colorectal cancer development. FIGS. 21A, 21B, 21C: Genome browser images of three loci with the highest concentration per gene of adenoma specific peaks, cancer specific, and / or adenoma and cancer shared peaks.
[0064] FIGS. 22A-22B show the genomic features of colorectal organoids. FIG. 22A:Heatmap of accessibility in all organoids at colorectal chromatin peaks which have differential accessibility in at least one phase of CRC development and are located near cancer driver or chromatin modifier gene sets. FIG. 22B: Summary of mutations identified in cancer driver and chromatin modifier genes through whole genome sequencing of colorectal adenoma (n=23) and cancer (n=3) organoids.
[0065] FIGS. 23A-23D show the chromatin accessibility dynamics of colorectal cancer development. FIG. 23A: Enrichment of transcription factor binding sequences within top 2,000 peak in differential accessibility sets. FIG. 23B: Proposed model of chromatin accessibility regulation throughout colon cancer development. FIG. 23C-23D: Locus visualization of median aggregated ATAC-seq profiles at the (FIG.23C) KLF5 / KLF12 and (FIG. 23D) BICC1 loci, with pathophysiological phase-specific chromatin peaks highlighted in pink.
[0066] FIGS. 24A-24F demonstrate the detection of colorectal cancer chromatin accessibility in cfDNA. FIG.24A: Volcano plot showing selection of 24,036 peaks with enriched accessibility in colon organoids compared to PBMC. FIG. 24B Summary of cfDNA fragmentation metrics (mean ± 1 SD) at colon-enriched peaks in the CAIRO5 cohort of 47 samples from patients with detectable cancer (RAS / BRAF ddPCR mutant allele fraction (MAF) > 0%) and 25 samples from patients without detectable cancer (RAS / BRAF ddPCR MAF = 0%), faceted by genomic context. FIG. 24C: Receiver operating characteristic curves using fragment coverage at distal elements position 0 (±10bp) distinguished patient samples with or without detectable ctDNA with high performance. FIG. 24D: Normalized cfDNA fragment end measurements from distal element genomic positions (0 ± lObp) predicted ddPCR MAF. FIG. 24E: Summary of cfDNA17179217191.1DOCKET: 348358.19102coverage (mean ± 1 SD) at differentially accessible distal elements identified altered fragment coverage at colorectal cancer specific peaks compared to adenoma-specific peaks. FIG. 24F:(Left) Summary of cfDNA coverage (mean ± 1 SD) at healthy colon distal accessibility loci and (right) Boxplots summarizing fragment coverage at healthy colon distal accessibility loci genomic position (0 ± lObp) indicated decreased normalized coverage in patients with detectable ctDNA, consistent with fragments originating from healthy colon tissue in cfDNA of patients with cancer. The center line in the boxplots represents the median, the upper limit of the boxplots represents the third quartile (75th percentile), the lower limit of the boxplots represents the first quartile (25 th percentile), the upper whiskers is the maximum value of the data that is within 1.5 times the interquartile range over the 75th percentile, and the lower whisker is the minimum value of the data that is within 1.5 times the interquartile range under the 25th percentile.
[0067] FIGS. 25A-25E show the genome-wide distribution of colon consensus peaks.FIG. 25A: Genomic annotation of 65,054 colon consensus peaks shown in FIG. 2B. FIG. 25B:Histogram showing the distribution of distances between each colon consensus peak to the nearest transcription start site. FIG. 25C: Circos plot showing the correspondence between gene density and consensus peak density in 5 Mbp bins tiled across the genome. FIG.25D: Histogram showing total variance explained by each principal component in PCA analysis of 65,054 colon consensus peaks. FIG. 25E: PCA of accessibility data from 46 colorectal organoid ATAC-seq analyses at 54,032 consensus peaks constructed using randomized group assignments for each sample.
[0068] FIGS. 26A-26E show the genome-wide distribution of multi-tissue consensus peaks. FIG. 26A: Genomic annotation of 90,028 multi-tissue consensus peaks used to generate FIG.3A. FIG. 26B: Histogram showing the distribution of distances between each consensus peaks from the multi-tissue consensus peak set to the nearest transcription start site. FIG. 26C:Circos plot showing the correspondence between gene density and consensus peak density in 5 Mbp bins tiled across the genome. FIG.26D: Histogram showing total variance explained by each principal component. FIG. 26E: Principal components 2 and 3 of multi-tissue PCA from FIG.3A.
[0069] FIGS. 27A-27C show the pathophysiological phase-specific chromatin peaks preferentially target differentially expressed genes and cancer driver genes. FIG. 27A: Volcano plot identifying 1,045 genes with differential RNA-seq expression (|Fold change] > 2, FDR<0.05)18179217191.1DOCKET: 348358.19102between colon adenoma and cancer organoids. Highlighted points are differentially expressed genes where the majority of nearby peaks are accessible in either (blue) adenoma (adenoma specific or adenoma / healthy shared peaks) or (red) cancer (cancer specific or cancer / healthy shared peaks). FIGS. 27B-27C: Line graphs showing the odds ratio and 95% confidence interval that (FIG. 27B) a gene is differentially expressed between adenoma and cancer, or (FIG. 27C) a gene is a cancer driver, given the presence and quantity of nearby pathophysiological phase-specific chromatin peaks.
[0070] FIGS. 28A-28L show the gapped k-mer analyses of ATAC-seq peaks reveal differential activity of transcription regulators. FIGS. 28A-28B: Receiving operator characteristic (FIG. 28A) and precision-recall curves (FIG. 28B) for gkm-SVM models trained on top 10,000 accessibility peaks of each colorectal organoid sample to distinguish regulatory elements from random genomic sequences. FIG. 28C: PCA of gkm-SVM weight vector resulting from training using top 10,000 peaks from 43 organoid samples. FIGS. 28D-28E: Receiver operating characteristic (FIG. 28D) and precision-recall curves (FIG. 28E) for gkm-SVM models trained using differentially accessible peaks in all 1,035 possible pairs of colorectal organoid samples.FIG. 28F: Scatterplot of the 2,000 most upregulated (red) and downregulated (blue) ATAC-seq peaks from an example cancer vs healthy comparison. FIGS. 28G-28I: Receiver operating characteristic curves showing classification performance by gkm-SVM and linear regression in cancer vs healthy (FIG. 28G), adenoma vs healthy (FIG. 28H), and adenoma vs cancer (FIG.281) samples explained by a linear regression of the seven selected transcriptions factors. FIGS.28J-28L: gkm-PWM rankings of transcription factor recognition sequences determined to be differentially active between (FIG. 28J) cancer vs healthy, (FIG. 28K) adenoma vs healthy, and (FIG. 28L) cancer vs adenoma colon organoid samples.
[0071] FIGS. 29A-29B show the dynamic accessibility peaks cluster near putative regulators of colorectal cancer development. FIGS. 29A-29B: Genome browser images of two loci with the highest concentration per gene of adenoma specific peaks, cancer specific, and / or adenoma and cancer shared peaks.
[0072] FIG. 30 shows the fragmentation features at distal element accessibility peaks. Scatterplot comparing fragmentation features at colon-specific peak centers (±25 bp) to locus19179217191.1DOCKET: 348358.19102flanks (1-2.5 Kbp from peak center) for 25 samples without detectable cancer (MAF = 0%) and 26 samples with detectable cancer (MAF > 25%).
[0073] FIGS. 31A-31E show the genomic context influences cfDNA fragmentation features at colon-specific peaks. FIG. 31A: Receiver operating characteristic curves for identifying individuals with cancer based on each cfDNA fragmentation metric, at each genomic context. FIG. 31B: Heatmap of normalized cfDNA fragmentation metric values for all cfDNA samples, faceted by genomic context. FIGS. 31C, 31D: Heatmap of Pearson correlations of cfDNA fragmentation metrics at each genomic context for all plasma samples from healthy individuals (FIG. 31C) and all those from individuals with cancer (FIG. 31D) FIG. 31E:Regression plots comparing fragmentation metrics at colon-specific peak centers (±10bp) to mutant allele fraction of 47 individuals with cancer, faceted by genomic context.
[0074] FIGS. 32A-32B show the chromatin accessibility signature detection in cfDNA varies by ctDNA burden. FIG. 32A: Cohort-level summary of cfDNA coverage (mean ± 1 SD) at varying levels of ctDNA burden for differentially accessible distal elements characteristic of different pathophysiological phases in CRC development, or specific to PBMCs. FIG. 32B:Cohort-level boxplot summary of normalized fragment coverage distributions in samples with varying levels of ctDNA burden for peak sets in FIG. 32A. Asterisks represent Bonferroni corrected p-values of T-test comparisons to MAF=0% samples.
[0075] FIGS. 33A-33F show the correspondence between PBMC ATAC-seq and cfDNA fragmentation in individuals without cancer. FIG. 33A: Median signal of 10 PBMC ATAC-seq runs in 73,036 consensus peaks (±2.5 Kbp) formed by resolving peak overlaps in all PBMC and colon organoid samples. All peaks are centered at genomic position 0 and ranked from top to bottom by ATAC-seq signal intensity in PBMC samples. FIG. 33B: Metrics for pooled cfDNA data from 306 healthy individuals at loci shown in FIG. 33A. FIG.33C: Boxplot of fragmentation metrics for 306 healthy cfDNA samples. Comparing mean values of cfDNA fragmentation metrics at the center (position 0 ± 25bp) of the 10,000 loci which are most accessible in PBMCs compared to those of the 10,000 loci which are the least accessible. FIG. 33D: Mean ATAC-seq of PBMC samples at 60,000 randomly selected regions, ordered by strength of accessibility at position 0.FIG. 33E: cfDNA fragmentation metrics at randomly selected regions shown in FIG. 33D. FIG.33F: Boxplot of fragmentation metrics for 306 healthy cfDNA samples. Comparing mean values20179217191.1DOCKET: 348358.19102of cfDNA fragmentation metrics at the center (position 0 ± 25bp) of the 10,000 loci which are most accessible in PBMCs compared to those of the 10,000 loci which are the least accessible.DETAILED DESCRIPTION
[0076] This disclosure is based in part, on the finding that changes in chromatin remodeling is a central feature of colorectal tumorigenesis. Briefly, using 43 patient-derived colorectal organoids derived from healthy, adenoma, and cancer tissues, as well as 210 primary cancer and normal tissues from other studies, the chromatin landscape of colorectal tumorigenesis was examined using transposase accessible chromatin analyses with next generation sequencing. A set of 62,708 genomic loci were identified that defined tissue-specific signatures of chromatin accessibility. Within these regions unique colorectal tumor-related signatures were identified comprising genes of known and novel pathways, including those involved in WNT, hippo, and RAS signaling. While the majority of chromatin accessibility differences were shared between adenomas and cancers, 895 chromatin changes were identified in adenomas that were not present in either healthy or malignant tissues. Analysis of differential loci in the circulating cell-free DNA (cfDNA) of 250 healthy individuals and 51 patients with metastatic CRC revealed that chromatin accessibility profiles correlated with genome-wide cell-free DNA fragmentation features and distinguished healthy individuals from those with colorectal cancer (AUC=0.94). These analyses provide evidence that changes in chromatin remodeling are a central feature of colorectal tumorigenesis and also provide an avenue for evaluating chromatin accessibility through evaluation of genome- wide cfDNA fragmentation.
[0077] CHROMATIN ACCESSIBILITY
[0078] Mapping alterations in cell states is a key aspect of understanding biological systems. Whether in development, differentiation or disease, cell state is governed by changes in gene expression that are, in turn, orchestrated by changes in gene regulatory programs. In recent years, it has become increasingly clear that these gene regulator}' programs are established and controlled by the activity of transcription factors (TFs) that both interpret and alter the underlying epigenetic state of chromatin. The epigenetic state of chromatin can be regulated by a variety of mechanisms, including chemical modification of both DNA and histone proteins that,21179217191.1DOCKET: 348358.19102in turn, alter chromatin dynamics and high-dimensional chromatin structure. It is now recognized that chromatin can exist in several different states that are defined by combinations of different epigenetic modifications and are associated with particular gene regulatory patterns. At the two ends of the spectrum are (i) active gene regulatory elements such as enhancers, promoters and insulators, which are bound by DNA-binding proteins, and (ii) inactive regions of silenced or poised chromatin, which are generally refractory to gene expression machinery. Understanding the epigenetic state of chromatin in a certain biological context can shed light onto the molecular mechanisms underlying the observed gene expression patterns.
[0079] ATAC-seq: The original DNase-seq and MNase-seq assays traditionally had complex, time-consuming library preparation protocols and required large numbers of cells as starting material. To address some of these limitations, while keeping the agnostic profiling of chromatin, the assay for transposase-accessible chromatin using sequencing (ATAC-seq) was developed. ATAC-seq uses the activity of an engineered, hyperactive Tn5 transposase preloaded with sequencing adapters to determine the sites of accessible chromatin. The development of ATAC-seq was based on two observations: (i) a transposase had previously been used to generate ‘tagmentation’ libraries, in which a Tn5 transposase was preloaded with sequencing adapters and used to simultaneously fragment and tag genomic DNA for high-throughput sequencing library preparation and (ii) the observation that in vivo Tn5 could efficiently insert into nucleosome-free regions. ATAC-seq generates genome-wide regulatory maps that are highly similar to those derived from DNase-seq and MNase-seq, while reducing library preparation complexity and hands-on time. ATAC-seq has been widely adopted owing to its low input material requirements (<50,000 cells) and short processing time, which facilitates data generation from large numbers of samples.
[0080] Extensive profiling efforts have shown that regions of Tn5-accessbile chromatin can be found at promoters, located proximal to the transcription start site (TSS), and at intergenic regions of the genome largely corresponding to enhancers, insulators or silencers (Grandi, F.C., Modi, H., Kampman, L. et al. Chromatin accessibility profiling by ATAC-seq. Nat Protoc 17, 1518-1552 (2022). doi.org / 10.1038 / s41596-022-00692-9). These patterns and locations of Tn5 accessibility, especially those at distal elements, are often cell type or cell state specific. Thus, ATAC-seq represents a valuable tool to understand how cells control gene expression, by22179217191.1DOCKET: 348358.19102mapping the location of putative gene regulatory elements. After processing and alignment of ATAC-seq fragments, enrichment of Tn5 transposition events at specific genomic regions is used to identify peaks of Tn5-accessible chromatin in each sample. These are often termed ‘ATAC-seq peaks’. Chromatin accessibility signal within these peak regions can be compared between different sample types using established pipelines and serve as the starting point for a variety of downstream analyses. For example, peaks can be linked to putative gene targets by using orthogonal chromatin conformation capture datasets or by naively assigning each peak to the nearest gene. These predicted gene regulatory interactions can provide a hint as to the functional importance of a given peak. Often, genes with several ATAC-seq peaks in their promoter and gene body are inferred to be actively expressed in that cell type. While gene expression is more accurately measured by RNA sequencing (RNA-seq), ATAC-seq can explain the mechanism behind how gene expression is regulated or why it might be different between two cell types or conditions.
[0081] cfDNA Fragmentation Profiles. A cfDNA fragmentation profile can include one or more cfDNA fragmentation patterns. A cfDNA fragmentation pattern can include any appropriate cfDNA fragmentation pattern. Examples of cfDNA fragmentation patterns include, without limitation, median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and the coverage of cfDNA fragments. In some embodiments, a cfDNA fragmentation pattern includes two or more (e.g., two, three, or four) of median fragment size, fragment size distribution, ratio of small cfDNA fragments to large cfDNA fragments, and the coverage of cfDNA fragments. In some embodiments, cfDNA fragmentation profile can be a genome-wide cfDNA profile (e.g., a genome-wide cfDNA profile in windows across the genome). In some embodiments, cfDNA fragmentation profile can be a targeted region profile. A targeted region can be any appropriate portion of the genome (e.g., a chromosomal region). Examples of chromosomal regions for which a cfDNA fragmentation profile can be determined as described herein include, without limitation, a portion of a chromosome (e.g., a portion of 2q, 4p, 5p, 6q, 7p, 8q, 9q, lOq, 1 Iq, 12q, and / or 14q) and a chromosomal arm (e.g., a chromosomal mm of 8q, 13q, 1 Iq, and / or 3p). In some embodiments, a cfDNA fragmentation profile can include two or more targeted region profiles.23179217191.1DOCKET: 348358.19102
[0082] In some embodiments, a cfDNA fragmentation profile can be used to identify changes (e.g., alterations) in cfDNA fragment lengths. An alteration can be a genome-wide alteration or an alteration in one or more targeted regions / loci. A target region can be any region containing one or more cancer-specific alterations. In some embodiments, a cfDNA fragmentation profile can be used to identify (e.g., simultaneously identify) from about 10 alterations to about 500 alterations (e.g., from about 25 to about 500, from about 50 to about 500, from about 100 to about 500, from about 200 to about 500, from about 300 to about 500, from about 10 to about 400, from about 10 to about 300, from about 10 to about 200, from about 10 to about 100, from about 10 to about 50, from about 20 to about 400, from about 30 to about 300, from about 40 to about 200, from about 50 to about 100, from about 20 to about 100, from about 25 to about 75, from about 50 to about 250, or from about 100 to about 200, alterations).
[0083] A cfDNA fragmentation profile can be obtained using any appropriate method. In some embodiments, cfDNA from a mammal (e.g., a mammal having, or suspected of having, cancer) can be processed into sequencing libraries which can be subjected to whole genome sequencing (e.g., low-coverage whole genome sequencing), mapped to the genome, and analyzed to determine cfDNA fragment lengths. Mapped sequences can be analyzed in non-overlapping windows covering the genome. Windows can be any appropriate size. For example, windows can be from thousands to millions of bases in length. As one non-limiting example, a window can be about 5 megabases (Mb) long. Any appropriate number of windows can be mapped. For example, tens to thousands of windows can be mapped in the genome. For example, hundreds to thousands of windows can be mapped in the genome. A cfDNA fragmentation profile can be determined within each window.
[0084] In the early detection of cancers, any of the systems or methods herein described, including mutation detection or copy number variation detection may be utilized to detect cancers. These system and methods may be used to detect any number of genetic aberrations that may cause or result from cancers. These may include but are not limited to cfDNA chromatin accessible profiles, cfDNA mutation profiles, frequency of mutations, cfDNA fragmentation profiles, mutations, mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability,24179217191.1DOCKET: 348358.19102chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation infection and cancer.
[0085] In some embodiments, methods and materials described herein also can include machine learning. For example, machine learning can be used for identifying chromatin accessible gene loci, mutation frequencies, altered fragmentation profile (e.g., using coverage of cfDNA fragments, fragment size of cfDNA fragments, coverage of chromosomes, and mtDNA).
[0086] METHODS OF TREATMENT
[0087] The methods embodied herein, include identifying a mammal as having cancer. The methods include, extracting cell-free DNA (cfDNA) from a subject’s biological sample; generating genomic libraries from the extracted cfDNA; sequencing individual cfDNA molecules to obtain tissue specific chromatin accessibility regions in patients with and without cancer, wherein the subject with colorectal cancer comprises a cfDNA fragmentation profile distinct to other cancer types; and treating the subject with colorectal specific therapeutics. In certain embodiments, the cfDNA fragmentation profiles comprise cfDNA fragmentation profiles of consensus peaks. In certain embodiments, evaluation of cfDNA features to produce a cfDNA fragmentation profile comprises evaluating sequence coverage, fragment size, and frequency of fragments ending at specific positions normalized by sequence coverage. In certain embodiments, higher chromatin accessibility at a center of consensus peaks comprises decreased cfDNA sequence coverage, increased fragment end frequency, a reduction in size of cfDNA fragments or combinations thereof.
[0088] In one aspect, a method of early detection of cancer and treatment of a subject comprises (i) determining a cell free DNA (cfDNA) fragmentation profile of the subject, the method comprising: extracting and enriching cell free DNA (cfDNA) from a subject’s biological sample; conducting low coverage whole genome sequencing of the subject’s isolated cfDNA to generate genomic libraries of sequenced cfDNA fragments; determining a cfDNA fragmentome profile of the subject’s sequenced DNA fragments comprising evaluating one or more cfDNA fragment characteristics; (ii) comparing the cfDNA fragmentome profiles of the subject to25179217191.1DOCKET: 348358.19102reference cfDNA fragmentome profiles from healthy subjects to detect and diagnose the cancer; comparing the subject’s cfDNA fragmentation profile and with normal reference non-cancer subjects; and, treating the subject diagnosed with cancer with a cancer specific therapy. In certain embodiments, cfDNA fragmentation profile comprises chromatin accessible regions.
[0089] In certain embodiments, cfDNA fragmentation profiles of healthy subjects comprise similar cfDNA fragmentation profiles. In certain embodiments, cfDNA fragmentation profiles of subjects with cancer are similar based on cancer type. In certain embodiments, cfDNA fragmentation profiles of subjects with colorectal cancer have similar cfDNA fragmentation profiles. In certain embodiments, cfDNA fragmentation profiles of subjects with colorectal cancer are distinct with respect to another type of cancer.
[0090] In certain embodiments, a subject is diagnosed as having cancer, e.g. early stage cancer. In certain embodiments, the type of cancer is identified, and the cancer is treated by various therapeutics, including therapeutics specific for the type of cancer. In certain embodiments, the cancer comprises colorectal cancer, lung cancer, breast cancer, gastric cancers, pancreatic cancers, bile duct cancers, brain cancer or ovarian cancer. In certain embodiments, the cancer is colorectal cancer.
[0091] In some embodiments, the therapy treats a tumor derived from a cancer, which is colorectal cancer. In some embodiments, the colorectal cancer is colon cancer. In other embodiments, the colorectal cancer is rectal cancer.
[0092] Colon cancer presents in five stages: Stage 0 (Carcinoma in situ), Stage I, Stage II, Stage III and Stage IV. Six types of standard treatment are used for colon cancer: 1) surgery, including a local excision, resection of the colon with anastomosis, or resection of the colon with colostomy; 2) radiofrequency ablation; 3) cryosurgery; 4) chemotherapy; 5) radiation therapy; and 6) targeted therapies, including monoclonal antibodies and angiogenesis inhibitors. In some embodiments, the combination therapy of the disclosure treats a colon cancer along with a standard of care therapy.
[0093] Rectal cancer presents in five stages: Stage 0 (Carcinoma in situ), Stage I, Stage II, Stage III and Stage IV. Six types of standard treatment are used for rectal cancer: 1) Surgery, including polypectomy, local excision, resection, radiofrequency ablation, cryosurgery, and 26179217191.1DOCKET: 348358.19102pelvic exenteration; 2) radiation therapy; 3) chemotherapy; and 4) targeted therapy, including monoclonal antibody therapy. In some embodiments, the methods of the disclosure treats a rectal cancer along with a standard of care therapy.
[0094] In certain embodiments, the cancer treatment can be surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, or any combinations thereof. The method also can include administering to the mammal a cancer treatment (e.g., surgery, adjuvant chemotherapy, neoadjuvant chemotherapy, radiation therapy, hormone therapy, cytotoxic therapy, immunotherapy, adoptive T cell therapy, targeted therapy, or any combinations thereof). The mammal can be monitored for the presence of cancer after administration of the cancer treatment.
[0095] Cancer therapies in general also include a variety of combination therapies with both chemical and radiation-based treatments. Combination chemotherapies include, for example, cisplatin (CDDP), carboplatin, procarbazine, mechlorethamine, cyclophosphamide, camptothecin, ifosfamide, melphalan, chlorambucil, busulfan, nitrosurea, dactinomycin, daunorubicin, doxorubicin, bleomycin, plicomycin, mitomycin, etoposide (VP 16), tamoxifen, raloxifene, estrogen receptor binding agents, taxol, gemcitabien, navelbine, famesyl-protein transferase inhibitors, transplatinum, 5-fluorouracil, vincristine, vinblastine and methotrexate, Temazolomide (an aqueous form of DTIC), or any analog or derivative variant of the foregoing. The combination of chemotherapy with biological therapy is known as biochemotherapy. The chemotherapy may also be administered at low, continuous doses which is known as metronomic chemotherapy.
[0096] Yet further combination chemotherapies include, for example, alkylating agents such as thiotepa and cyclosphosphamide; alkyl sulfonates such as busulfan, improsulfan and piposulfan; aziridines such as benzodopa, carboquone, meturedopa, and uredopa; ethylenimines and methylamelamines including altretamine, triethylenemelamine, trietylenephosphoramide, triethiylenethiophosphoramide and trimethylolomelamine; acetogenins (especially bullatacin and bullatacinone); a camptothecin (including the synthetic analogue topotecan); bryostatin; cally statin; CC-1065 (including its adozelesin, carzelesin and bizelesin synthetic analogues);27179217191.1DOCKET: 348358.19102cryptophycins (particularly cryptophycin 1 and cryptophycin 8); dolastatin; duocarmycin (including the synthetic analogues, KW-2189 and CB1-TM1); eleutherobin; pancrati statin; a sarcodictyin; spongistatin; nitrogen mustards such as chlorambucil, chlomaphazine, cholophosphamide, estramustine, ifosfamide, mechlorethamine, mechlorethamine oxide hydrochloride, melphalan, novembichin, phenesterine, prednimustine, trofosfamide, uracil mustard; nitrosureas such as carmustine, chlorozotocin, fotemustine, lomustine, nimustine, and ranimnustine; antibiotics such as the enediyne antibiotics (e.g., calicheamicin, especially calicheamicin gammall and calicheamicin omegall; dynemicin, including dynemicin A; bisphosphonates, such as clodronate; an esperamicin; as well as neocarzinostatin chromophore and related chromoprotein enediyne antiobiotic chromophores, aclacinomysins, actinomycin, authramycin, azaserine, bleomycins, cactinomycin, carabicin, carminomycin, carzinophilin, chromomycinis, dactinomycin, daunorubicin, detorubicin, 6-diazo-5-oxo-L-norleucine, doxorubicin (including morpholino-doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino-doxorubicin and deoxydoxorubicin), epirubicin, esorubicin, idarubicin, marcellomycin, mitomycins such as mitomycin C, mycophenolic acid, nogalarnycin, olivomycins, peplomycin, potfiromycin, puromycin, quelamycin, rodornbicin, streptonigrin, streptozocin, tubercidin, ubenimex, zinostatin, zombicin; anti-metabolites such as methotrexate and 5-fluorouracil (5-FU); folic acid analogues such as denopterin, pteropterin, trimetrexate; purine analogs such as fludarabine, 6-mercaptopurine, thiamiprine, thioguanine; pyrimidine analogs such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, floxuridine; androgens such as calusterone, dromostanolone propionate, epitiostanol, mepitiostane, testolactone; anti-adrenals such as rnitotane, trilostane; folic acid replenisher such as frolinic acid; aceglatone; aldophosphamide glycoside; aminolevulinic acid; eniluracil; arnsacrine; bestrabucil; bisantrene; edatraxate; defofamine; demecolcine; diaziquone; elformithine; elliptinium acetate; an epothilone; etoglucid; gallium nitrate; hydroxyurea; lentinan; lonidainine; maytansinoids such as maytansine and ansamitocins; mitoguazone; mitoxantrone; mopidanmol; nitraerine; pentostatin; phenamet; pirarubicin; losoxantrone; podophyllinic acid; 2-ethylhydrazide; procarbazine; PSK polysaccharide complex; razoxane; rhizoxin; sizofiran; spirogermanium; tenuazonic acid; triaziquone; 2, 2’, 2” -tri chlorotri ethylamine; trichothecenes (especially T-2 toxin, verracurin A, roridin A and anguidine); urethan; vindesine;28179217191.1DOCKET: 348358.19102dacarbazine; mannomustine; mitobronitol; mitolactol; pipobroman; gacytosine; arabinoside (“Ara-C”); cyclophosphamide; taxoids, e.g., paclitaxel and docetaxel gemcitabine; 6-thioguanine; mercaptopurine; platinum coordination complexes such as cisplatin, oxaliplatin and carboplatin; vinblastine; platinum; etoposide (VP- 16); ifosfamide; mitoxantrone; vincristine; vinorelbine; novantrone; teniposide; edatrexate; daunomycin; aminopterin; xeloda; ibandronate; irinotecan (e.g., CPT-11); topoisomerase inhibitor RPS 2000; difluorometlhylornithine (DMFO); retinoids such as retinoic acid; capecitabine; carboplatin, procarbazine, plicomycin, gemcitabien, navelbine, farnesyl-protein transferase inhibitors, transplatinum; and pharmaceutically acceptable salts, acids or derivatives of any of the above.
[0097] Immunotherapeutics, generally, rely on the use of immune effector cells and molecules to target and destroy cancer cells. The immune effector may be, for example, an antibody specific for some marker on the surface of a tumor cell. The antibody alone may serve as an effector of therapy, or it may recruit other cells to actually effect cell killing. The antibody also may be conjugated to a drug or toxin (chemotherapeutic, radionuclide, ricin A chain, cholera toxin, pertussis toxin, etc.) and serve merely as a targeting agent. Alternatively, the effector may be a lymphocyte carrying a surface molecule that interacts, either directly or indirectly, with a tumor cell target. Various effector cells include cytotoxic T cells and NK cells as well as genetically engineered variants of these cell types modified to express chimeric antigen receptors.
[0098] The immunotherapy may comprise suppression of T regulatory cells (Tregs), myeloid derived suppressor cells (MDSCs) and cancer associated fibroblasts (CAFs). In some embodiments, the immunotherapy is a tumor vaccine (e.g., whole tumor cell vaccines, peptides, and recombinant tumor associated antigen vaccines), or adoptive cellular therapies (ACT) (e.g., T cells, natural killer cells, TILs, and LAK cells). The T cells may be engineered with chimeric antigen receptors (CARs) or T cell receptors (TCRs) to specific tumor antigens. As used herein, a chimeric antigen receptor (or CAR) may refer to any engineered receptor specific for an antigen of interest that, when expressed in a T cell, confers the specificity of the CAR onto the T cell. Once created using standard molecular techniques, a T cell expressing a chimeric antigen receptor may be introduced into a patient, as with a technique such as adoptive cell transfer. In29179217191.1DOCKET: 348358.19102some aspects, the T cells are activated CD4 and / or CD8 T cells in the individual which are characterized by y-lFN- producing CD4 and / or CD8 T cells and / or enhanced cytolytic activity relative to prior to the administration of the combination. The CD4 and / or CD8 T cells may exhibit increased release of cytokines selected from the group consisting of IFN-y, TNF-rz and interleukins. The CD4 and / or CD8 T cells can be effector memory T cells. In certain embodiments, the CD4 and / or CDS effector memory T cells are characterized by having the expression of CD44hlghCD62Llow.
[0099] The immunotherapy may be a cancer vaccine comprising one or more cancer antigens, in particular a protein or an immunogenic fragment thereof, DNA or RNA encoding said cancer antigen, in particular a protein or an immunogenic fragment thereof, cancer cell lysates, and / or protein preparations from tumor cells. As used herein, a cancer antigen is an antigenic substance present in cancer cells. In principle, any protein produced in a cancer cell that has an abnormal structure due to mutation can act as a cancer antigen. In principle, cancer antigens can be products of mutated Oncogenes and tumor suppressor genes, products of other mutated genes, overexpressed or aberrantly expressed cellular proteins, cancer antigens produced by oncogenic viruses, oncofetal antigens, altered cell surface glycolipids and glycoproteins, or cell type-specific differentiation antigens. Examples of cancer antigens include the abnormal products of ras and p53 genes. Other examples include tissue differentiation antigens, mutant protein antigens, oncogenic viral antigens, cancer-testis antigens and vascular or stromal specific antigens. Tissue differentiation antigens are those that are specific to a certain type of tissue. Mutant protein antigens are likely to be much more specific to cancer cells because normal cells shouldn’t contain these proteins. Normal cells will display the normal protein antigen on their MHC molecules, whereas cancer cells will display the mutant version. Some viral proteins are implicated in forming cancer, and some viral antigens are also cancer antigens. Cancer-testis antigens are antigens expressed primarily in the germ cells of the testes, but also in fetal ovaries and the trophoblast. Some cancer cells aberrantly express these proteins and therefore present these antigens, allowing attack by T-cells specific to these antigens. Exemplary antigens of this type are CTAG1 B and MAGEA1 as well as Rindopepimut, a 14-mer intradermal injectable peptide vaccine targeted against epidermal growth factor receptor vlll (EGFRvlll; deletion of exons 2-7) variant. Rindopepimut is particularly suitable for treating glioblastoma when used in 30179217191.1DOCKET: 348358.19102combination with an inhibitor of the CD95 / CD95L signaling system as described herein. Also, proteins that are normally produced in very low quantities, but whose production is dramatically increased in cancer cells, may trigger an immune response. An example of such a protein is the enzyme tyrosinase, which is required for melanin production. Normally tyrosinase is produced in minute quantities but its levels are very much elevated in melanoma cells. Oncofetal antigens are another important class of cancer antigens. Examples are alphafetoprotein (AFP) and carcinoembryonic antigen (CEA). These proteins are normally produced in the early stages of embryonic development and disappear by the time the immune system is fully developed. Thus, self-tolerance does not develop against these antigens. Abnormal proteins are also produced by cells infected with oncoviruses, e.g. EBV and HPV. Cells infected by these viruses contain latent viral DNA which is transcribed, and the resulting protein produces an immune response. A cancer vaccine may include a peptide cancer vaccine, which in some embodiments is a personalized peptide vaccine. In some embodiments, the peptide cancer vaccine is a multivalent long peptide vaccine, a multi-peptide vaccine, a peptide cocktail vaccine, a hybrid peptide vaccine, or a peptide-pulsed dendritic cell vaccine
[0100] The immunotherapy may be an antibody, such as part of a polyclonal antibody preparation, or may be a monoclonal antibody. The antibody may be a humanized antibody, a chimeric antibody, an antibody fragment, a bispecific antibody or a single chain antibody. An antibody as disclosed herein includes an antibody fragment, such as, but not limited to, Fab, Fab’ and F(ab’)2, Fd, single-chain Fvs (scFv), single-chain antibodies, disulfide-linked Fvs (sdfv) and fragments including either a VL or VH domain. In some aspects, the antibody or fragment thereof specifically binds epidermal growth factor receptor (EGFR1, Erb-Bl), HER2 / neu (Erb-62), CD20, Vascular endothelial growth factor (VEGF), insulin-like growth factor receptor (IGF-1R), TRAIL-receptor, epithelial cell adhesion molecule, carcinoembryonic antigen, Prostate-specific membrane antigen, Mucin-1, CD30, CD33, or CD40.
[0101] Examples of monoclonal antibodies include, without limitation, trastuzumab (anti-HER2 / neu antibody); Pertuzumab (anti-HER2 mAb); cetuximab (chimeric monoclonal antibody to epidermal growth factor receptor EGFR); panitumumab (anti-EGFR antibody); nimotuzumab (anti-EGFR antibody); Zalutumumab (anti-EGFR mAb); Necitumumab (anti-31179217191.1DOCKET: 348358.19102EGFR mAb); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-210 (humanized anti-HER-2 bispecific antibody); MDX-447 (humanized anti-EGF receptor bispecific antibody); Rituximab (chimeric murine / human anti-CD20 mAb); Obinutuzumab (anti-CD20 mAb);Ofatumumab (anti-CD20 mAb); Tositumumab-1131 (anti-CD20 mAb); Ibritumomab tiuxetan (anti-CD20 mAb); Bevacizumab (anti-VEGF mAb); Ramucirumab (anti-VEGFR2 mAb);Ranibizumab (anti-VEGF mAb); Aflibercept (extracellular domains of VEGFR1 and VEGFR2 fused to IgGl Fc); AMG386 (angiopoietin-1 and -2 binding peptide fused to IgGl Fc);Dalotuzumab (anti-IGF-lR mAb); Gemtuzumab ozogamicin (anti-CD33 mAb); Alemtuzumab (anti-Campath- 1 / CD52 mAb); Brentuximab vedotin (anti-CD30 mAb); Catumaxomab (bispecific mAb that targets epithelial cell adhesion molecule and CD3); Naptumomab (anti-5T4 mAb); Girentuximab (anti-Carbonic anhydrase ix); or Farletuzumab (anti-folate receptor). Other examples include antibodies such as Panorex™ (17-1 A) (murine monoclonal antibody); Panorex (MAbl7-lA) (chimeric murine monoclonal antibody); BEC2 (ami-idiotypic mAb, mimics the GD epitope) (with BCG); Oncolym (Lym-1 monoclonal antibody); SMART Ml 95 Ab, humanized 13’ 1 LYM-1 (Oncolym), Ovarex (B43.13, anti -idiotypic mouse mAb); 3622W94 mAb that binds to EGP40 (17-1 A) pancarcinoma antigen on adenocarcinomas; Zenapax (SMART Anti-Tac (IL-2 receptor); SMART Ml 95 Ab, humanized Ab, humanized); NovoMAb-G2 (pancarcinoma specific Ab); TNT (chimeric mAb to histone antigens); TNT (chimeric mAb to histone antigens); Gliomab-H (Monoclonals-Humanized Abs); GNI-250 Mab; EMD-72000 (chimeric-EGF antagonist); LymphoCide (humanized IL.L.2 antibody); and MDX-260 bispecific, targets GD-2, ANA Ab, SMART IDIO Ab, SMART ABL 364 Ab or ImmuRAIT-CEA. Further examples of antibodies include Zanulimumab (anti-CD4 mAb), Keliximab (anti-CD4 mAb); Ipilimumab (MDX-101; anti-CTLA-4 mAb); Tremilimumab (anti-CTLA-4 mAb); (Daclizumab (anti-CD25 / IL-2R mAb); Basiliximab (anti-CD25 / IL-2R mAb); MDX-1106 (anti-PDl mAb); antibody to GITR; GC1008 (anti-TGF- antibody); metelimumab / CAT-192 (anti- TGF-P antibody); lerdelimumab / CAT-152 (anti-TGF-0 antibody); ID11 (anti-TGF-P antibody); Denosumab (anti-RANKL mAb); BMS-663513 (humanized anti-4-lBB mAb); SGN-40 (humanized anti-CD40 mAb); CP870,893 (human anti-CD40 mAb); Infliximab (chimeric anti-TNF mAb; Adalimumab (human anti-TNF mAb); Certolizumab (humanized Fab anti-TNF); Golimumab (anti-TNF); Etanercept (Extracellular domain of TNFR fused to IgGl Fc);32179217191.1DOCKET: 348358.19102Belatacept (Extracellular domain of CTLA-4 fused to Fe); Abatacept (Extracellular domain of CTLA-4 fused to Fe); Belimumab (anti-B Lymphocyte stimulator); Murom onab-CD3 (anti-CD3 mAb); Otelixizumab (anti-CD3 mAb); Teplizumab (anti-CD3 mAb); Tocilizumab (anti-IL6R mAb); REGN88 (anti-IL6R mAb); Ustekinumab (anti-IL- 12 / 23 mAb); Briakinumab (anti-IL-12 / 23 mAb); Natalizumab (anti-a4 integrin); Vedolizumab (anti-a4 [37 integrin mAb); T1 h (anti-CD6 mAb); Epratuzumab (anti-CD22 mAb); Efalizumab (anti-CDl la mAb); and Atacicept (extracellular domain of transmembrane activator and calcium-modulating ligand interactor fused with Fc).EXAMPLES
[0102] EXAMPLE 1 : CHROMATIN LANDSCAPES
[0103] There is an increasing realization that insights into cancer chromatin dynamics may be useful for the understanding of tumor-derived cell-free DNA (cfDNA) fragmentation. cfDNA liquid biopsy methods of blood or other bodily fluids provide an opportunity for noninvasive disease detection, therapeutic stratification, and monitoring during therapy20 28. In healthy individuals, the majority of cfDNA in the blood originates from peripheral blood mononuclear cells (PBMCs) which have undergone apoptosis as a part of their normal life cycle29. Growing evidence suggests that cfDNA fragmentation patterns are determined by tissue-specific factors such as overall chromatin organization and transcription factor binding30,31. In patients with cancer, alterations in cfDNA fragmentation have been linked to changes in cancer at transcription factor binding sites or select regions of chromatin accessibility26,30. In adenomas, it is currently unclear whether overall chromatin organization sufficiently differs from normal colonic epithelium to result in different fragmentation patterns to be detectable in cfDNA, and whether these patterns differ from those in adenocarcinomas. In this study, ATAC-seq was used to identify genome-wide chromatin changes during colorectal cancer development. Also, the characteristics of cfDNA fragmentation at tissue-specific accessibility regions in patients with and without cancer were analyzed.
[0104] METHODS
[0105] Sample collection
[0106] Colonic tissue was obtained from material resected during a colonoscopy procedure either at the VU University Medical Center or MC Slotervaart, Amsterdam, or a surgical procedure 33179217191.1DOCKET: 348358.19102at the Netherlands Cancer Institute (NKT). Briefly, after resection, tissues were sent to the pathology department of the respective center. After macroscopic examination of the specimen, the central part of the lesion was secured for diagnostic purposes. The left-over tissue was split for research purposes, one portion of which was used for the establishment of an organoid culture. The protocols for sample collection were approved by the institutional review board of the NKI under IRBm20-079. Peripheral blood mononuclear cell (PBMC) samples were purchased from i Specimen.
[0107] Plasma samples from healthy individuals were procured from the Danish Endoscopy III trial65and the colonoscopy arm of the Cocos trial (COCOS, Netherlands Trial Register ID NTR1829)66. The protocols for the Danish Endoscopy III Project were approved by The Danish Data Protection Agency (2007-58-0015 / HVH-2013-022) and the Regional Ethics Committee (H-4-2013-050), while ethical approval for the COCOS trial was granted by the Dutch Health Council. Plasma samples of individuals with colorectal cancer were procured as a part of Arm 1 of the CAIRO5 trial (NCT02162563). All these individuals had a somatic mutation in RAS or BRAF which was used to measure mutant allele fraction by droplet digital PCR (ddPCR).
[0108] Organoid culture
[0109] Organoid cultures were established and cultured as previously described in Martens-de Kemp et al., which built upon protocols established by Van de Wetering et al.6932. Briefly, organoid cultures were inspected every 2-3 days, upon which either the culture medium was refreshed or the organoids were passaged. For passaging, the culture medium was removed, 500 pl ice-cold ADF+++ was added and the Matrigel® Matrix was broken up by pipetting. To wash away the Matrigel® Matrix, organoids were collected in a cold 50 ml tube and cold ADF+++ was added to a final volume of 15 ml. After spinning, the pellet was resuspended in 1 ml of TrypLE Express (Gibco, Waltham, MA, USA) and incubated at 37°C for 5 minutes. Then, ADF+++ was added and the tube was centrifuged again at 1050 rpm for 5 minutes. The pellet was resuspended in Matrigel® Matrix and the cells were plated in droplets in a pre-warmed 24-wells culture plate. After allowing the Matrigel® Matrix to solidify at 37°C for 10 minutes, 500 pl of ADF+++ / BCM (supplemented with 10 pM LY27632) was added to each well and cells were incubated at 37°C.
[0110] ATAC-seq of organoid samples34179217191.1DOCKET: 348358.19102
[0111] Colorectal organoids that were cultured from normal colonic mucosa, adenomas or colorectal cancer tissue were analyzed using Omni ATAC-seq75with some modifications. Each organoid culture was aliquoted into two 50,000 cell, technical replicates, lysis incubation time was 5 minutes, nuclei were pelleted for 10 minutes and tagmentation enzyme (Illumina) concentration was 400nM. Paired-end sequencing was performed on an Illumina NovaSeq6000, with a read length of 151 bp. A sequencing replicate refers to a genomic library from a single donor sequenced in two separate batches.
[0112] ATAC-seq data processing
[0113] ATAC-seq data were uniformly processed by PEPATAC vO.10.076using the same tools and parameters as Corces et al. 2018, unless otherwise noted. Skewer vO.2.277was used to trim adaptors from the raw reads. For samples with paired-end reads, PEPATAC was used to remove reads with no matching pair.
[0114] To filter out non-salient reads, a “prealignment” step was taken to identify and exclude all reads that mapped to any of the following reference sequences: chrM (revised Cambridge Reference Sequence78), human alpha satellite regions, human Alu repeats, human rDNA, and human repeat sequences. Finally, the resulting reads were mapped to the hgl9 reference sequence.
[0115] Bowtie2 v2.9.279was used for all alignments and Refgenie vO.12.O80was used to download and manage all reference sequences. Note, our bowtie2 parameters for the prealignment (-X 2000 — rg-id) deviate from those used in Corces 2018 (-k 1 -D 20 -R 3 -N 1 -L 20 -i S, 1,0.50 -X 2000 -rg-id).
[0116] SAMtools v 1.1381was used to sort and filter (samtools view -b -q 10 -@ 1) the resulting bam files. Duplicated reads were removed by Picard v 2.20.3 MarkDuplicates (broadinstitute.github.io / picard / ). MACS2 v2.2.7.182callpeaks was used to call peaks on the deduplicated bam files. These peaks were centered and padded on each side by 250bp, to a final width of 501 bp. Peaks were excluded if they fell within the boundaries of the ENCODE exclusion list83(Accession ENCFF001TDO). This filtering was performed with bedtools v2.24.084.
[0117] Only those samples that passed ENCODE quality control thresholds were included in the analysis: fraction of reads in peaks (FRiP) >0.2, PCR bottlenecking coefficient 1 (PBC1) >0.7, and PBC2>1, transcription start site (TSS) score >6, Peak count >100 thousand, and35179217191.1DOCKET: 348358.19102Deduplicated Aligned Reads >50 million83. The original standards suggested that non-redundant fraction (NRF) would become a less meaningful metric at high sequencing depths, and therefore we relaxed the threshold to NRF>0.25 for the organoid samples. Concordance between replicates was evaluated by collating peaks called from all sequencing or technical replicates of a given sample and calculating Pearson correlations between all pairwise combinations of those replicates. Samples with Pearson correlation <0.6 in all replicate comparisons were removed from the analysis.
[0118] Consensus peak construction, annotation, and analysis
[0119] ATAC-seq peaks from multiple samples were coalesced into consensus peaks following the iterative overlap peak merging algorithm described in Grandi 202286. These consensus peaks were generated using their provided code with the default parameters (spm=5, rule = “(n+l) / 2”, extend=250). Consensus peaks were assigned one of three annotations in the following order of priority: promoter, exon, and CTCF binding site. TSS and exon coordinates were taken from the Gencode v26 basic annotations for hgl9, while CTCF binding site coordinates for hgl9 were taken from the Homer software v4.11.1 known motifs database87,88. Promoter coordinates were defined as the region 1 Kbp upstream of a TSS to lOObp downstream of the TSS. All consensus peaks which did not map to any genomic annotations with an overlap of at least Ibp were assigned the category of “distal element”. Genome-wide peak and gene density were calculated and visualized as Circos plots using the circleize Bioconductor package version 0.4.1689,90.
[0120] Statistical analyses of consensus tables were performed in the R programming language v4.3.291. First, the read counts for all peaks of a given sample were centered to have mean 0 and scaled to have standard deviation 1. Principal component analysis (PCA) was then performed on the ATAC-seq read count consensus table of all samples using prcomp function, with scaling and centering arguments set to TRUE. Similarly, all heatmap visualizations labeled “Accessibility (Normalized)” were scaled in this way, first by sample, then by peak.
[0121] T o understand whether the PCI and PC2 axes were a consequence of our supervi sed approach to peak selection, we repeated our analysis using randomized grouping labels in the place of tissue labels and found that PCA of these consensus peaks revealed that samples continued to cluster by their true tissue of origin (FIG. 25E).36179217191.1DOCKET: 348358.19102
[0122] Selecting chromatin signatures of colorectal cancer development
[0123] The PCA of the colon consensus table (62,890 peaks) was used to evaluate the main sources of variation between the stages of colorectal cancer development. The percent contribution of individual peaks to each principal component (PC) were evaluated by the get_pca_var function in the factoextra package vl.0.792. Elements of the eigenvectors having the highest percent contribution to the principal component were selected, and that jointly summed to 90% of the total variation explained by that principal component. We refer to the collection of these elements that correspond to peaks on ATAC-seq axes as chromatin signatures. The PC coordinate values we report for each peak were calculated as follows by factoextra:Singular valuePC Coordinates = Read count eigenvector xg# Samples — 1
[0124] Normalized accessibility at chromatin signature sites were visualized using ComplexHeatmap v2.18.093.
[0125] Prediction of tissue type from chromatin signatures
[0126] To understand whether chromatin signatures would be a useful feature for predicting colorectal disease state (Healthy, Adenoma, or Cancer), we cross-validated the procedure for deriving chromatin signatures and evaluated confusion matrices of linear discriminant models for the classification accuracy. Briefly, we randomized the assignment of each organoid sample into one of seven training folds, ensuring that (1) each fold included healthy, adenoma, and cancer samples, and (2) all samples from the same donor were assigned to the same fold. Leaving out one fold, we constructed the consensus peaks using the remaining six folds and performed PCA on the resulting table of read counts as previously described. Using the Ida function from the R package MASS v7.3.60.0.194, we trained a linear discriminant model on the first two PCs of each training consensus table with colorectal tissue type as the response variable. We used this model to predict disease state of the held-out donors, and evaluated the model performance with respect to accuracy, precision, and recall of these disease state assignments. We repeated this procedure for each of the seven possible hold-out folds.
[0127] DNA binding element motif analysis
[0128] All peaks in the chromatin signatures were analyzed for de novo transcription factor motifs by Homer software v4.11.188function fmdMotifsGenome.pl. For each chromatin signature,37179217191.1DOCKET: 348358.19102peaks which had opposite PC coordinate sign (+ or -) were handled separately. All peaks with a positive PC coordinate were provided as the target BED file to the Homer function, while the corresponding peaks with a negative contribution were provided as the background peak set. This was repeated with negatively contributing peaks as the target, and positively contributing peaks as the background. Motifs were ranked by the magnitude of their enrichment within these regions, and only motifs with an enrichment p-value < le-10 were reported.
[0129] Chromatin signature gene set enrichment analysis
[0130] Sets of genes potentially regulated by each chromatin signature were identified by selecting all transcription start sites within ±2 Kbp of any peak in the chromatin signature. Peaks with a positive PC coordinate value were analyzed separately from those with a negative value. We provided each gene set to the Homer function annotatePeaks.pl and assessed seven of the provided ontologies: cosmic, kegg, pathwaylnteractionDB, biological_process, smpdb, smart, and reactome. For each gene set in each ontology, we performed a two-sided Fisher’s exact test to calculate the odds of finding the genes from the gene set within 2 Kbp of the chromatin signatures relative to the odds of finding all other genes within the same region. P-values were adjusted for multiple testing by Benjamini-Hochberg95. The resulting statistically significant enrichments (FDR<0.01) were selected and ranked by odds ratio. The Jaccard similarity coefficient96was calculated in a pairwise fashion between the top 100 enriched gene sets. If two gene sets had a Jaccard similarity coefficient > 0.25, then the smaller of the two gene sets was removed from the analysis. The top 25 resulting enrichments for each chromatin signature were reported.
[0131] RNA-seq of organoid samples
[0132] The NEBNext Ultra Directional RNA Library Prep Kit for Illumina with rRNA reduction was used to process the samples. Sample preparation was performed according to the protocol "NEBNext Ultra Directional RNA Library Prep Kit for Illumina" (NEB E7420S / L and NEB #E6310S / L / X). Briefly, rRNA was reduced using a ribonuclease H-based method. Then, fragmentation of the rRNA-reduced RNA and a cDNA synthesis was performed. This was used for ligation with the sequencing adapters and PCR amplification of the resulting product. The quality and yield after sample preparation was measured with the Fragment Analyzer (Advanced Analytical). Clustering and DNA sequencing using the Illumina cBot and HiSeq 2500 were performed according to manufacturer’s protocols. A concentration of 16.0 pM of DNA was used38179217191.1DOCKET: 348358.19102as input. HiSeq control software HCS v2.2.58 was used. Image analysis, base calling, and quality check were performed with the Illumina data analysis pipeline RTAvl.18.64 andBcl2fastq v2.17. On average 67 million reads were obtained per sample.
[0133] Quality and quantity of the total RNA were assessed by the 2100 Bioanalyzer using a Nano chip (Agilent). Total RNA samples having a RIN > 8 were subjected to library generation. Strand-specific libraries were generated using the TruSeq Stranded mRNA sample preparation kit (Illumina Inc., Cat no RS-122-2101 / 2) according to the manufacturer's instructions (Illumina Inc., Cat no 15031047 Rev. E). Briefly, polyadenylated RNA from intact total RNA was purified using oligo(dT) beads. Following purification, the RNA was fragmented, random-primed and reverse transcribed using SuperScript II Reverse Transcriptase (Invitrogen, Cat no 18064-014) with the addition of actinomycin D. Second strand synthesis was performed using polymerase I and ribonuclease H with replacement of dTTP for dUTP. The generated cDNA fragments were 3 '-end adenylated and ligated to Illumina paired-end sequencing adapters and subsequently amplified by 12 cycles of PCR. The libraries were analyzed on a 2100 Bioanalyzer using a 7500 chip (Agilent), diluted and pooled equimolar into a 10 nM multiplex sequencing pool, containing 18 samples per pool. RNA sequencing was performed on an Illumina HiSeq V42500, using a 125 bases paired-end run. On average 26 million reads were obtained per sample. Initial fastq quality control was performed with fastp97v0.20.1 without adapter trimming. Alignment to the hgl9 lift-over of the Gencode87v26 transcript reference and count generation was performed using STAR98v2.7.10a. MultiQC" vl.17 was used to summarize both fastp and STAR reports. Outlier evaluation was performed using both Pearson correlation and PCA on VST normalized gene expression (DESeq2100vl.42.0) across all genes with no outliers or quality issues detected. DESeq2 was used for differential expression.
[0134] Using accessibility loci from the PCI and PC2 chromatin signatures, we subset the results of the differential expression analysis to include only those genes with a peak overlapping the gene promoter or an exon. Each targeted gene was assigned to the “negative” or a “positive” PC coordinate group based on the sign of the PC coordinate of the targeting peak. In the case of genes with multiple targeting peaks, the peak with the maximum absolute value was used for group assignment.
[0135] cfDNA sequencing and processing39179217191.1DOCKET: 348358.19102
[0136] cfDNA sequencing data for these cohorts was generated in previous studies24,101. Briefly, cfDNA was extracted from healthy individuals and 4 cycles of polymerase chain reaction (PCR) were performed prior to sequencing. Paired-end sequencing of these were performed on a NovaSeq6000 at a read length of 101 bp. Samples with less than 10 million reads were excluded from the analysis. The remaining samples had a mean sequencing coverage of 1.5x. cfDNA was extracted from individuals with cancer prior to any treatment and 6 cycles of PCR were performed prior to sequencing. Paired-end sequencing of these were performed on a NovaSeq6000 at a read length of 101 bp. Samples with less than 10 million reads were excluded from the analysis. The remaining samples had a mean sequencing coverage of 5.2x. One third of reads were randomly selected from the fastq of each sample for further analysis in order to match the sequencing depths of the cancer samples to that of the healthy samples.
[0137] cfDNA fragmentation features
[0138] Given that PCR is known to introduce bias into fragment abundance due in part to GC content of the fragments102, cfDNA fragments were reweighted based on fragment length and over / under representation of GC content as previously described24. All cfDNA fragmentation metrics were calculated using these GC weights. Fragments between 100-250bp were subsequently selected for downstream analysis. Three cfDNA fragmentation metrics were considered in this study: fragment coverage, mean fragment length, and fragment end positioning. Fragment coverage was calculated as the total number of cfDNA fragments mapping to any given genomic position. At positions with no fragments, the mean fragment length was assigned as NA and dropped from downstream aggregation calculations. Fragment end positions were calculated as the number of fragments that end within 50bp of a given genomic position. The fragment end position counts were then corrected by the mean fragment length and the mean coverage at that position. These cfDNA fragmentation features were calculated at each consensus peak and the surrounding region (-2.5 Kbp), and tiled into 5 bp bins. BEDOPS103(v2.4.41) was used to map cfDNA fragments to loci, and all calculated using GC weighted values.
[0139] In order to summarize cfDNA fragmentation at individual loci across a cohort, fragments from all samples in the cohort were pooled together. These fragments retained their GC correction weights as initially calculate, but were otherwise treated as a single sample and followed the mapping process as outlined above. The fragmentation metrics wree normalized at each locus40179217191.1DOCKET: 348358.19102by first calculating the mean value for up- and downstream flanking region (1-2.5 Kbp from genomic position 0), then subtracting that value from the entire locus, and finally centering about 0 and scaling to standard deviation 1. In order to summarize cfDNA fragmentation of an individual participant at all selected loci, the median value of each fragmentation metric was calculated at each genomic position across all loci, and the resulting values were centered about 0 and scaled to standard deviation 1.
[0140] Differential accessibility of consensus peaks
[0141] Using the colon organoid consensus peak set consisting of 65,054 peaks, we used DESeq299 to calculate differential accessibility between the PBMC, colon healthy, colon adenoma, and colon cancer sample sets separately. In order to account for possible effects of aneuploidy, these analyses were performed independently on each chromosome arm. The Benjamini -Hochberg false discovery rate was implemented to adjust p-values for multiple hypothesis testing. For samples with multiple replicates passing all quality control thresholds, only the replicate with the highest FRiP (or TSS score in the case of ties) was included in the differential accessibility analysis. Statistically significant differentially accessible peaks were those with a median increase or decrease in read counts within the 500bp peak window greater than two-fold between sample groups and an adjusted p-value < 0.05.
[0142] A peak was assigned to one of six pathophysiological phase-specific peak sets if it was differentially accessible in all colon-colon comparisons (Healthy vs Adenoma, Healthy vs Cancer, Adenoma vs Cancer), as well as upregulated in at least one colon-PBMC comparison (Healthy vs PBMC, Adenoma vs PBMC, Cancer vs PBMC). We undertook a systematic literature search for genes at loci with five or more pathophysiological phase-specific peaks. To determine whether a given gene (e.g. KLF5) had prior evidence of involvement in colorectal adenocarcinoma development, we searched PubMed for the following terms: (KLF5) AND (Colon OR Colorectal OR Intestine OR Rectal OR Gastrointestinal) AND (Cancer OR Tumor OR Neoplasm OR Adenoma OR Adenocarcinoma OR Carcinoma). The results were reviewed and scored for evidence of involvement in colorectal adenocarcinoma development or in cancer development more broadly on a scale from 0 to 3, where a score of 0 signified no prior evidence that the gene is involved, a score of 1 signified minimal prior evidence of involvement, a score of 2 signified41179217191.1DOCKET: 348358.19102more than minimal prior evidence of involvement, and 3 signified substantial prior evidence of involvement, typically including multiple independent studies and mechanistic validation.
[0143] Whole genome sequencing of organoids
[0144] Whole genome sequencing was performed as previously described (Huber, A. R. et al. Improved detection of colibactin-induced mutations by genotoxic E. coli in organoids and colorectal cancer. Cancer Cell 42, 487-496. e6 (2024). Briefly, genomic DNA was isolated from organoid pellets using the Qiagen DNeasy Blood & Tissue kit. DNA sequencing libraries were made with a TruSeq Nano kit (Illumina) from 50 ng of genomic DNA using manufacturers’ instructions. These libraries were sequenced at a depth of 30x using a NovaSeq 6000 (Illumina). Mutation calls were assessed for the cancer driver and chromatin modifier gene sets defined in Heide, et al. (The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733-743 (2022) using Strelka (v2.9.10) (Kim, S. et al. Strelka2: fast and accurate calling of germline and somatic variants.Nat. Methods 15, 591-594 (2018) in germline mode. All variant calls were required to be supported by 4 or more sequencing reads, represent a mutant allele fraction >10%, and not listed as a common variant in dbSNP vl38 or vl56. We analyzed variants annotated by SnpEff (v4.3t)102 as missense, frameshift, stop gained, stop lost, start lost, conservative inframe insertion, conservative inframe deletion, disruptive inframe insertion, disruptive inframe deletion, initiator codon variant, stop retained variant, inframe deletion, inframe insertion, splice acceptor variant, or splice donor variant. To minimize artifactual variants associated with poor mappability regions, we used BL AT (v35)103 to ensure that the lOObp surrounding the variant was required to have no more than a single match in the hgl9 reference genome, and that the same region with the variant present had no matching regions. Stop gain or nonsense mutations that pass these filters were analyzed. All other variants were annotated by OpenCravat (v2.11.1) and additionally required to have >25 mutations in large intestine samples in the COSMIC database (vlOO), be an OncoKB (vl.1.3) Hotspot, or be classified by ClinVar (2025.04.01) as pathogenic or likely pathogenic.
[0145] Identifying transcription factor drivers of chromatin accessibility
[0146] Following our standard pipeline, we trained gkm-SVM on ATAC-seq data after removing promoter peaks within 2kb of an annotated TSS, non-cell specific peaks accessible in more than 30% of ENCODE samples (mostly CTCF binding sites). We removed promoters as they42179217191.1DOCKET: 348358.19102have similar TFBS and are usually active across all cell types . Gkm-SVM uses counts of gapped kmer features under the 300bp sequences centered on the ATAC-seq peaks, as gapped kmers are effective at describing regulatory sequence features and do not require previous knowledge of TFBS. Gkm-SVM is trained to distinguish these 300bp sequences from a negative control set of sequences sample from the reference genome which have been matched for length, GC content, and repeat fraction. The resulting weight vector indicates the importance of each gapped k-mer in classifying these positive and negative control sequence sets. Following the default training method in the gkmSVM-R package, the AUROC was evaluated using 5-fold cross validation (CV) and the average AUROC was computed from 5 CV test sets. All cross-fold validation sets produced similar sets of features and had very similar weight vectors. We extracted TFBS motifs from this weight vector using gkm-PWM109 , which produces a lasso weight reflecting the importance of each motif in explaining the full weight vector (W). Other motif metrics Z and I describe the Z-score for kmers mapping to the motif and the error induced by removing the motif from the list. To produce a reduced set of TFBS motifs describing the differences between any two samples, we trained gkm-SVM on differentially active peaks in two samples as the positive and negative sets53 . We performed an enrichment analysis of the 30 motifs most frequently detected with gkm-PWM, scoring the sets of differentially active peaks with ScanACE. A systematic literature search for involvement in colorectal adenocarcinoma development, or cancer more broadly, was conducted as described above. The motifs of the top seven transcription factors (TCF7L2, TCF3, RUNX, TEAD, KLF4, API, and HNF4A) were used to train a linear model to distinguish positive and negative control sequence sets, as described above, based on ScanACE scores for these motifs. The AUROC was evaluated using 5-fold CV and was compared to the gkm-SVM model.
[0147] Results
[0148] 54 individuals with colorectal cancer or adenomas from the Netherlands Cancer Institute, MC Slotervaart, or the VU University Medical Center32were analyzed. From this group of patients, 59 organoid cultures were established from healthy colon, advanced colorectal adenoma, or primary colorectal cancer tissues and confirmed the epithelial purity of these samples through immunohistochemical staining (FIGS. 7A-7C). An assay for transposase-accessible chromatin was performed using sequencing (ATAC-seq) on 52 of these organoid samples,43179217191.1DOCKET: 348358.19102representing 49 individuals, as well as 10 freshly collected and isolated peripheral blood mononuclear cell (PBMC) control samples. ATAC-seq was performed in duplicate on each sample, totaling 124 transposase treated libraries with 9 additional sequencing replicates. A median of 255 million reads were analyzed at a sequencing depth of 6x across all the ATAC-seq libraries (FIG. 8A). Following stringent quality control analyses, 46 ATAC-seq libraries with low complexity, sequencing quality, signal-to-noise ratio, or replicate concordance were excluded (FIGS. 8A-8F). The remaining 78 libraries from 41 patients, representing 43 organoids with normal, adenoma or adenocarcinoma phenotypes, were further assessed identifying ~35 million peaks across these samples (FIG. 1; Table 1).
[0149] To directly compare ATAC-seq peaks across colorectal samples, a set of 62,890 peaks were identified from locations that were observed to be accessible in a consensus of each of the sets of normal colonic epithelial, adenoma, or carcinoma organoids (FIGS. 2A, 2B). These consensus peaks demarcate accessible regions in healthy, premalignant, or malignant colorectal neoplastic epithelial cells. ATAC-seq read counts at these peaks distinguished colorectal cells from PBMC samples, while preserving the heterogeneity of accessibility in colorectal samples in different stages of cancer development (FIGS. 9A, 9B). Accessibility peaks among samples within the same tissue type (colorectal or PBMC) and disease state (healthy, adenoma, or cancer) were more correlated (e g. colorectal healthy median Pearson r=0.47) than accessibility peaks between tissue types or disease states (median Pearson r=-0.09 between tissues and disease states) (FIG.9B).
[0150] Genome-wide examination of the colorectal consensus ATAC-seq peaks revealed that a subset was located within gene promoters (18.3%), exons (11.8%), or CTCF binding sites (9.3%), with the remainder (60.6%) at distal elements (intergenic or intronic regions) (FIG. 10A). Exons and CTCF binding sites, which comprise less than 4% of the hg!9 reference genome, harbor a disproportionately large share of these consensus peaks. The frequency of consensus peaks was correlated with gene density, with more than three quarters of consensus peaks within 100 Kbp of transcription start sites (TSSs) (FIGS. 10B, 10C).
[0151] To determine whether the chromatin accessibility in organoid samples were similar to those of primary tissues, the 78 colorectal organoid and 17 PBMC samples were compared with44179217191.1DOCKET: 348358.19102210 additional colon, liver, lung andPBMC samples from previous studies'233 35(FIGS. 11 A-l IE, 12A and 12B). Using principal component analysis, it was found that the first two principal components, which collectively explained 42% of the variance across samples, clearly differentiated samples by tissue type (FIG. 3 A), with colorectal organoid peak profiles grouping together with colon tissue peak profiles from previous studies12,33.
[0152] Principal component analysis of the 62,890 colorectal organoid consensus peaks revealed that principal component 1 (PCI, 19% of variance) separated adenomas from other samples, while PC2 (14% of variance) represented a transition from healthy, to adenoma, to cancer (FIG. 3B). To understand whether the PCI and PC2 axes were a consequence of our supervised approach to peak selection, the analysis was repeated using randomized grouping labels in the place of tissue labels and found that PCA of these consensus peaks revealed that samples continued to cluster by their true tissue of origin (FIGS. 12A, 12B). Even with an unsupervised approach to consensus peak selection and dimensionality reduction, the largest source of variation between samples were biological differences in chromatin accessibility between different tissues.
[0153] To identify the consensus peaks that were most closely associated with functional differences in disease state, peaks were ranked by the magnitude of their contribution to each PC (FIG. 13A). Of the 30,556 peaks comprising 90% of the variance of PC2, 53% were accessible in cancer samples and relatively inaccessible in healthy samples, while the remaining peaks showing the opposite pattern (FIG. 3D). Similarly, PCI exhibited a chromatin accessibility signature in which 41% of the selected peaks were accessible in adenoma samples, and the remainder were relatively inaccessible in the healthy and cancer samples (FIG. 3C). The most highly ranked peaks from either component were adjacent to well-known tumor suppressor genes such as NF1, FANCC, and SMAD3 (FIGS. 13A, 13B)36 40.
[0154] It was sought to identify the regulatory factors associated with differences in chromatin accessibility. The peaks comprising each chromatin signature were analyzed for DNA sequence motifs corresponding to transcription factor recognition sequences. It was found that nearly a quarter of these motifs corresponded to regulatory factors which have been reported to play a role in colorectal adenoma or cancer development, including transcription factor 7 TCF7), mothers against decapentaplegic homolog 2 (SMAD2), Runt-related transcription factor 1 (RUNXF), CAMP responsive element binding protein 1 (CREB1) and Hypermethylated-in-cancer45179217191.1DOCKET: 348358.191021 (HIC1) genes, as well as the important signaling pathways they control, such as JEV7 / p-catenin, TGF-P, Hippo, MAPK pathways (FIGS. 3C, 3D, 13C, 13D)18,41 52. The majority of the genes we identified had not been previously detected in nucleosome accessibility studies of colorectal primary tissues17,18, presumably because of the admixture of cell types analyzed in these samples. Additionally, enrichment was found for a variety of regulatory factors not previously shown to be involved in colorectal carcinogenesis. These included, ELF2, a hypoxia response gene that has shown to physically interact with 7?MVA753,54, a transcription factor which promotes tumor metastasis by activating the WNT I P-catenin signaling pathway and EMT in colorectal cancer55, and the zinc finger transcription factor ZNF165, which has shown to be expressed in normal colon tissue and is a SMAD3 cofactor which together promote TGF-P signaling56 58.
[0155] To identify the genes and pathways that may be regulated by these chromatin changes, chromatin accessibility peaks were mapped to their nearby genes and analyzed these for enrichment of gene sets or pathways (FIGS. 14A, 14B). Enrichment for pathways known to be dysregulated in colorectal adenocarcinoma were found, which were implicated in this analysis of regulatory factors, including WNT, Hippo, RAS, and MAPK pathways5'4642'59 <>4. Additionally, enrichment for pathways known to be more generally involved in cancer were found, including related to tissue development, cell structure and organization, cell signaling, cell differentiation, and the maintenance of cell number (FIGS. 3E-3F).
[0156] To determine if differences in chromatin accessibility corresponded with expression changes of nearby genes, RNA-seq was performed on 21 of the analyzed organoids, as well as 7 additional colorectal organoids where RNA samples were available, representing 17 adenoma and 11 cancer samples. Genes containing peaks which were more accessible in adenoma (PCI) were 3.1-fold (95% 0=2.3-4.2; p<0.001, Fisher exact test) more likely to be more highly expressed in adenoma than cancer compared to genes lacking such peaks, while genes with peaks that were more accessible in cancer (PC2) were 3.3-fold (95% CI=2.6-4.1; p<0.001, Fisher exact test) more likely to be more highly expressed in cancer than adenoma compared to genes lacking such peaks (FIGS. 15A-15D). Analyses of genes in pathways enriched for chromatin changes identified expression changes that tended to agree with those predicted by nearby DNA accessibility (FIGS. 15E-15H). Overall, these analyses are consistent with the notion that46179217191.1DOCKET: 348358.19102chromatin accessibility changes observed in the colorectal organoids have functional consequences that result in global downstream changes in gene expression.
[0157] Given the importance of genomic and epigenomic characteristics on cfDNA fragmentation30, it was sought to define the features of cfDNA fragmentation that may be influenced by chromatin accessibility. Because white blood cells are the source of the vast majority of cfDNA in the blood of healthy individuals29, cfDNA fragmentation features were first examined at genomic regions that were accessible in PBMCs. To do so, a combined set of 71,270 consensus peaks were identified for PBMCs and colorectal organoid tissues as described above (FIG. 4A). This consensus set of peaks from PBMC / colorectal samples had similar overall genomic distributions to those of colorectal organoids (FIGS. 16A-16C).
[0158] Features of cfDNA fragmentation were evaluated at these consensus peaks in low-coverage whole genome sequencing data of plasma samples from individuals without cancer. Samples from 251 individuals age 50-75 determined to be free of colorectal cancer were analyzed by fecal immunochemical testing (FIT) screening from the Danish Endoscopy III Project65and COCOS (Netherlands Trial Register ID NTR1829) cohorts66. The sequencing reads from all 251 individuals were aggregated, totalling 10.8 billion reads, performed a fragment-level GC correction for these sequences, and evaluated the key fragmentation features of the subset of cfDNA fragments overlapping the PBMC / colorectal organoid consensus peaks or their surrounding regions (+ / - 2.5 Kbp). The cfDNA features evaluated included sequence coverage, fragment size, and the frequency of fragments ending at specific positions normalized by sequence coverage. When regions of chromatin accessibility were sorted from most to least accessible by the normalized mean ATAC-seq signal of the peak, it was found that all these fragmentation features showed a clear signal corresponding to the position of the consensus peak (FIG. 4B). Higher chromatin accessibility at the center of the peak (-25bp) was correlated with decreased cfDNA sequence coverage (Spearman p= -0.57, p<0.001), increased fragment end frequency (Spearman p=0.42, p<0.001), and a reduction in size of cfDNA fragments (Spearman p=0.42, p<0.001). Analyses of the 10,000 most accessible loci to the 10,000 least accessible loci revealed an increase in average fragment coverage from 1.64x (IQR=1.56-1.73) up to 2.20x (IQR=2.04-2.37, p<0.001, two-sided t-test), with a concomitant reduction in median fragment lengths from 174 bp (IQR=173-176bp) to 169 bp (IQR=168-171, p<0.001, two-sided t-test), and an increase in47179217191.1DOCKET: 348358.19102relative fragment end position coverage from 1.82x (IQR=1.73-1.93) to 2.50x (IQR=2.29-2.74; p<0.001, two-sided t-test)(FIG. 4C). This correspondence between accessibility in PBMCs and cfDNA fragmentation metrics was not observed at a similarly sized set of 60,000 randomly selected loci (FIGS. 17A-17C). These observations provide evidence that the PBMC consensus peaks, as defined here, are effective in identifying genomic regions that are particularly susceptible to fragmentation in the course of white blood cell death that are ultimately reflected in cfDNA.
[0159] In order to gain insight into the effects of genome-wide white blood cell chromatin accessibility on cfDNA of each indivdual, cfDNA sequences were aggregated across all 71,720 consensus peak regions for each of the 251 individuals and their fragmentation characteristics were examined. Aggregating these peaks on a per-sample basis, it was found that the fragmentation features observed for each ATAC-seq peak, when combined across all peaks, formed a robust signal that is reproducible across individuals without cancer (FIG. 4C). As expected, these analyses demonstrate that white blood cell specific chromatin accessibility results in an important determinant of cfDNA fragmentation, and aggregation of these regions genome-wide provide a reproducible metric of cfDNA fragmentation in individuals without disease.
[0160] Having characterized the connection between chromatin accessibility of PBMCs and cfDNA fragmentation in healthy individuals, it was next evaluated whether colorectal organoid specific chromatin accessibility peaks were associated with changes in cfDNA in patients with colorectal cancer. 23,911 colorectal consensus peaks were first identified which were differentially accessible between the colorectal organoid and PBMC samples (FIG. 5A). As chromatin accessibility patterns can be distinct to certain types of genomic features25,3031, the differential consensus peaks were categorized based on four genomic contexts across the genome (promoter, CTCF binding site, exonic regulatory element, distal elements). By aggregating cfDNA metrics according to each of these genomic contexts, cfDNA fragmentation was profiled at all 23,911 colorectal-enriched accessibility peaks as four meta-loci for a single individual. Low-coverage WGS (2x) of cfDNA was analyzed from baseline liquid biopsy draws of 51 individuals from the CAIRO5 cohort (NCT02162563)67, a clinical trial with a cohort of colorectal cancer patients with unresectable liver-only metastases. cfDNA reads were categorized by mapping to each peak and surrounding regions to the four meta-loci across all 251 healthy and 51 colorectal cancer samples. Fragmentation features were calculated for each individual and found differences48179217191.1DOCKET: 348358.19102in median fragment coverage (on average 0.86x lower coverage, 95% CI=0.70-1.0x coverage), length (0.25bp reduction, 95% CI=1.2-0.58 bp) and end positions (1.05x higher coverage, 95% CI=0.96-1.15x coverage) when comparing healthy individuals to those with cancer at distal elements, and that these differences were directionally consistent in other genomic contexts (FIGS. 5B, 18).
[0161] It was evaluated whether these differences in cfDNA fragmentation metrics at these meta-loci could be used to classify individuals as healthy or as having cancer (FIG. 19A). Analysis of the center lObp of each meta-locus found that fragmentation features across the colorectal accessibility peaks tended to be uncorrelated with one another in healthy individuals, while these were strongly correlated in individuals with cancer (FIGS. 19B-19D). Evaluation of all types of meta-loci cfDNA features for their predictive capability revealed that cfDNA fragment coverage had high accuracy for identifying cancer patients across different genomic contexts (AUC 0.85-0.94; FIGS. 5C, 19A). It was hypothesized that for patients with colorectal cancer the observed changes in cfDNA features were a direct result of circulating tumor-derived DNA (ctDNA). To quantify levels of ctDNA, the mutant allele fraction (MAF) of each sample was calculated as determined by droplet digital PCR of KRAS alterations that are presumed to be clonal in these cancers and a strong linear relationship was found between fragmentation metrics and MAFs across all meta-loci, particularly at distal elements (R2=0.69, P<0.001; FIGS. 5D, 19E). These observations across a range of MAF levels provide evidence that changes in these features are a direct result of the presence of tumor derived cfDNA.
[0162] As the main source of observed variation in chromatin accessibility during colorectal cancer development differentiated adenoma from healthy and cancer samples, it was evaluated whether these changes could be detected in cfDNA profiles. To select for ATAC-seq peaks which would be more likely to lead to changes in cfDNA, a stringent differential accessibility analysis (>2 fold change, FDR<0.05) was performed between one or two organoid sample types (e.g. colon healthy, adenoma, or cancer sample groups) and the remaining groups, as well as parallel analysis between the organoid sample types and PMBCs, selecting the differential peaks that were shared between these analyses. It was found that 3,360 the 62,890 consensus peaks (5.3%) were differentially accessible in at least one comparison (FIGS. 20A-20C) and represented49179217191.1DOCKET: 348358.19102six classes of colorectal-specific accessibility loci which could be expected to exhibit informative cfDNA profiles (FIGS. 6A, 20C).
[0163] Focusing on genes with more than five intragenic or nearby peaks that were differentially accessible in the tissues analyzed, it was found that 80% of these peaks were associated with genes known to be involved in colon cancer development. By associating all adenoma-specific peaks with the nearest TSS, the two genes with the highest number of associated peaks were adenomatosis polyposis coli downregulated 1 (APCDD1) and bicaudal C homolog 1 (BICCP) (FIGS. 6B, 21A). APCDD1 has previously been shown to be a target of P-catenin / Tcf4 and is upregulated in colon cancer compared to healthy tissue68,69. While BICC1 was one of the most highly differentially expressed genes in our sample set (19.7x higher expression in adenoma compared to cancer; FIG. 15 A), there is only limited prior evidence of its involvement in colon cancer development70,71, and no evidence of its involvement in the development of colon adenomas. Overall, disease state-specific peaks tended to be far from genes (>10 Kbp), with those having the highest accessibility between lOOkb and 1Mb from nearby genes (FIG. 6C).
[0164] Interestingly, a large number of peaks (783) were observed to be accessible only in adenomas and inaccessible in normal epithelium or cancer tissue. These observations provide evidence of a process of colorectal cancer development where adenomas have a unique accessibility state, with changes that are later reversed in cancer. This model is different from previously proposed models of colorectal cancer development which rely simply on accumulation of genetic or epigenetic changes3,7,72. To validate this model, cfDNA features of peak dynamics from organoids in the separate CRC liquid biopsy cohort were evaluated. The cfDNA fragment coverage profiles of disease-specific distal elements to those of PBMC-specific distal elements were compared and it was found that these peak classes exhibited patterns of cfDNA coverage that were characteristic of the accessibility patterns observed in the organoid cohort. In healthy individuals (n=37; MAF = 0%), and those with cancer (n=7; MAF > 50%) a strong depletion of fragments at the center of the white blood cell-specific meta locus was observed, providing evidence that most cells contributing cfDNA had a strongly accessible peak at the corresponding loci. Increased fragment coverage was observed at the center of the colon-derived meta-loci in healthy individuals consistent with the absence of cfDNA from adenoma or cancer cells in this cohort. In contrast, in the cohort of individuals with colon cancer, fragment coverage depletion50179217191.1DOCKET: 348358.19102was observed in the colon-derived meta-loci, but not adenoma loci, providing evidence of an increase in cfDNA originating from colon cancer or healthy colon in these individuals but not from adenoma specific peaks (FIG. 6D). Taken together, these observed cfDNA profiles are consistent with the adenoma-specific accessibility patterns of the model for colon cancer development.
[0165] DISCUSSION
[0166] Epigenetic modifications of healthy and pre-malignant tissues are a central component to the development and progression of cancer73, and have been shown to be regulated over the development of colorectal cancer17 18. In this study, chromatin accessibility was analyzed in colorectal normal, adenoma, and carcinoma organoids in order to evaluate changes in the colorectal epithelium at different stages of tumor development. New chromatin signatures were analyzed that distinguished healthy, adenoma, and carcinoma organoids from one another, and found that those signatures were associated with transcription factors previously known to play a role in colorectal cancer development as well as new factors whose role remains to be elucidated. The observed chromatin signatures identified genes involved in pathways known to be dysregulated in colorectal cancer development, most notably those involved in cell growth, cell structure and organization, and regulation of cell populations. As it has been unclear whether overall chromatin organization in adenomas differs from that of normal colonic epithelium or from those in adenocarcinomas, these observations deepen our understanding of chromatin changes in colorectal cancer development and suggest that adenomas have a distinct chromatin state, that may change during malignant transformation into an adenocarcinoma chromatin state.
[0167] Separately, this study showed that accessibility signatures of colorectal epithelial tissue and blood cells provided insights into the origins of fragmentation profdes of cfDNA. Characteristics of nucleosomal cfDNA were altered at regions of chromatin accessibility, and changes in these features at colon-specific regions were associated with altered cfDNA profiles and could be evaluated to distinguish individuals with colorectal cancer from those who are healthy. Some of the fragmentation characteristics explored here are similar to those observed at transcription factor binding sites26’31’74, raising the possibility that such changes in these features may be attributable to chromatin accessibility, rather than transcription factor occupancy alone. Analyses of cfDNA fragmentation at regions of chromatin accessibility have provided an in vivo validation of the altered chromatin states of colorectal adenomas and adenocarcinomas.51179217191.1DOCKET: 348358.19102
[0168] Limitations of this study include the use of organoids as a model system that may not fully represent the native context of normal or colon cancer tissues. This concern is mitigated by previous studies that have shown that organoids provide an accurate in vitro representation of the tissue from where they are derived69and the observation that chromatin accessibility peaks of colorectal organoids had similarities to earlier accessibility analyses of impure primary tissues. Another concern may be that the cfDNA analyses were not performed using matched samples from the same patients from which the organoid samples were derived. However, the validation of organoid-informed regions of chromatin accessibility through independent changes in cfDNA features in the circulation of colon cancer patients suggest that this model system reflects important physiologic characteristics of chromatin packaging in vivo. Although the scale of the analyses was modest (78 samples from 41 patients), these nevertheless represent the largest patient cohort and the largest sequencing analyses to date (255 million reads per sample, and a total of 24 billion reads sequenced) of chromatin accessibility dynamics of colorectal adenomas and cancer. Larger studies will be needed to validate the changes in cfDNA resulting from alterations in chromatin accessibility and to determine whether these may be useful in clinical settings.
[0169] This work provides a significant step in understanding chromatin dysregulation in colorectal adenomas and carcinoma. These efforts lay the groundwork for future mechanistic studies into the roles of specific transcription factors in colorectal cancer development, and for the application of chromatin signatures for cfDNA-based early cancer detection and monitoring.Table 11 Colon and samples analyzed with ATAOSeq„ , . . Number of ATAC-seq Sample type Tissue Disease state „ . , Samples analyses Organoid Colon Healthy 7 11 Organoid Colon Adenoma 21 41 Organoid Colon Cancer 15 26 Primary Tissue PBMC Healthy 10 17
[0170] REFERENCES1. Bray, F. et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA. Cancer J. Clin. 74, 229-263 (2024).52179217191.1DOCKET: 348358.191022. Lockhart-Mummery, J. & Dukes, C. The pre-cancerous changes in the rectum and colon. Surg Gynecol Obstet 46, 591-596 (1928).3. Muto, T., Bussey, H. J. & Morson, B. C. The evolution of cancer of the colon and rectum. Cancer 36, 2251-2270 (1975).4. Stryker, S. J. et al. Natural history of untreated colonic polyps. Gastroenterology’ 93, 1009- 1013 (1987).5. Fearon, E. R. & Vogelstein, B. A genetic model for colorectal tumorigenesis. Cell 61, 759-767 (1990).6. Toyota, M. et al. CpG island methylator phenotype in colorectal cancer. Proc. Natl. Acad. Set. U. S. A. 96, 8681-8686 (1999).7. Hermsen, M. et al. Colorectal adenoma to carcinoma progression follows multiple pathways of chromosomal instability. Gastroenterology 123, 1109-1119 (2002).8. Sjoblom, T. etal. The consensus coding sequences of human breast and colorectal cancers. Science 314, 268-274 (2006).9. Wood, L. D. etal. The genomic landscapes of human breast and colorectal cancers. Science 318, 1108-1113 (2007).10. Carvalho, B. et al. Multiple putative oncogenes at the chromosome 20q amplicon contribute to colorectal adenoma to carcinoma progression. Gut 58, 79-89 (2009).11. Jones, S. et al. Personalized genomic analyses for cancer mutation discovery and interpretation. Sci. Transl. Med. 7, 283ra53 (2015).12. Corces, M. R. et al. The chromatin accessibility landscape of primary human cancers. Science 362, eaavl898 (2018).13. Song, L. et al. Open chromatin defined by DNasel and FAIRE identifies regulatory elements that shape cell-type identity. Genome Res. 21, 1757-1767 (2011).14. Flavahan, W. A., Gaskell, E. & Bernstein, B. E. Epigenetic plasticity and the hallmarks of cancer. Science 357, eaal2380 (2017).15. Buenrostro, J. D., Giresi, P. G., Zaba, L. C., Chang, H. Y. & Greenleaf, W. J. Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat. Methods 10, 1213-1218 (2013).16. Buenrostro, J. D. et al. Single-cell chromatin accessibility reveals principles of regulatory variation. Nature 523, 486-490 (2015).17. Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733-743 (2022).18. Becker, W. R. et al. Single-cell analyses define a continuum of cell state and composition changes in the malignant transformation of polyps to colorectal cancer. Nat. Genet. 54, 985-995 (2022).19. Cusanovich, D. A. et al. Multiplex single cell profiling of chromatin accessibility by combinatorial cellular indexing. Science 348, 910-914 (2015).20. Haber, D. A. & Velculescu, V. E. Blood-based analyses of cancer: circulating tumor cells and circulating tumor DNA. Cancer Discov. 4, 650-661 (2014).21. Cristiano, S. et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385-389 (2019).22. Phallen, J. et al. Early Noninvasive Detection of Response to Targeted Therapy in NonSmall Cell Lung Cancer. Cancer Res. 79, 1204-1213 (2019).53179217191.1DOCKET: 348358.1910223. Anagnostou, V. et al. Dynamics of Tumor and Immune Responses during Immune Checkpoint Blockade inNon-Small Cell Lung Cancer. Cancer Res. 19, 1214-1225 (2019).24. Mathios, D. et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat. Commun. 12, 5060 (2021).25. Zhu, G. et al. Tissue-specific cell-free DNA degradation quantifies circulating tumor DNA burden. Nat. Commun. 12, 2229 (2021).26. Foda, Z. H. et al. Detecting liver cancer using cell-free DNA fragmentomes. Cancer Discov. CD-22-0659 (2022) doi:10.1158 / 2159-8290.CD-22-0659.27. van ’t Erve, I. et al. Metastatic Colorectal Cancer Treatment Response Evaluation by UltraDeep Sequencing of Cell-Free DNA and Matched White Blood Cells. Clin. Cancer Res. Off. J. Am. Assoc. Cancer Res. 29, 899-909 (2023).28. Chung, D. C. et al. A Cell-free DNA Blood-Based Test for Colorectal Cancer Screening. N. Engl. J. Med. 390, 973-983 (2024).29. Moss, J. et al. Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA in health and disease. Nat. Commun. 9, 5068 (2018).30. Snyder, M. W ., Kircher, M., Hill, A. J., Daza, R. M. & Shendure, J. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. Cell 164, 57-68 (2016).31. Ulz, P. et al. Inference of transcription factor binding from cell-free DNA enables tumor subtype prediction and early detection. Nat. Commun. 10, 4666 (2019).32. Martens-de Kemp, S. R. et al. Overexpression of the miR-17-92 cluster in colorectal adenoma organoids causes a carcinoma-like gene expression signature. Neoplasia N. Y. N 32, 100820 (2022).33. ENCODE Project Consortium. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57-74 (2012).34. Marquez, E. J. et al. Sexual-dimorphism in human immune system aging. Nat. Commun.11, 751 (2020).35. Currin, K. W. et al. Genetic effects on liver chromatin accessibility identify disease regulatory variants. Am. J. Hum. Genet. 108, 1169-1189 (2021).36. Li, Y. etal. Somatic mutations in the neurofibromatosis 1 gene in human tumors. Cell 69, 275-281 (1992).37. Esteban- Jurado, C. etal. The Fanconi anemia DNA damage repair pathway in the spotlight for germline predisposition to colorectal cancer. Eur. J. Hum. Genet. EJHG 24, 1501-1505 (2016).38. Zhu, Y ., Richardson, J. A., Parada, L. F. & Graff, J. M. Smad3 mutant mice develop metastatic colorectal cancer. Cell 94, 703-714 (1998).39. Sodir, N. M. et al. Smad3 deficiency promotes tumorigenesis in the distal colon of ApcMin / + mice. Cancer Res. 66, 8430-8438 (2006).40. Fleming, N. I. etal. SMAD2, SMAD3 and SMAD4 mutations in colorectal cancer. Cancer Res. 73, 725-735 (2013).41. Korinek, V. et al. Constitutive transcriptional activation by a beta-catenin-Tcf complex in APC- / - colon carcinoma. Science 275, 1784-1787 (1997).42. Morin, P. J. etal. Activation of beta-catenin-Tcf signaling in colon cancer by mutations in beta-catenin or APC. Science 275, 1787-1790 (1997).43. Roose, J. et al. Synergy between tumor suppressor APC and the beta-catenin-Tcf4 target Tcfl. Science 285, 1923-1926 (1999).54179217191.1DOCKET: 348358.1910244. Fijneman, R. J. A. et al. Runxl is a tumor suppressor gene in the mouse gastrointestinal tract. Cancer Set. 103, 593-599 (2012).45. Slattery, M. L. etal. Associations between genetic variation in RUNX1, RUNX2, RUNX3, MAPK1 and eIF4E and riskof colon and rectal cancer: additional support for a TGF-P-signaling pathway. Carcinogenesis 32, 318-326 (2011).46. Hamamoto, T. et al. Compound disruption of smad2 accelerates malignant progression of intestinal tumors in ape knockout mice. Cancer Res. 62, 5955-5961 (2002).47. Wales, M. M. etal. p53 activates expression of HIC-1, a new candidate tumour suppressor gene on 17pl3.3. Nat. Med. 1, 570-577 (1995).48. Ahuja, N. et al. Association between CpG island methylation and microsatellite instability in colorectal cancer. Cancer Res. 57, 3370-3374 (1997).49. Mohammad, H. P. et al. Loss of a single Hicl allele accelerates polyp formation in Apc(A716) mice. Oncogene 30, 2659-2669 (2011).50. Najdi, R. etal. A Wnt kinase network alters nuclear localization of TCF-1 in colon cancer. Oncogene 28, 4133-4146 (2009).51. Druliner, B. R. et al. Molecular characterization of colorectal adenomas with and without malignancy reveals distinguishing genome, transcriptome and methylome alterations. Sci. Rep. 8, 3161 (2018).52. Liu, L. et al. Genome-wide DNA methylation profiling and gut flora analysis in intestinal polyps patients. Fair. J. Gastroenterol. Hepatol. 33, 1071-1081 (2021).53. Christensen, R. A., Fujikawa, K., Madore, R., Oettgen, P. & Varticovski, L. NERF2, a member of the Ets family of transcription factors, is increased in response to hypoxia and angiopoietin-1: a potential mechanism for Tie2 regulation during hypoxia. J. Cell. Biochem. 85, 505-515 (2002).54. Cho, J.-Y. et al. Isoforms of the Ets transcription factor NERF / ELF-2 physically interact with AML1 and mediate opposing effects on AML 1 -mediated transcription of the B cell-specific blk gene. J. Biol. Chem. 279, 19512-19522 (2004).55. Li, Q. et al. RUNX1 promotes tumour metastasis by activating the Wnt / p-catenin signalling pathway and EMT in colorectal cancer. J. Exp. Clin. Cancer Res. CR 38, 334 (2019).56. Dong, X.-Y., Yang, X.-A., Wang, Y.-D. & Chen, W.-F. Zinc-finger protein ZNF165 is a novel cancer-testis antigen capable of eliciting antibody response in hepatocellular carcinoma patients. Br. J. Cancer 91, 1566-1570 (2004).57. Maxfield, K. E. et al. Comprehensive functional characterization of cancer-testis antigens defines obligate participation in multiple hallmarks of cancer. Nat. Commun. 6, 8840 (2015).58. Gibbs, Z. A. et al. The testis protein ZNF165 is a SMAD3 cofactor that coordinates oncogenic TGFp signaling in triple-negative breast cancer. eLife 9, e57679 (2020).59. Roose, J. et al. Synergy between tumor suppressor APC and the beta-catenin-Tcf4 target Tcfl. Science 285, 1923-1926 (1999).60. Oshima, H., Oshima, M., Kobayashi, M., Tsutsumi, M. & Taketo, M. M. Morphological and molecular processes of polyp formation in Apc(delta716) knockout mice. Cancer Res. 57, 1644-1649 (1997).61. Barry, E. R. et al. Restriction of intestinal stem cell expansion and the regenerative response by YAP. Nature 493, 106-110 (2013).62. Cheung, P. etal. Regenerative Reprogramming of the Intestinal Stem Cell State via Hippo Signaling Suppresses Metastatic Colorectal Cancer. Cell Stem Cell 27, 590-604. e9 (2020).55179217191.1DOCKET: 348358.1910263. Bos, J. L. et al. Prevalence of ras gene mutations in human colorectal cancers. Nature 327, 293-297 (1987).64. Fang, J. Y. & Richardson, B. C. The MAPK signalling pathways and colorectal cancer. Lancet Oncol. 6, 322-327 (2005).65. Rasmussen, L. et al. Protocol Outlines for Parts 1 and 2 of the Prospective Endoscopy III Study for the Early Detection of Colorectal Cancer: Validation of a Concept Based on Blood Biomarkers. JMIRRes. Protoc. 5, el 82 (2016).66. Stoop, E. M. et al. Participation and yield of colonoscopy versus non-cathartic CT coIonography in population-based screening for colorectal cancer: a randomised controlled trial. Lancet Oncol. 13, 55-64 (2012).67. Huiskens, J. et al. Treatment strategies in colorectal cancer patients with initially unresectable liver-only metastases, a study protocol of the randomised phase 3 CAIRO5 study of the Dutch Colorectal Cancer Group (DCCG). BMC Cancer 15, 365 (2015).68. Takahashi, M. et al. Isolation of a novel human gene, APCDD1, as a direct target of the beta-Catenin / T-cell factor 4 complex with probable involvement in colorectal carcinogenesis. Cancer Res. 62, 5651-5656 (2002).69. van de Wetering, M. et al. Prospective derivation of a living organoid biobank of colorectal cancer patients. Cell 161, 933-945 (2015).70. Arai, Y. et al. Fibroblast growth factor receptor 2 tyrosine kinase fusions define a unique molecular subtype of cholangiocarcinoma. Hepatol. Baltim. Md59, 1427-1434 (2014).71. Lv, J., Wang, J., Shang, X., Liu, F. & Guo, S. Survival prediction in patients with colon adenocarcinoma via multi -omics data integration using a deep learning algorithm. Biosci. Rep. 40, BSR20201482 (2020).72. Issa, J.-P. CpG island methylator phenotype in cancer. Nat. Rev. Cancer 4, 988-993 (2004).73. Feinberg, A. P., Koldobskiy, M. A. & Gondor, A. Epigenetic modulators, modifiers and mediators in cancer aetiology and progression. Nat. Rev. Genet. 17, 284-299 (2016).74. Doebley, A.-L. etal. A framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. Nat. Commun. 13, 7475 (2022).75. Corces, M. R. et al. An improved ATAC-seq protocol reduces background and enables interrogation of frozen tissues. Nat. Methods 14, 959-962 (2017).76. Smith, J. P. et al. PEP AT AC: an optimized pipeline for ATAC-seq data analysis with serial alignments. NAR Genomics Bioinforma. 3, IqablOl (2021).77. Jiang, H., Lei, R., Ding, S.-W. & Zhu, S. Skewer: a fast and accurate adapter trimmer for next-generation sequencing paired-end reads. BMC Bioinformatics 15, 182 (2014).78. Andrews, R. M. et al. Reanalysis and revision of the Cambridge reference sequence for human mitochondrial DNA. Nat. Genet. 23, 147 (1999).79. Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357-359 (2012).80. Stolarczyk, M., Reuter, V. P., Smith, J. P., Magee, N. E. & Sheffield, N. C. Refgenie: a reference genome resource manager. GigaScience 9, gizl49 (2020).81. Li, H. et al. The Sequence Alignment / Map format and SAMtools. Bioinforma. Oxf Engl.25, 2078-2079 (2009).82. Zhang, Y. etal. Model-based analysis of ChlP-Seq (MACS). Genome Biol. 9, R137 (2008).83. Amemiya, H. M., Kundaje, A. & Boyle, A. P. The ENCODE Blacklist: Identification of Problematic Regions of the Genome. Sci. Rep. 9, 9354 (2019).56179217191.1DOCKET: 348358.1910284. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinforma. Oxf. Engl. 26, 841-842 (2010).85. Landt, S. G. et al. ChlP-seq guidelines and practices of the ENCODE and modENCODE consortia. Genome Res. 22, 1813-1831 (2012).86. Grandi, F. C., Modi, H., Kampman, L. & Corces, M. R. Chromatin accessibility profiling by ATAC-seq. Nat. Protoc. 17, 1518-1552 (2022).87. Frankish, A. et al. GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Res. 47, D766-D773 (2019).88. Heinz, S. et al. Simple combinations of lineage-determining transcription factors prime cis-regulatoiy elements required for macrophage and B cell identities. Mol. Cell 38, 576-589 (2010).89. Huber, W. et al. Orchestrating high-throughput genomic analysis with Bioconductor. Nat. Methods 12, 115-121 (2015).90. Gu, Z., Gu, L., Eils, R., Schlesner, M. & Brors, B. circlize Implements and enhances circular visualization in R. Bioinforma. Oxf. Engl. 30, 2811-2812 (2014).91. R Core Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing (2021).92. Kassambara, A. & Mundt, F. Factoextra: Extract and Visualize the Results of Multivariate Data Analyses. R Package Version 1.0.7. (2020).93. Gu, Z., Eils, R. & Schlesner, M. Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinforma. Oxf. Engl. 32, 2847-2849 (2016).94. Venables, W. N. & Ripley, B. D. Modern Applied Statistics with S. (Springer, New York, 2002).95. Benjamini, Y. & Hochberg, Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. J. R. Stat. Soc. Ser. B Methodol. 57, 289-300 (1995).96. Jaccard, P. Nouvelles recherches sur la distribution florale. Bull. Societe Vaudoise Sci. Nat.44, 223-270 (1908).97. Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinforma. Oxf. Engl. 34, i884— i890 (2018).98. Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinforma. Oxf. Engl. 29, 15-21 (2013).99. Ewels, P., Magnusson, M., Lundin, S. & Kaller, M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinforma. Oxf. Engl. 32, 3047-3048 (2016).100. Love, M. I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 15, 550 (2014).101. van ’t Erve, I. et al. Cancer treatment monitoring using cell-free DNA fragmentomes. Nat. Commun. 15, 8801 (2024).102. Benjamini, Y. & Speed, T. P. Summarizing and correcting the GC content bias in high-throughput sequencing. Nucleic Acids Res. 40, e72 (2012).103. Neph, S. et al. BEDOPS: high-performance genomic feature operations. Bioinforma. Oxf Engl. 28, 1919-1920 (2012).OTHER EMBODIMENTS57179217191.1DOCKET: 348358.19102
[0171] From the foregoing description, it will be apparent that variations and modifications may be made to the disclosure described herein to adopt it to various usages and conditions. Such embodiments are also within the scope of the following claims.
[0172] All citations to sequences, patents and publications in this specification are herein incorporated by reference to the same extent as if each independent patent and publication was specifically and individually indicated to be incorporated by reference. By their citation of various references in this document, Applicants do not admit any particular reference is “prior art” to their disclosure.58179217191.1
Claims
DOCKET: 348358.19102What is claimed:
1. A method of diagnosing and treating a subject diagnosed with cancer comprising:a) producing a genome-wide tissue specific chromatin accessibility profile from a sample obtained from a subject;b) comparing the genome-wide tissue specific chromatin accessibility profile of the subject to a normal healthy control sample;c) identifying changes in the genome-wide tissue specific chromatin accessibility profile of the subject and correlating these changes to a genome-wide tissue specific cancer chromatin accessibility profiles, thereby diagnosing the subject with cancer; and,d) treating the subject with a cancer specific therapy.
2. The method of claim 1, wherein the genome-wide tissue specific chromatin accessibility profile comprises identifying genomic loci comprising tissue-specific signatures of chromatin accessibility.
3. The method of claim of claim 2, wherein the genetic loci are identified by conducting assays for transposase accessible chromatin by sequencing (ATAC-seq) thereby generating libraries comprising peaks across the sample.
4. The method of claim 3, wherein the ATAC-seq generated libraries exclude libraries with low complexity, sequencing quality, signal-to-noise ratio, or replicate concordance.
5. The method of claim 3, wherein overlap of each location of peaks across samples identify accessible consensus peaks.
6. The method of claim 5, wherein the accessible consensus peaks obtained from the subject’s samples are compared to accessible consensus peaks obtained from cancer samples and normal samples.
7. The method of claim 6, wherein the accessible consensus peaks are diagnostic of the type of cancer.
8. The method of claim 7, wherein the accessible consensus peaks among samples within the same tissue type or disease state comprise a higher correlation or similarity.59179217191.1DOCKET: 348358.191029. The method of claim 7, wherein the accessible consensus peaks among normal samples within the same tissue type comprise a higher correlation or similarity.
10. The method of claim 9, wherein the accessible consensus peaks among normal samples are inaccessible in healthy samples as compared to the accessible consensus peaks among cancer samples.
11. The method of claim 10, wherein identifying genes and pathways regulated by chromatin changes comprises mapping accessible consensus peaks to nearby genes and analyzed these for enrichment of gene sets or pathways.
12. The method of claim 1, wherein the subject is diagnosed with colorectal cancer.
13. The method of any one of claims 1-11, further comprising producing a cell-free DNA (cfDNA) fragmentation profile at tissue specific chromatin accessibility regions in patients with and without cancer.
14. The method of claim 13, wherein the cfDNA fragmentation profiles comprise cfDNA fragmentation profiles of consensus peaks.
15. The method of claim 14, wherein evaluation of cfDNA features to produce a cfDNA fragmentation profile comprises evaluating sequence coverage, fragment size, and frequency of fragments ending at specific positions normalized by sequence coverage.
16. The method of claim 15, wherein higher chromatin accessibility at a center of consensus peaks comprises decreased cfDNA sequence coverage, increased fragment end frequency, a reduction in size of cfDNA fragments or combinations thereof.
17. The method of any one of claims 13-16, wherein cfDNA fragmentation profiles of healthy subjects are similar.
18. The method of any one of claims 13-16, wherein cfDNA fragmentation profiles of subjects with cancer are similar based on cancer type.
19. The method of claim 18, wherein a subject with colorectal cancer comprises a cfDNA fragmentation profile distinct to other cancer types.60179217191.1DOCKET: 348358.1910220. A method of diagnosing and treating colorectal cancer in a subject, comprising a) producing a cell-free DNA (cfDNA) fragmentation profde at tissue specific chromatin accessibility regions in patients with and without cancer, wherein the subject with colorectal cancer comprises a cfDNA fragmentation profile distinct to other cancer types; andb) treating the subject with colorectal specific therapeutics.
21. The method of claim 20, wherein the cfDNA fragmentation profiles comprise cfDNA fragmentation profiles of consensus peaks.
22. The method of claim 20, wherein evaluation of cfDNA features to produce a cfDNA fragmentation profile comprises evaluating sequence coverage, fragment size, and frequency of fragments ending at specific positions normalized by sequence coverage.
23. The method of claim 22, wherein higher chromatin accessibility at a center of consensus peaks comprises decreased cfDNA sequence coverage, increased fragment end frequency, a reduction in size of cfDNA fragments or combinations thereof.
24. The method of any one of claims 20-23, wherein cfDNA fragmentation profiles of healthy subjects are similar.
25. The method of any one of claims 20-23, wherein cfDNA fragmentation profiles of subjects with cancer are similar based on cancer type.179217191.1