Methods for analyzing chromatin architecture in tissue to boost detection of cancer associated signals in cell-free DNA
A neural network-based method integrates chromatin conformation and genetic sequence data to profile tumor-specific epigenetic landscapes, addressing the neglect of epigenetic changes in MRD detection, enhancing cancer signal detection in cell-free DNA and improving personalized cancer therapies.
Patent Information
- Application Number
- PCT/US2025/031047
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-27
- Publication Date
- 2025-12-04
AI Technical Summary
Current methods for detecting minimal residual disease (MRD) in cancer patients focus on tumor-specific genomic alterations, neglecting cancer-specific epigenetic changes that occur in conjunction with genetic alterations, leading to a need for methods that profile epigenetic alterations in tumor tissue to enhance detection of cancer-associated signals in cell-free DNA.
A computer-implemented method using a neural network to analyze chromatin conformation data from tumor tissue and genetic sequence information from cell-free DNA, integrating epigenetic and genetic features to determine cancer recurrence, with potential input features including tumor-specific somatic genetic variation and methylation patterns.
Enhances the detection of cancer-associated signals in cell-free DNA by profiling tumor-specific epigenetic landscapes, enabling data-driven healthcare options and uncovering epigenetic changes during disease development, aging, and health, thereby improving personalized adjuvant therapies.
Smart Images

Figure US2025031047_04122025_PF_FP_ABST
Abstract
Description
METHODS FOR ANALYZING CHROMATIN ARCHITECTURE IN TISSUE TO BOOST DETECTION OF CANCER ASSOCIATED SIGNALS IN CELL-FREE DNACROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of, and relies on the filing date of, U.S. provisional patent application number 63 / 654,631, which was filed May 31, 2024, the entire disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION
[0002] The present disclosure relates generally to methods for detecting minimal residual disease (MRD) using tissue profiling methods to boost detection of cancer associated signals in cell-free DNABACKGROUND
[0003] Minimal residual disease (MRD) refers to the small fraction of cancer cells that may persist even after a patient exhibits a complete remission of the disease following treatment. Despite standard clinical measures indicating a successful response to therapy, MRD underscores the importance of sensitive detection methods such as next-generation sequencing (NGS) in combination with library preparation strategies to profile the genome and its epigenetic modifications, which can be pivotal in identifying and monitoring these residual cancer cells.
[0004] Increasing clinical evidence suggests that circulating tumor DNA (ctDNA) is a valuable biomarker for detecting MRD in cancer patients. For example, clinical trials, such as the lung TRACERx study, have utilized ctDNA profiling following surgery to identify patients with early- stage NSCLC who are at high risk for disease recurrence, demonstrating the potential of ctDNA as an adjuvant biomarker (Abbosh et al., 2020). Additionally, in colorectal cancer, ctDNA analysis post-curative intent surgery has been shown to be a strong prognostic marker of recurrence-free survival, emphasizing its role as a significant and independent predictor of recurrence. Therefore, the use of ctDNA for MRD detection could enable oncologists to better personalize adjuvant therapies, potentially sparing patients from unnecessary toxic treatments.
[0005] Current methods to profile ctDNA in the context of MRD have largely focused on tumorspecific genomic alterations such as nucleic acid mutations that can be traced in the blood after treatment. These assays first profile genetic mutations on the primary tumor tissue to identify somatic clonal variants, which are then followed using targeted resequencing in plasma.
[0006] However, these methods ignore cancer specific epigenetic changes that occur in conjunction with genetic alterations. Therefore, there is a need for methods that profile epigenetic alterations in tumor tissue to boost detection of cancer associated signals in cell free DNA based MRD assays.SUMMARY
[0007] The present disclosure provides methods and systems to profile tumor specific epigenetic landscapes in tissue to boost detection of cancer associated signals in cell-free DNA based MRD assays. The methods can enable delivery of data-driven healthcare options to subjects experiencing a disease. Additionally, the methods may help uncover the epigenetic and chromatin architecture changes that occur during development, aging, health, and disease.
[0008] In one aspect, the disclosure provides a computer-implemented method configured to use a previously generated neural network, trained to determine cancer recurrence in a patient, comprising: (a) extracting a plurality of molecules from a tumor tissue sample obtained from the patient to acquire chromatin conformation data from the tumor tissue sample;(b)extracting a plurality of cell-free DNA (cfDNA) from a blood sample obtained from the patient, wherein the blood sample is collected from the same patient at a later timepoint following tumor resection;(c) sequencing a portion of the cfDNA molecules extracted from the blood sample, to acquire sequencing reads, wherein the plurality of sequencing reads comprise genetic sequence information from the cfDNA molecules; (d) inputting a plurality of features to the previously trained neural network, wherein the plurality of features comprises at least a portion of the chromatin conformation data from the tissues sample and at least a portion of the genetic sequence information from the cfDNA molecules; and (e) outputting from the classifier a determination of whether or not the cancer reoccurred in the cancer patient.
[0009] In some embodiments, the tumor tissue sample from the subject is genetically assayed to further obtain tumor specific somatic genetic variation. In some embodiments, the input features for the neural network comprise at least a portion of the somatic genetic variation obtained from the tumor. In some embodiments, the tumor tissue sample from the subject is epigenetically assayed to further obtain tumor specific methylation patterns. In some embodiments, the input features for the neural network comprise at least a portion of the information of tumor specific methylation patterns. In some embodiments, the input features for the neural network comprise at least a portion of the tumor specific somatic genetic variation and at least a portion of the tumor specific methylation patterns. In some embodiments, prior to sequencing, the cfDNA moleculesare enriched for genomic regions that comprise tumor specific somatic genetic variation obtained in the genetic assay.
[0010] In another aspect, the disclosure provides a neural network for determining the reoccurrence of a tumor in a subject previously treated to remove the tumor, wherein the neural network has been trained with at least one of a plurality of feature vectors, each vector comprising at least one of: (i). chromatin conformation data obtained from the tumor prior to resecting the tumor from the subject, (ii). tumor specific somatic genetic variation data obtained from the tumor prior to resecting the tumor from the subject, (iii). DNA sequence information from a cell-free DNA (cfDNA) sample obtained at a later time point after resecting the tumor from the subject; and (iv). the tumor recurrence status of the patient at the time the cfDNA sample was obtained.
[0011] In an additional aspect, the disclosure provides a method for determining a plurality of epigenetic landscapes in a sample from a subject having a disease, comprising: (a) determining a plurality of classes of molecules in the sample using a plurality of assays, wherein the sample comprises cell-free DNA (cfDNA); (b) generating a set of target regions by capturing each of the plurality of classes of molecules isolated from the sample of the subject, wherein each target region of the set of target regions spans an epigenetic landscape from the plurality of epigenetic landscapes known to be associated with the disease in the subject; and (c) for each of the plurality of epigenetic landscapes, determining the sequence of at least a segment of each target region for each of the plurality of classes of molecules, wherein the segment comprises an epigenetic biomarker or a portion thereof, thereby determining the plurality of epigenetic landscapes in the sample from the subject having a disease.
[0012] In some embodiments, the method further comprises isolating the plurality of classes of molecules from a tumor sample from the subject, and identifying the plurality of epigenetic landscapes in the tumor before determining the sequence in the sample of blood. In some embodiments, the epigenetic landscapes further comprise information layers comprising biochemical states of each of the plurality of classes of molecules. In some embodiments, the biochemical states comprise cytosine methylation, transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. In some embodiments, the epigenetic landscapes in theplurality of epigenetic landscapes comprises higher-order chromatin states associated with the disease in the subject.
[0013] In additional embodiments, the sample comprises cell-free DNA molecules isolated from blood.
[0014] In some embodiments, the disease is cancer comprising at least one of adenocarcinoma, basal cell carcinoma, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, cholangiocarcinoma, colorectal cancer, endometrial cancer, esophageal cancer, gallbladder cancer, gastric cancer, germ cell tumors, glioma, head and neck cancer, hepatocellular carcinoma, Kaposi sarcoma, kidney cancer, lip and oral cavity cancer, liver cancer, lung cancer, melanoma, mesothelioma, neuroendocrine tumors, ovarian cancer, pancreatic cancer, penile cancer, prostate cancer, sarcoma, skin cancer, small cell lung cancer, squamous cell carcinoma, stomach cancer, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vulvar cancer.
[0015] In some embodiments, the plurality of classes of molecules comprises modified DNA, a DNA-protein complex, DNA-histone complex, and / or mRNA.
[0016] In some embodiments, the plurality of assays comprise TAM-ChIP, CUT&RUN, CUT&Tag, ChlP-seq, and / or RNA-seq.
[0017] In some embodiments, the determining the sequence of at least a segment of each target region for each of the plurality of classes of molecules comprises sequencing the DNA or mRNA molecules to generate sequencing reads.
[0018] In some embodiments, the epigenetic biomarker further comprises a plurality of fragmentome biomarkers. In some embodiments, the fragmentome biomarkers comprise fragment length score, fragment end-densities, the ratio of short (100-150 bp) to long (151-220 bp) fragments, and / or fragment length signatures.
[0019] In some embodiments, the assays are performed at a plurality of time points. In some embodiments, the subject has not undergone surgery to remove disease. In some embodiments, the subject has undergone surgery to remove disease.
[0020] In additional embodiments, the method further comprises administering a therapy to the individual, where the therapy is effective in treating the disease having at least one of the epigenetic landscapes from the plurality of epigenetic landscapes determined in the sample from the subject. In some embodiments, the therapy comprises an epigenetic therapy, targeting DNA methylation and histone modification mechanisms. In some embodiments, the epigenetic therapy comprises azacitidine, decitabine, vorinostat, and / or romidepsin. In some embodiments, the therapycomprises a BET protein inhibitors, a histone methyltransferase inhibitors, a lysine-specific demethylase inhibitors, and / or Bromodomain inhibitors.
[0021] In some embodiments, the capturing comprises hybrid capture using a set of capture probes, multiplexed PCR amplification, or multiplexed qPCR using sets of primers.
[0022] In an additional aspect, the disclosure provides a method to remove noise in epigenetic datasets using chromatin interaction data comprising (a) providing a first sample from a subject comprising cellular genomic DNA and processing the first sample to determine a set of chromatin interaction partners in the cellular genomic DNA of the subject; (b) providing a second sample from the subject comprising cell free DNA(cfDNA) molecules processing the second sample to determine a set of features in the cfDNA molecules of the subject, wherein the features comprise genetic, epigenetic, and / or fragmentomic markers; and (c) applying a machine learning model trained on characteristics and relationships between interaction partners to identify and correct discordant features in the second sample comprising cfDNA, thereby remove noise in epigenetic datasets using chromatin interaction data.
[0023] In another aspect, the disclosure provides a method for enriching disease associated signals comprising (a) providing a first sample from a subject comprising cellular genomic DNA and processing the first sample to determine a set of chromatin interaction partners in the cellular genomic DNA of the subject; (b) providing a second sample from the subject comprising cell-free DNA (cfDNA) molecules processing the second sample to determine a set of features in the cfDNA molecules of the subject, wherein the features comprise genetic, epigenetic, and / or fragmentomic markers; and (c) annotating the set of chromatin interaction partners in the first sample with one or more features in the cfDNA molecules in the second sample, thereby integrating information from the first sample comprising cellular genomic DNA and the second sample comprising cell-free DNA (cfDNA) molecules from the subject to enrich disease associated signals in the subject based on the combined genetic, epigenetic and fragmentomic landscape of the interaction partners, wherein the landscape is indicative of a biological function or disfunction.
[0024] In yet another aspect, the disclosure provides a method for determining disease associated signals comprising: (a) providing a first sample from a subject comprising cellular genomic DNA and processing the sample by (i) linking the cellular genomic DNA to produce linked genomic DNA; (ii) contacting the linked genomic DNA with one or more reagents that fragment and attach at least one label to the genomic DNA to produce labeled nucleic acid molecules; (iii) contactingthe labeled nucleic acid molecules with one or more reagents that proximity ligates the labeled nucleic acid molecules to generate ligated nucleic acid molecules; (iv) shearing the ligated nucleic acid molecules to generate sheared nucleic acid molecules; (v) sequencing the sheared nucleic acid molecules to produce a first dataset comprising a first plurality of sequencing reads; (b) providing a second sample from the subject comprising cell-free DNA (cfDNA) molecules and processing the sample by (i) partitioning the cfDNA molecules on the basis of at least one feature thereby generating a plurality of subsamples comprising at least a first and a second subsample; (ii) attaching a set of adapters to the cfDNA molecules in each of the subsamples to generate adapter- ligated cfDNA molecules in the first and the second subsamples, wherein the molecular barcodes attached to the cfDNA molecules in the first subsample differ from the molecular barcodes attached to the cfDNA molecules in the second subsample, wherein the adapters are attached to both ends of the cfDNA molecules in the first and second sub samples; (iii) sequencing at least one subsample to generate a second dataset comprising a second plurality of sequencing reads; (c) processing the first dataset comprising a first plurality of sequencing reads and the second dataset comprising a second plurality of sequencing reads to determine overlapping genomic regions between the first and second datasets; and (d) determining from the overlapping genomic regions disease associated signals in the subject.
[0025] In an additional aspect, the disclosure provides a method for enriching disease associated signals comprising: (a) providing a first sample from a subject comprising cellular genomic DNA and processing the sample by: (i) cross-linking the cellular genomic DNA to produced crosslinked nucleic acid molecules; (ii) digesting the cross-linked nucleic acid molecules to produce digested, cross-linked nucleic acid molecules comprising at least one 3' overhang sequence and at least one 5' overhang sequence; (iii) repairing the at least one 3’ and 5’ overhang sequences; (iv) incorporating a capture label on the least 3’ and 5’ overhang sequences thereby creating labeled overhang sequences; (v) ligating the labeled overhang sequences to generate a junction marker between the at least one 3’ end overhang sequence and at least one 5’ overhang sequence in the nucleic acid molecules; (vi) fragmenting the nucleic acid molecules in (v) generate, fragmented nucleic acid molecules comprising a junction marker and fragmented nucleic acid molecules lacking a junction maker; (vii) separating the fragmented nucleic acid molecules by contacting the molecules in (iv) with a capture molecule with affinity for the capture label to generate fragmented nucleic acid molecules comprising a junction marker, wherein the junction marker comprises a capture label bound to the capture molecule and fragmented nucleic acid molecules lacking ajunction marker; and (viii) sequencing the fragmented nucleic acid molecules comprising a junction marker to generate a first set of sequencing reads comprising chromatin interaction partners; (b) providing a second sample comprising cell free DNA(cfDNA) molecules and processing the sample by: (i) partitioning the cfDNA molecules on the basis of at least one genetic, epigenetic and / or fragmentomic biomarkers thereby generating a plurality of subsamples comprising at least a first and a second subsample; (ii) ligating a first set of adapters to the plurality of the subsamples, wherein the molecular barcodes for the first subsample differ from the molecular barcodes of the second subsample to generate adapter-ligated DNA molecules in the first and the second subsamples, wherein the adapters are ligated at both ends of the doublestranded DNA molecules; (iii) pooling the adapter ligated molecules from the first subsample and the adapter ligated molecules from the second subsample to generate a mixture comprising adapter-ligated molecules from the first subsample and adapter ligated molecules from the second subsample; (iv) amplifying the mixture in (iii) to generate amplified molecules; (v) enriching the amplified molecules in (iv) to generate a plurality of enriched molecules; (vi) amplifying the enriched molecules in (v) to generate amplified enrich molecules; (vii) sequencing the amplified enrich molecules to generate a second set of sequencing reads; (c) processing the second set of sequencing reads to generate a second dataset comprising a plurality of genetic, epigenetic, and / or fragmentomic markers; (d) determining for the chromatin interaction partners in the first dataset one or more genetic, epigenetic, and / or fragmentomic states using the second dataset thereby generating annotated chromatin interaction partners, thereby integrating information from the first dataset and the second datasets; and (e) determining from the annotated chromatin interaction partners in (d) a plurality of cancer associated signals thereby enriching disease associated signals.
[0026] In an additional aspect, the disclosure provides a computer-implemented method to enrich disease associated signals in a plurality of samples from a subject, comprising: (a) providing, in computer memory, a first dataset from a first sample comprising a plurality of chromatin interaction partners and a second dataset from a second sample comprising epigenetic data; (b) determining from the second dataset an epigenetic state for the plurality of chromatin interaction partners in the first dataset; and (c) determining a plurality of disease associated signals from the epigenetic state of the plurality of interaction partners, thereby enriching disease associated signals in the plurality of samples from the subject.
[0027] In some embodiments, the method further comprises determining that the chromatin interaction partners exhibit similar epigenetic state, and using the similarity information to error correct epigenetic states between interaction partners.
[0028] In some embodiments, the first and second samples are taken from the subject at different time points. In some embodiments, the first and second samples are taken at a same time point, wherein the first and second samples are obtained from a whole blood sample.
[0029] In some embodiments, the first sample comprising cellular genomic DNA is collected from buffy coat and the second sample comprising cell free DNA is collected from plasma.
[0030] In some embodiments, the method further comprises annotating chromatin interaction partners with functional genomic data to enrich functionally relevant genomic loci.
[0031] In some embodiments, the first dataset comprises promoter capture Hi-C.
[0032] In some embodiments, the first sample comprises a tissue sample and the second sample comprises cell free DNA (cfDNA).
[0033] In additional aspects, the disclosure provides a computer system for diagnosing a disease in a subject, the system comprising: a sequencing system configured to receive and process a sample comprising nucleic acid molecules collected from the subject, the sequencing system comprising: a sequencing pipeline having one or more sequencing devices for associating the nucleic acid molecules in the sample with sequence reads; a processor programmed to: receive, via a sequence analysis pipeline, the sequence reads from the sequencing system; determine a plurality of features encoded in the sequencing reads, the features comprising chromatin conformation data from a tumor in the subject and genetic sequence information from cell free DNA molecules obtained from the subject after curative-intent treatment of the tumor; input the plurality of features into a previously trained classifier wherein the classifier is configured to determine whether the disease is present in the subject; and determine, based on an output of the classifier, whether or not the disease is present in the subject, thereby diagnosing the disease in the subject.
[0034] In some embodiments, the disease diagnosed by the system comprises minimal residual disease.
[0035] In other aspects the disclosure provides a system for diagnosing a disease in a subject, the system comprising: a sequencing system configured to receive and process a sample comprising nucleic acid molecules collected from the subject, the sequencing system comprising: a sequencing pipeline having one or more sequencing devices for associating the nucleic acid molecules in thesample with sequence reads; a processor programmed to: receive, via a sequence analysis pipeline, the sequence reads from the sequencing system; determine a plurality of features encoded in the sequencing reads, the features comprising chromatin state data from a tumor in the subject and genetic and epigenetic sequence information from cell free DNA molecules obtained from the subject after curative-intent treatment of the tumor; input the plurality of features into a previously trained classifier wherein the classifier is configured to determine whether the disease is present in the subject; and determine, based on an output of the classifier, whether or not the disease is present in the subject, thereby diagnosing the disease in the subject.
[0036] In some embodiments, the chromatic state used within the system comprises, chromatin interaction data, chromatin accessibility data, and / or transcription factor occupancy. In additional embodiments, the system comprises, determining minimal residual disease status in the subject.BRIEF DESCRIPTON OF THE DRAWINGS
[0037] Fig. 1A. The figure shows an example flowchart illustrating a method for identifying cancer associated epigenetic landscapes in tissue and targeted resequencing of selected targets in a cell-free DNA sample.
[0038] Fig. IB. The figure shows an example schematic diagram illustrating methods to profile cancer associated epigenetic landscapes in a tissue sample.
[0039] Fig. 1C. The figure shows an example schematic diagram illustrating methods to profile cancer associated epigenetic landscapes in a cell-free DNA sample.DETAILED DESCRIPTION
[0040] The concept of an “epigenetic landscape” was originally introduced by Conrad Waddington in the 1940s to visualize the developmental pathways a cell might follow to differentiate into different cell types. The model describes the complex process of cellular development and differentiation, where a cell’s fate is determined by a combination of genetic and environmental factors. Importantly, the metaphorical landscape illustrates how cells can differentiate into various types based on epigenetic biochemical marks such as DNA methylation and histone modifications. These marks are modifiable by environmental factors like diet, stress, and toxins, thereby shaping the “landscape” through which cells “travel” during development. The idea underscores the importance of gene regulation mechanisms beyond the genetic sequence.
[0041] The epigenetic landscape in cancer cells is characterized by significant epigenomic alterations compared to normal cells. Some of the most well studied epigenomic alterations include DNA methylation, histone modifications, chromatin remodeling, nucleosome positioning, and non-coding RNAs expression. These epigenetic mechanisms influence chromatin architecture and ultimately gene regulation. These epigenetic changes contribute to the initiation and progression of cancer cells by altering gene expression patterns and leading to malignant cellular transformation.
[0042] The methods described herein are directed to the integration of a plurality of sequencing datasets comprising a plurality of epigenetic states to identify disease specific epigenetic landscapes in cell-free DNA. Data integration is achieved through a variety of artificial intelligence (Al), machine learning, and deep learning methods. Combined, these methods can elucidate molecular changes in the multilayer epigenome structure in health and disease.
[0043] Epigenetic landscapes can comprise sets of epigenetic data. In some embodiments, of the disclosure the epigenetic data can comprise methylation data, histone modification data, chromatin conformation capture data, nucleosome positioning, histone variants for example, replacement of the canonical histone H2A with the variant H2A.Z, RNA methylation, chromatin accessibility, DNA hydromethylation, DNA phosphorylation, acetylation, transcription factor binding sites, and / or chromatin looping. In some embodiments, of the disclosure the TFBS are directly assayed using methods such as ChlP-seq or CUT&Tag, in alternative embodiments the TFBS are inferred based on epigenetic data sets, machine learning algorithms.
[0044] Additionally, the disclosure provides methods to screen, diagnose, stage, and / or subtype a disease by monitoring changes in disease specific epigenetic landscapes comprising sets of epigenetic biochemical modifications such as DNA methylation and histone modifications, DNA accessibility changes and chromatin structure e.g., DNA looping.
[0045] In certain embodiments of the disclosure the disease comprises a type of cancer. In yet other embodiments the disease comprises a form of Alzheimer’s disease (AD). In additional embodiments, the disease comprises autoimmune disorders including Rheumatoid Arthritis (RA), Crohn’s disease and / or ulcerative colitis.
[0046] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with such embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention isintended to cover all alternatives, modifications, and equivalents, which may be included within the invention as defined by the appended claims.
[0047] Before describing the present teachings in detail, it is to be understood that the disclosure is not limited to specific compositions or process steps, as such may vary. It should be noted that, as used in this specification and the appended claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, reference to “a nucleic acid” includes a plurality of nucleic acids, reference to “a cell” includes a plurality of cells, and the like.
[0048] Numeric ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant digits and the error associated with the measurement. Also, the use of “comprise”, “comprises”, “comprising”, “contain”, “contains”, “containing”, “include”, “includes”, and “including” are not intended to be limiting. It is to be understood that both the foregoing general description and detailed description are exemplary and explanatory only and are not restrictive of the teachings.
[0049] Unless specifically noted in the above specification, embodiments in the specification that recite “comprising” various components are also contemplated as “consisting of’ or “consisting essentially of’ the recited components; embodiments in the specification that recite “consisting of’ various components are also contemplated as “comprising” or “consisting essentially of’ the recited components; and embodiments in the specification that recite “consisting essentially of’ various components are also contemplated as “consisting of’ or “comprising” the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0050] The section headings used herein are for organizational purposes and are not to be construed as limiting the disclosed subject matter in any way. In the event that any document or other material incorporated by reference contradicts any explicit content of this specification, including definitions, this specification controls.Definitions
[0051] “Cell-free DNA,” “cfDNA molecules,” or simply “cfDNA” include DNA molecules that naturally occur in a subject in extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). While the cfDNA originally existed in a cell or cells in a large complex biological organism, e.g., a mammal, it has undergone release from the cell(s) into a fluid found in the organism, and may be obtained by obtaining a sample of the fluid without the need to perform an in vitro cell lysis step.
[0052] As used herein, a modification or other feature is present in “a greater proportion” in a first subsample or population of nucleic acid than in a second subsample or population when the fraction of nucleotides with the modification or other feature is higher in the first subsample or population than in the second population. For example, if in a first subsample, one tenth of the nucleotides are mC, and in a second subsample, one twentieth of the nucleotides are mC, then the first subsample comprises the cytosine modification of 5-methylation in a greater proportion than the second subsample.
[0053] As used herein, “machine learning model” (or “model”) refers to a collection of parameters and functions, where the parameters are trained on a set of training samples or individual data points or instances used to train a machine learning model. These samples are part of the dataset that provides the model with examples of input data along with the corresponding output (for supervised learning), or just input data (for unsupervised learning). The parameters and functions may be a collection of linear algebra operations, non-linear algebra operations, and tensor algebra operations. The parameters and functions may include statistical functions, tests, and probability models. The training samples can correspond to samples having measured properties of the sample (e.g., genomic, epigenomic, transcriptomic, metabolites etc. data and other subject data, such as histology, imaging data and / or electronic medical health records, or insurance claim data), as well as known patient / sample metadata including classifications or labels for example molecular phenotypes or specific cancer or disease therapies. Other phenotypes can include patient biomedical information including “cardiovascular phenotypes” or “cardiovascular risk factors” such as weight, height, Body Mass Index(BMI) , and other physical characteristics. Yet other phenotypes can include cancer risk factors including smoking, excessive alcohol consumption, poor diet, physical inactivity, obesity, genetic predispositions, exposure to harmful chemicals and radiation, chronic inflammation, certain infections (such as human papillomavirus, hepatitis B and C), hormonal imbalances, and advanced age. The model can learn from the training samples in a training process that optimizes the parameters (and potentially the functions) to provide an optimal quality metric (e.g., accuracy) for classifying new samples. A variety of advanced statistical and computational methods that can be employed as training functions including Expectation Maximization(EM) to find maximum likelihood estimates of parameters in probabilistic models, especially for models with latent variables, Maximum Likelihood Estimation (MLE) to estimate the parameters of a statistical model. MLE methods select the set of parameters that maximize the likelihood function i.e., the parameters under which the observed data is most probable. BayesianParameter Estimation Methods which incorporate prior knowledge in addition to the data at hand through the use of probability distributions. These include Markov Chain Monte Carlo (MCMC), Gibbs Sampling, Hamiltonian Monte Carlo (HMC), and Variational Inference (VI), or Gradient- Based Methods including Stochastic Gradient Descent (SGD) and the Broyden-Fletcher-Goldfarb- Shanno (BFGS) algorithm.
[0054] As used herein, “Deep learning models” refers to a collection of “architectures” or “algorithms” useful in scenarios where machine learning approaches may fall short, for example due to the complexity of the data (e.g., high dimensional genomics, epigenomics datasets used alone or in one or more combinations). Additional high dimensional data can include medical images such as MRIs, or histology reports. Deep learning models, especially CNNs, can automatically extract relevant features without manual feature engineering.
[0055] Exemplary applications for deep learning models include complex pattern recognition tasks for example, recognizing specific functional elements in genomics data, functional elements can include transcription factor binding sites (TFBS) or chromatin interaction sites for example promoter-enhancer interactions. Additionally, recognition of disease specific features in histology images including patters or specific features of PDL-1 expression.
[0056] As used herein, “without substantially altering base-pairing specificity” of a given nucleobase means that a majority of molecules comprising that nucleobase that can be sequenced do not have alterations of the base pairing specificity of the second nucleobase relative to its base pairing specificity as it was in the originally isolated sample. In some embodiments, 75%, 90%, 95%, or 99% of molecules comprising that nucleobase that can be sequenced do not have alterations of the base pairing specificity of the second nucleobase relative to its base pairing specificity as it was in the originally isolated sample.
[0057] As used herein, “base pairing specificity” refers to the standard DNA base (A, C, G, or T) for which a given base most preferentially pairs. Thus, for example, unmodified cytosine and 5- methylcytosine have the same base pairing specificity (i.e., specificity for G) whereas uracil and cytosine have different base pairing specificity because uracil has base pairing specificity for A while cytosine has base pairing specificity for G. The ability of uracil to form a wobble pair with G is irrelevant because uracil nonetheless most preferentially pairs with A among the four standard DNA bases.
[0058] As used herein, a “combination” comprising a plurality of members refers to either of a single composition comprising the members or a set of compositions in proximity, e.g., in separatecontainers or compartments within a larger container, such as a multiwell plate, tube rack, refrigerator, freezer, incubator, water bath, ice bucket, machine, or other form of storage.
[0059] The “capture yield” of a collection of probes for a given target set refers to the amount (e.g., amount relative to another target set or an absolute amount) of nucleic acid corresponding to the target set that the collection of probes captures under typical conditions. Exemplary typical capture conditions are an incubation of the sample nucleic acid and probes at 65°C for 10-18 hours in a small reaction volume (about 20 |1L) containing stringent hybridization buffer. The capture yield may be expressed in absolute terms or, for a plurality of collections of probes, relative terms. When capture yields for a plurality of sets of target regions are compared, they are normalized for the footprint size of the target region set (e.g., on a per-kilobase basis). Thus, for example, if the footprint sizes of first and second target regions are 50 kb and 500 kb, respectively (giving a normalization factor of 0.1), then the DNA corresponding to the first target region set is captured with a higher yield than DNA corresponding to the second target region set when the mass per volume concentration of the captured DNA corresponding to the first target region set is more than 0.1 times the mass per volume concentration of the captured DNA corresponding to the second target region set. As a further example, using the same footprint sizes, if the captured DNA corresponding to the first target region set has a mass per volume concentration of 0.2 times the mass per volume concentration of the captured DNA corresponding to the second target region set, then the DNA corresponding to the first target region set was captured with a two-fold greater capture yield than the DNA corresponding to the second target region set.
[0060] “Capturing” one or more target nucleic acids refers to preferentially isolating or separating the one or more target nucleic acids from non-target nucleic acids.
[0061] A “captured set” of nucleic acids refers to nucleic acids that have undergone capture.
[0062] A “target-region set” or “set of target regions” refers to a plurality of genomic loci targeted for capture and / or targeted by a set of probes (e.g., through sequence complementarity).
[0063] “Corresponding to a target region set” means that a nucleic acid, such as cfDNA, originated from a locus in the target region set or specifically binds one or more probes for the target-region set.
[0064] “Specifically binds” in the context of an probe or other oligonucleotide and a target sequence means that under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence, or replicates thereof, to form a stable probe:target hybrid, while at the same time formation of stable probemon-target hybrids is minimized. Thus, a probehybridizes to a target sequence or replicate thereof to a sufficiently greater extent than to a nontarget sequence, to enable capture or detection of the target sequence. Appropriate hybridization conditions are well-known in the art, may be predicted based on sequence composition, or can be determined by using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) at §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, particularly §§ 9.50-9.51, 11.12- 11.13, 11.45-11.47 and 11.55-11.57, incorporated by reference herein).
[0065] “Sequence-variable target region set” refers to a set of target regions that may exhibit changes in sequence such as nucleotide substitutions (i.e., single nucleotide variations), insertions, deletions, or gene fusions or transpositions in neoplastic cells (e.g., tumor cells and cancer cells).
[0066] “Epigenetic target region set” refers to a set of target regions that may show sequenceindependent changes in neoplastic cells (e.g., tumor cells and cancer cells) or that may show sequence-independent changes in cfDNA from subjects having cancer relative to cfDNA from healthy subjects. Examples of sequence-independent changes include, but not limited to, changes in methylation (increases or decreases), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. For present purposes, loci susceptible to neoplasia-, tumor-, or cancer-associated focal amplifications and / or gene fusions may also be included in an epigenetic target region set because detection of a change in copy number by sequencing or a fused sequence that maps to more than one locus in a reference genome tends to be more similar to detection of exemplary epigenetic changes discussed above than detection of nucleotide substitutions, insertions, or deletions, e.g., in that the focal amplifications and / or gene fusions can be detected at a relatively shallow depth of sequencing because their detection does not depend on the accuracy of base calls at one or a few individual positions. In some embodiments, the epigenetic target region set includes one or more genomic regions, where the epigenetic state (e.g., methylation state) of cfDNA molecules in these regions is unchanged in cancer, but their presence / quantity in blood indicates increased, aberrant presentation of cfDNA from certain tissue (e.g. cancer origin) into circulation.
[0067] A nucleic acid is “produced by a tumor” or ctDNA or circulating tumor DNA, if it originated from a tumor cell. Tumor cells are neoplastic cells that originated from a tumor, regardless of whether they remain in the tumor or become separated from the tumor (as in the cases, e.g., of metastatic cancer cells and circulating tumor cells).
[0068] The term “methylation” or “DNA methylation” refers to addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to addition of a methyl group to a cytosine at a CpG site (cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in a 5’ -> 3’ direction of the nucleic acid sequence). In some embodiments, DNA methylation refers to addition of a methyl group to adenine, such as in N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the 5th carbon of cytosine). In some embodiments, 5-methylation refers to addition of a methyl group to the 5C position of the cytosine to create 5 -methylcytosine (5mC). In some embodiments, methylation comprises a derivative of 5mC. Derivatives of 5mC include, but are not limited to, 5- hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the 3rd carbon of cytosine). In some embodiments, 3C methylation comprises addition of a methyl group to the 3C position of the cytosine to generate 3 -methylcytosine (3mC). Methylation can also occur at non CpG sites, for example, methylation can occur at a CpA, CpT, or CpC site. DNA methylation can change the activity of methylated DNA region. For example, when DNA in a promoter region is methylated, transcription of the gene may be repressed. DNA methylation is critical for normal development and abnormality in methylation may disrupt epigenetic regulation. The disruption, e.g., repression, in epigenetic regulation may cause diseases, such as cancer. Promoter methylation in DNA may be indicative of cancer.
[0069] The term “hypermethylation” refers to an increased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypermethylated DNA can include DNA molecules comprising at least 1 methylated residue, at least 2 methylated residues, at least 3 methylated residues, at least 5 methylated residues, or at least 10 methylated residues.
[0070] The term “hypomethylation” refers to a decreased level or degree of methylation of nucleic acid molecule(s) relative to the other nucleic acid molecules within a population (e.g., sample) of nucleic acid molecules. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA can include DNA molecules comprising 0 methylated residues, at most 1 methylated residue, at most 2 methylated residues, at most 3 methylated residues, at most 4 methylated residues, or at most 5 methylated residues.
[0071] The terms “or a combination thereof’ and “or combinations thereof’ as used herein refers to any and all permutations and combinations of the listed terms preceding the term. For example,“A, B, C, or combinations thereof’ is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CAB ABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.
[0072] “Or” is used in the inclusive sense, i.e., equivalent to “and / or,” unless the context requires otherwise.Exemplary methods
[0073] FIG. 1A-C show an exemplary method for determining cancer associated epigenetic landscapes in tissue and subsequent targeted resequencing of selected targets in cell-free DNA sample. The workflow begins by analyzing a patient’s tissue sample with a plurality of assays to identify disease specific epigenetic landscapes. Such epigenetic landscapes comprise datasets comprising layers of molecular phonotypes such as chromatin state, TFBS etc. Using a variety of methods such as a neural network or a deep learning network to recognize underlying relationships in the datasets these datasets of molecular phenotypes can be integrated to generate an output including but not limited to a classification i.e., whether a specific epigenetic landscape is associated with the disease , a pattern recognition e.g., a specific methylation patter or TFBS associated with the disease, and / or data clustering. Alone or in combination, these outputs can help determine disease associated epigenetic landscapes which are then targeted sequenced in cell free DNA. This initial stage of sample processing is illustrated in FIG. IB. In a subsequent stage, as illustrated in FIG. 1C, the epigenetic landscapes are profiled in cell free DNA using targeted resequencing for each of the individual types of molecules for example, histone marks, TFBS etc. Such data is then used to computationally infer chromatin interactions in cell free DNA. In some embodiments, the epigenetic landscapes may be identified based on one or more features e.g., TFBS, methylation status. In additional embodiments, although the epigenetic landscapes identified in tissue may comprise several layers of biochemical information, such information may not be recovered or be partially recovered in cell free DNA. Nevertheless, the epigenetic landscapes may still be identifiable in cell free DNA based on one or more layers or biochemical information including histone marks and / or TFBS. In yet other embodiments, it may be useful to remove datasets or features within datasets for example to reduce dimensionality of the data. Similarly, in some embodiments it may be useful to remove correlated features in the datasets.Moreover, in some embodiments, layers of biochemical information may be used to infer features that are not directly detected in the assays. For example, transcriptional activity can be inferred based on histone marks including acetylation of H3K9, H3K14, H3K27, and H4K16, or trimethylation of histone H3 at lysine 4 (H3K4me3) located after first exons and around transcription start sites. In other embodiments the epigenetic landscapes are profiled using genome-wide assays in tissue; however; in additional embodiments the epigenetic landscapes may be profiled using exome sequencing or targeted sequencing in tissue. Any number of epigenetic landscapes can be targeted resequenced in cell free DNA. For example, one or more than one up to all landscapes originally identified in tissue, can be targeted sequenced in cell free DNA. In a preferred embodiment between 10-20 epigenetic landscapes originally identified in tissue can be targeted resequenced in cell free DNA.Biochemical assaysHi-C technique to capture chromatin conformation
[0074] Three-dimensional chromatin organization varies among different cell types and plays a crucial role in gene regulation by bringing distant functional elements into close spatial proximity. These functional elements contribute to maintaining homeostasis in health and play a pivotal role in disease regulation. The implementations of techniques, architectures, frameworks, systems, processes, and computer-readable instructions described herein are directed to the analysis of three-dimensional chromatin organization as captured by chromatin conformation capture (3C) techniques including, 4C (Circular Chromosome Conformation Capture), 5C (Chromosome Conformation Capture Carbon Copy), Hi-C, ChlA-PET (Chromatin Interaction Analysis by Paired-End Tag Sequencing), Capture-C, Capture Hi-C, HiChIP, Micro-C.
[0075] Following is an exemplary workflow that outlines the steps in preparing a Hi-C library, from cell culture to final library preparation and quality control.
[0076] In exemplary workflows for the Hi-C technique cells are cultured and chromatin is crosslinked e g., fix cells in 1% formaldehyde, quench with glycine, harvest by centrifugation, and store at -80°C for future use. Crosslinking conditions are typically standardized to ensure consistency across experiments. Cells are then lysed, and chromatin is digested. For example, cells can be lysed with a Dounce homogenizer in the presence of cold hypotonic buffer supplemented with protease inhibitors and IGEPAL CA-630. Lysates are wash and resuspended in restriction enzyme buffer, chromatin is then solubilized with SDS and incubated at 65°C, quench SDS with Triton X-100, and digested with a restriction enzyme (e.g., Hindlll). Biotin marking of DNA endsand blunt end ligation: DNA ends are labeled with biotin- 14-dCTP using the Klenow fragment of DNA polymerase I. Then, DNA fragments are ligated in a diluted condition at 16°C to favor intramolecular ligation. DNA is then purified with the following steps, degradation of proteins with Proteinase K, extraction with phenol: chloroform, DNA precipitation, resuspension, and RNase A treatment to yield high-quality DNA. Following DNA purification several quality control steps can be performed to ensure the library meets quality metrics. For example, by profiling fragment size distribution and quantify DNA using Agilent Bioanalyzer or Agilent TapeStation systems. Biotin removal from un-ligated ends: T4 DNA polymerase is used to remove biotin-labeled ends that have not been ligated. DNA is then precipitated and washed to prepare for downstream applications. DNA fragmentation and size fractionation: DNA is fractionated based on size using the Covaris 8700 and AMPure XP beads to achieve the desired size distribution for sequencing. End repair and “A” tailing: DNA molecules are repaired for asymmetric breaks and prepare for Illumina adapter ligation by filling in overhangs, and adenylating the 3’ ends. Streptavidin pulldown of biotinylated Hi-C ligation products: Mix biotinylated Hi-C ligation products with streptavidin beads, wash, and prepare for Illumina adapter ligation. Paired-end adapter ligation and library amplification: Illumina paired-end adapters are then ligated while the DNA is bound to streptavidin beads, libraries are PCR amplified with minimal cycles to avoid PCR artifacts the resulting amplified DNA is purified. Final quality control and library quantification: Assess the quality of the final amplified library and quantify it, ideally using a Bioanalyzer, before sequencing.
[0077] Moreover, the chromatin conformation capture experiment (Hi-C, 3C, 4C, capture Hi-C etc.) might utilize chromatin fragmentation methods that do not depend on sequence specificity. Exemplary protocols include TopoLink™ (Catalog #: 21010) from Dovetail Genomics (part of Cantata Bio) . In certain embodiments, a targeted approach might be more desirable, for example, when the research question is focused on promoter-enhancer interactions. Exemplary protocols include a capture Hi-C protocol using Dovetail Genomics Dovetail® Targeted Enrichment Panels (Catalog # 25013). In some embodiments, experimental workflows can combine one or more experiments in a single workflow using a variety of methods. For example, the ChlP-seq and the Hi-C workflows can be combined in one workflow using Dovetail® HiChIP MNase Kit (Catalog #: 21007) see, for example, Yang, Jae-Hyun et al. “Loss of epigenetic information as a cause of mammalian aging.” Cell vol. 186,2 (2023): 305-326. e27. doi : 10.1016 / j . cell.2022.12.027, which is incorporated herein by reference. This approach can investigate chromatin interactions mediatedby specific proteins of interest. In additional embodiments, the methods disclosed herein can be used alone or in combination with Dovetail® Micro-C Kit (Catalog # 21006) or similar methods, to generate uniform fragments that capture nucleosome positioning information, and maintain even coverage across the genome. This approaches obtain ultra-high-resolution topology mapping down to the mono-nucleosome level (150-200 bp conformation), see, for example, Bayanjargal, Ariunaa et al. “The DBD-a4 helix of EWS: :FLI is required for GGAA microsatellite binding that underlies genome regulation in Ewing sarcoma.” bioRxiv the preprint server for biology 2024.01.31.578127. 31 Jan. 2024, doi: 10.1101 / 2024.01.31.578127. Preprint which is incorporated herein by reference. Additionally, Hi-C experiments can improve and / or obtain haplotype phasing, genome assembly, and / or variant detection. Protocols and reagents to obtain such information include for example, Dovetail® Omni-C® Kit (Catalog # 21006). See, for example, Milevskiy, Michael J G et al. “Three-dimensional genome architecture coordinates key regulators of lineage specification in mammary epithelial cells.” Cell genomics vol. 3,11 100424. 16 Oct. 2023, doi: 10.1016 / j.xgen.2023.100424, which is incorporated herein by reference.Exemplary workflows for the capture Hi-C technique
[0078] Cross-linking of Chromatin: Cells are treated with formaldehyde or a combination of crosslinkers to preserve physical interactions between chromosomal regions. Digestion of DNA: The cross-linked chromatin is then digested using a restriction enzyme. This step is crucial for creating ends that can be ligated later. Some protocols might use a combination of enzymes for more efficient digestion. Ligation under Dilute Conditions: The digested chromatin is diluted and ligated, allowing for the ligation of interacting DNA ends that are in close proximity due to chromatin folding. Purification and Shearing: Cross-links are reversed, and the DNA is purified. The DNA may then be sheared into smaller fragments to prepare for library preparation. Capture Step: Biotinylated probes, designed to hybridize to regions of interest, are used to selectively capture specific fragments from the ligated DNA pool. This step enhances the resolution and specificity of interactions being analyzed. Library Preparation: Captured DNA fragments are processed into a sequencing library, including end repair, A-tailing, adapter ligation, and enrichment of targeted fragments through biotin-streptavidin pull-down. Sequencing: The prepared library is sequenced using high-throughput sequencing technology. Capture-Hi-C data analysis involves processing sequencing data is to identify chromatin interactions, including mapping reads to a reference genome, filtering, and identifying significant interactions within the captured regions.Detailed bioinformatic pipeline for Hi-C data
[0079] Various bioinformatic pipelines exists for the analysis of Hi-C datasets. Exemplary pipelines include, snHiC (Gregoricchio & Zwart, 2023). Briefly the analysis steps include (1) Generation of Contact Matrices: snHiC facilitates the creation of contact matrices at multiple resolutions in a single run, streamlining the initial analysis of Hi-C data. (2) Aggregation of Individual Samples: The pipeline allows for the aggregation of individual samples into user- specified groups, enabling comparative analyses across different conditions or time points. (3)Detection of Chromatin Features: It includes steps for the detection of domains, compartments, loops, and stripes, which are critical structural features of the genome organization revealed by Hi- C data. (4)Differential Analysis: snHiC supports differential compartment and chromatin interaction analyses, allowing users to identify changes in genome organization under different experimental conditions. The snHiC workflow can be automated using snakemake, which, makes it less prone to errors and more reproducible. To setup a yaml-formatted file is available to build a compatible conda environment, simplifying the setup and ensuring that users have all the necessary software and dependencies.
[0080] Similar analysis of Hi-C and Capture Hi-C data can be accomplished using publically available algorithms including HiCUP: A pipeline for mapping and processing Hi-C and CHi-C data, removing artefacts and producing quality control reports. CHICAGO: Capture Hi-C Analysis of Genomic Organization. HiC-bench, a comprehensive and reproducible Hi-C data analysis platform. HiC-Pro optimized pipeline for processing Hi-C data from raw reads to normalized contact maps. ChiCMaxima, a pipeline for detection and visualization of chromatin looping in CHi-C, which also allows integrating information from biological replicates.
[0081] DNA methylation plays a crucial role in regulating gene expression and maintaining genome stability. Disruption of DNA methylation control mechanisms can lead to various diseases, including cancer. A body of evidence supports that cancer cells exhibit largely different DNA methylation patterns compared to normal cells. In general, cancer cells are characterized by genome-wide hypomethylation, as well as hypermethylation of CpG islands associated with tumor suppressor genes and developmental regulators. Hypermethylation in the promoter regions of tumor suppressor genes leads to their silencing, while hypomethylation in the promoter regions of oncogenes can activate them, both mechanisms play a significant role in the development of cancer.DNA methylation sequencing workflows
[0082] Various wet lab methods exist to obtain bisulfite converted DNA the following is an exemplary wet lab workflow. DNA Extraction and Quantification: Genomic DNA is extracted from the sample of interest, followed by quantification and quality assessment. Bisulfite Conversion: DNA is treated with sodium bisulfite, which converts unmethylated cytosines to uracil while leaving methylated cytosines unchanged. Library Preparation: The bisulfite-treated DNA is then used to prepare a sequencing library. Library preparation steps include fragmentation of DNA, end-repair, A-tailing, adapter ligation, and size selection. The library preparation method should be compatible with bisulfite-converted DNA, as the conversion process can degrade the DNA. PCR Amplification: The library is amplified by PCR to enrich for fragments that have adaptors ligated at both ends. Sequencing: The prepared library is sequenced using high-throughput sequencing technology. The choice of sequencing platform (e.g., Illumina, PacBio, or Oxford Nanopore) can vary based on the desired read length, throughput, and cost considerations. Data Analysis: Sequencing data is processed to identify and quantify methylation patterns across the genome. This involves aligning reads to a reference genome, identifying methylated cytosines, and performing downstream analyses such as differential methylation analysis between samples or groups.DNA methylation data analysis tools
[0083] RnBeads is a software tool for large-scale analysis and interpretation of DNA methylation data Msuite: is an analysis toolkit for DNA methylation profiling, specifically optimized for emerging bisulfite-free methods. methylKit: An R package for the analysis of genome-wide DNA methylation profiles which also supports epigenome-wide association studies and biomarker discovery. BSmooth: Provides alignment, quality control, and analysis pipeline for whole-genome bisulfite sequencing. Similar open-source methods include MethLAB, MethCy and Methylation plotter. Additional methods include BEAT (BS-Seq Epimutation Analysis Toolkit), an R / Bioconductor package for quantitative analysis of DNA methylation from bisulfite sequencing data, utilizing a binomial mixture model. SINBAD, designed for pre-processing, quality assessment, and analysis of single-cell methylation data, starting from multiplexed sequencing reads. CpGtools, a Python package for analyzing DNA methylation data, offering a comprehensive suite for analyzing, annotating, QC, and visualizing the data and MethTools a toolbox for visualizing and analyzing DNA methylation data generated by the bisulfite sequencing.Single site methylation sequencing workflows
[0084] Various single-site methylation (SSM) sequencing methods exist, these are broadly categorized into RRBS-based and WGBS-based approaches. RRBS-based methods focus on GC- rich regions, optimizing cost and efficiency for single-cell research. WGBS-based methods offer genome-wide coverage, with variations to enhance DNA preservation and reduce loss. Both methodologies include adaptations for single-cell analysis, integrating technologies like microfluidics and unique molecular identifiers to improve accuracy and minimize DNA loss. Workflows for single-site methylation sequencing has seen significant recent advances, with methodologies focusing on high-throughput, cost-efficiency, and enhanced sensitivity. A typical outline of the workflow based on the most recent research includes the following steps: (1) Sample Preparation and DNA Isolation: Collection of target samples, followed by DNA extraction and purification. This initial step is crucial for ensuring the quality of DNA for subsequent methylation analysis. Bisulfite Treatment (Optional for Some Methods): Traditional workflows often involve bisulfite conversion, where unmethylated cytosines are converted to uracil, while methylated cytosines remain unchanged. However, newer methods like MLAD-seq and EAC-seq offer bi sulfite-free alternatives, providing single-base resolution and quantitative detection of 5mC without the DNA degradation associated with bisulfite treatment details of such protocols can be found in the literature for example Xiong et al., 2022 and Wang et al., 2022. Library Preparation and Sequencing: library preparation may involve enrichment of regions of interest or tagging DNA fragments with unique molecular identifiers (UMIs) before sequencing. Data Analysis and Interpretation: Post-sequencing, the data undergoes processing to identify methylated sites. Computational tools and algorithms, such as those incorporated in MethylScore, can accurately identify differentially methylated regions (DMRs) and predict phenotypes or disease states based on methylation profiles (Htither et al., 2022). Methods like EAC-seq be applied for direct and bi sulfite-free detection of DNA methylation.Enzymes used in single site methylation profiling
[0085] Bisulfite sequencing does not directly use enzymes for the conversion but relies on chemical treatment to differentiate between methylated and unmethylated cytosines. Methods that use enzymes include Ten-eleven translocation (TET)-assisted pyridine borane sequencing (TAPS): this method utilizes TET enzymes to oxidize 5mC to 5caC, which is then converted to thymine via pyridine borane reduction, allowing for sequencing without bisulfite conversion. Enzymatic methyl-seq (EM-seq) this method involves the use of enzymes for conversion, this approach is a less damaging alternative to bisulfite treatment for identifying methylated sites.Single cell methylation profiling
[0086] Additionally, single cell methylation profiling techniques can be applied to one or more of the methods in the disclosure, including for example, MLAD-seq a technique for single-base resolution and quantitative detection of 5mC in DNA, EAC-seq utilizes engineered proteins for bi sulfite-free, quantitative mapping of 5mC at single-base resolution, Digital-scRRBS a microfluidics-based platform for single-cell methylation sequencing, msRRBS a scalable singlecell reduced representation bisulfite sequencing technology that allows pooling of cell-specific barcoded DNA fragments before bisulfite conversion which improves efficiency and reduces cost.Methods to error correct DNA methylation data
[0087] Various methods exist to profile methylation status across the genome, including sodium bisulfite conversion and sequencing, differential enzymatic cleavage of DNA, and affinity- mediated capture of methylated DNA; however, these methods have limitations that may lead to misclassification, as the presence of unmethylated cytosines or DNA fragments may be erroneously recognized as methylated, and conversely, methylated cytosines may be detected as unmethylated. Additionally, the resolution of some techniques may not discern methylation variations in specific regions accurately. Therefore, there is a need for methods that can correct DNA methylation data, enhancing both specificity and sensitivity in analysis.
[0088] In additional aspects, the present disclosure provides methods to filter background noise in methylation data and to error correct DNA methylation data using chromatin interaction data to boost detection of disease associated signals. These methods work by leveraging the congruent methylation states of chromatin interaction partners Xi, Xii, Xiii...Xn, to identify and rectify inaccuracies in the methylation data. In essence, when interaction partners Xi and Xii are observed, they should consistently exhibit the same methylation status.
[0089] To address the noise in the methylation dataset, we construct a probabilistic model that captures the relationship between Xi and Xii and facilitates the removal of inaccuracies by emphasizing concordance and treating discordant instances as indicative of assay introduced errors not true biochemical states. An example of such model can comprise a linear model with an interaction or co-occurrence term. For example, given that Xi and Xii are interaction partners, methylation status:;:pO+pi Xi +p2*Xii +p3*(Xi xXii)+c. Where p3 represents the interaction or co-occurrence effect, indicating that the status of Xi depends on the status of Xii. In this model methylation status is the dependent variable and Xi status and Xii status are independent variables along with an interaction term. Individually, the terms in this linear model represent:
[0090] Methyl ation status: this is the dependent variable predicted by the model.
[0091] P0: this is the intercept term representing the predicted value of methylation_ status when both Xi and Xii are zero.
[0092] pi*Xi: this term represents the contribution of Xi to the predicted methylation_ status, pi is the coefficient associated the Xi determining the impact of a one-unit change on Xi on the predicted methylation_status.
[0093] P 2*Xii : similarly, this term represents the contribution of Xii to the predicted methylation_ status. P2 is the coefficient associated with Xii determining the impact of a one-unit change on Xii on the predicted methylation status.
[0094] P 3*(Xi *Xii): this is the interaction term, capturing the combined effect of Xi and Xii on the predicted methylation_ status when they interact.
[0095] E: represents the error or residual in the model, capturing the difference between the predicted methylation_status and the actual observed values.
[0096] In essence, this model describes how Xi is influenced by its own value, the value of Xii, and the interaction between them. The coefficients pi, P2, P3 determine the strength and direction of these influences.
[0097] Alternatively, the relationship between the status of Xi, Xii, Xiii. ..Xn and their impact on the target variable methylation status can be represented in a way that is not confined to a specific mathematical form. Exemplary models can include the function Methylation status = f(Xi, Xii)- c. Instead of using a predefined linear model, this is a more flexible approach using machine learning. By feeding a diverse set of data points into the model, we enable it to autonomously learn the relationship and dependencies between the features (Xi and Xii) and the targetvariable(methylation status). In this scenario, we want the model to learn the congruence between the statuses of Xi and Xii and use this learned relationship to identify and correct discordant methylation states between interaction partners in test data. In this context, training data can comprise a table listing all possible combinations of Xi and Xii such as Table 1. Additionally, a training data can comprise chromatin interaction data comprising multiple interaction partners and the corrected status. Furthermore, the training data can comprise data with labeled examples of Xi and Xii and their corrected status, specifically tailored to the relationship between the variables.
[0098] Methylation_status = f(Xi, Xii)- e, where:
[0099] Methyl ation status: represents the underlying relationship between Xi and Xii, it’s the dependent variable to be predicted or output of the model.
[0100] f(Xi, Xii) : represents the function that the machine learning model learns.
[0101] E: is the error term, accounting for any unexplained variability in the target variable that the model does not capture. The error term may also indicate that the observed value of the outcome variable may deviate from the predicted value due to random or unaccounted factors.
[0102] Such general function captures the relationship and interactions between Xi and Xii. Furthermore, this function can accommodate linear relationships, interactions, or more complex patterns based on the characteristics of the data. The target variable Methylation status is determined by a relationship involving independent variables Xi and Xii. In this model, the actual form of the function f is not explicitly defined; it is determined by the learning algorithm based on the patterns in the training data. This allows the model to adapt to the specific characteristics of the data without imposing rigid assumptions about the underlying relationships.
[0103] In some cases, a concordance filter may be sufficient to filter out discordant Xi, Xii, Xiii...Xn interaction partners. A concordance filter can be a rule-based approach that leverages the knowledge that Xi and Xii should always have the same status (methylated or unmethylated) to filter out data points where their statuses between the pairs are discordant. A data structure comprising columns Xi status, Xii_status, and filtered where:
[0104] Xi status: is a binary value (0 or 1) representing the status of Xi.
[0105] Xii_status: is a binary value (0 or 1) representing the status of Xii.
[0106] filtered: is a Boolean flag (True / False) indicating whether the data point is filtered or not (this flag can be initially set to False for all).
[0107] Additionally, a filtering rule can be applied to the chromatin interaction data whereby we iterate through each data point and apply the following rule: If Xi_status is equal to Xii_status, setfiltered to False (keep the data point). If Xi status is not equal to Xii status, set filtered to True (mark the data point as noise). This filtering rule approach assumes a perfect concordance between Xi and Xii, which might not always be true in chromatin interaction data and does not explicitly model the noise in the data collection process. However, assuming a perfect concordance between Xi and Xii, setting a filtering rule can improve the reliability of methylation assays based on the inherent relationship between one Xi, Xii or multiple interaction partners Xi, Xii, Xiii . . .Xn.As described earlier in this disclosure, to capture more complex relationships between Xi, Xii, Xiii....Xn, and other assay driven noise characteristics, machine learning models that learn the patterns in chromatin interaction data may be better suited. Other suitable options include Logistic Regression to predict the methylation status (methylated or unmethylated) of Xi based on both Xi status and Xii status. Alternately, a Decision Tree can be employed to learn a set of rules based on the methylation status or levels of methylation to classify data points as concordant or discordant. In additional embodiments, a Gaussian Mixture Models can be employed to model the joint distribution of Xi status and Xii status (methylation status or methylation level) to identify regions of high concordance and filter outliers.Methods to integrating sets of molecular data
[0108] Multiple methods can be employed to integrate multiple layers of epigenome information such as methylation patterns, transcription factor binding sites (TFBS), and histone modifications described earlier in the disclosure to predict regional features like the presence of a TFBS or chromatin interaction sites (e.g., enhancer promoter interaction) at a specific genomic location. For example, a neural network approach can integrate multiple layers of epigenomic information to predict regional features, using deep learning to handle high-dimensional data. Such model can undergo continuous iteration and validation against known biological insights can refine the model specificity and sensitivity.Neural Network Architecture; Multi-Modal Deep Learning Framework
[0109] Input Layer: the methods in the present disclosure provide datasets involving various types of epigenomic information, hence the input layer is designed to handle multiple data modalities. We achieve this by using separate input channels or sub-networks for each data type (methylation, TFBS, and histone modifications). Each channel can preprocess its respective data type, normalizing and encoding it in a form suitable for deep learning, that is including several steps to convert raw data into a format that neural networks can effectively process and learn from. These steps include data normalization / standardization, encoding, reshaping, handling missing values,feature engineering / selection, and data augmentation. Feature extraction layers: For each data modality, a convolutional neural network (CNNs) or recurrent neural network (RNN) can be used to capture spatial dependencies and patterns within the genomic sequences. CNNs are particularly useful for identifying patterns in histone modifications and TFBS, while RNNs or transformerbased models can effectively process sequential data, capturing long-range dependencies in methylation patterns for example. Integration Layer: After feature extraction, the outputs of the separate channels are integrated. We achieve this through concatenation, followed by dense layers, or by using more sophisticated integration techniques like attention mechanisms, which allow the model to weigh the importance of information from different epigenomic layers dynamically. Prediction Layer: We then feed the integrated features into one or more dense layers with nonlinear activation functions to enable the prediction of regional features, for example the presence of a TFBS at a specific genomic location( this holds true for other feature like histone marks and interaction sites). The output layer is designed according to the specific prediction task, for example binary classification for predicting the presence / absence of a TFBS. Training: The model should be trained on labeled datasets where the ground truth (e.g., presence or absence of TFBS in specific regions or other functional element or genomic feature) is known. We use a crossentropy loss function for classification tasks, and optimize the model using gradient descent algorithms like Adam or SGD. Regularization and Dropout: To prevent overfitting, given the complexity of epigenomic data and the deep architecture, we incorporate regularization techniques (L1 / L2 regularization) and dropout layers, particularly after dense layers in the network. Evaluation and Fine-tuning: We evaluate the model’ s performance using standard metrics like accuracy, precision, recall, and Fl score. Depending on the results, fine-tune the model by adjusting the architecture, hyperparameters, or training procedure. In some embodiments we use techniques like cross-validation for a more robust evaluation. Similarly, for enhanced per Performance methods such as data augmentation techniques specific to genomic data to increase the diversity of the training set. Other techniques include multi-task learning when for example predicting multiple regional features simultaneously, under this framework the network is designed to make several predictions at once, sharing representations between tasks to improve learning efficiency and prediction accuracy. Lastly techniques such as transfer learning can be employed to leverage pre-trained models on related tasks to improve performance, this method is especially useful when labeled data are limited.Neural Network; Training
[0110] In some embodiments, the neural network of the present invention is previously generated (i.e. previously trained). In other words, the neural network may be built and trained on known datasets prior to being used in the present invention. As previously explored, the model may be previously trained on labeled datasets. In some embodiments, ground truth of the labeled datasets may be known. For example, the labelled data sets may comprise information regarding a specific genomic region or the presence or absence of specific features within a specific genomic region.[0U1] In some embodiments, the labelled datasets may comprise information regarding the epigenetic landscape of a sample. In some embodiments, the labelled datasets may comprise information regarding the epigenetic landscape of a tumor. In some embodiments, the labelled datasets may comprise information regarding the epigenetic landscape of a specific genomic region. In some embodiments, information regarding the epigenetic landscape of a sample, tumor, or specific genomic region may comprise information regarding higher-order chromatin states.
[0112] In some embodiments, the labelled datasets may comprise information regarding a plurality of classes of molecules from a sample. The information may comprise the presence or absence of biochemical states of the plurality of classes of molecules from a sample. In some embodiments, biochemical states may comprise cytosine methylation, transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.
[0113] In some embodiments, the labelled datasets may comprise DNA sequence information of a sample. In some embodiments, the labelled datasets may comprise DNA sequence information of a tumor. In some embodiments, the labelled datasets may comprise DNA sequence information of a specific genomic region. In some embodiments, the labelled data sets may comprise tumor specific somatic genetic variation data.
[0114] In some embodiments, the labelled datasets used in training the neural network are obtained from a subject that is known to have a disease. In some embodiments, the disease is cancer. In some embodiments, the disease is cancer comprising at least one of adenocarcinoma, basal cell carcinoma, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, cholangiocarcinoma, colorectal cancer, endometrial cancer, esophageal cancer, gallbladder cancer, gastric cancer, germ cell tumors, glioma, head and neck cancer, hepatocellular carcinoma, Kaposisarcoma, kidney cancer, lip and oral cavity cancer, liver cancer, lung cancer, melanoma, mesothelioma, neuroendocrine tumors, ovarian cancer, pancreatic cancer, penile cancer, prostate cancer, sarcoma, skin cancer, small cell lung cancer, squamous cell carcinoma, stomach cancer, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vulvar cancer. In some embodiments, the subject is known to have previously had cancer. In some embodiments, the subject is known to have cancer recurrence.
[0115] In some embodiments, the subject is known to have been previously treated to remove a tumor. In some embodiments, wherein the subject is known to have been previously treated to remove a tumor, the neural network may be trained on labelled datasets obtained prior to resecting the tumor from the patient. In some embodiments, the labelled datasets may comprise chromatin conformation data, tumor specific somatic genetic variation data, DNA sequence information from cell free DNA (cfDNA). In some embodiments, the labelled datasets may comprise information regarding the tumor recurrence status of the patient at the time the sample was obtained
[0116] In some embodiments, the labelled datasets may be obtained from the same patient for which the generated neural network is used. In other words, the trained neural network may be patient specific. In some embodiments, the labelled datasets may not be exclusively obtained from the same patient for which the generated neural network is used. In other words, the trained neural network may not be patient specific. In some embodiments, the neural network may be trained on labelled datasets obtained from public databases. In some embodiments, the neural network may be trained on patterns or signatures that are known to be associated with a disease, such as cancer. Such patterns or signatures may be associated with genetic variation or epigenetic variation. Such patterns or signatures may be associated with the spatial organization of chromatin. Where the neural network is patient specific, the neural network may be tumor informed.
[0117] In some embodiments, the neural network may be trained on labelled datasets obtained from subjects that have undergone surgery to remove a disease. In some embodiments, the neural network may be trained on labelled datasets obtained from subjects that have not undergone surgery to remove a disease. In some embodiments, the neural network may be trained on labelled datasets obtained prior to the resection of a tumor. Therefore, the neural network may be tumor informed.
[0118] In some embodiments, the neural network may be trained on labelled datasets obtained from the patient prior to treatment or therapy. In some embodiments, neural network may be trained on labelled datasets obtained from patients who are undergoing treatment or therapy. Insome embodiments, the neural network may be trained on labelled datasets obtained from the patient who has previously undergone treatment of therapy. In some embodiments, the therapy may be effective in treating a disease having at least one of the epigenetic landscapes from the plurality of epigenetic landscapes determined in the sample from the subject. In some embodiments, the therapy may comprise an epigenetic therapy, targeting DNA methylation and histone modification mechanisms. In some embodiments, the epigenetic therapy may comprise Azacitidine, decitabine, vorinostat, and / or romidepsin. In some embodiments, the therapy may comprise a BET protein inhibitors, a histone methyltransferase inhibitors, a lysine-specific demethylase inhibitors, and / or Bromodomain inhibitors.
[0119] In some embodiments, the labelled datasets may comprise information regarding the prognosis of the disease, such as a specific cancer.
[0120] In additional embodiments, unsupervised learning approaches may be employed to identify patterns in unlabeled data using input feature variables without target or output variables. When applied to genomic datasets described in the present disclosure, unsupervised learning may be used to cluster samples based on specific attributes, detect anomalies, and / or dimensionality reduction. Unsupervised learning methods may include k-means, hierarchical clustering, PCA, t-SNE, similarity network fusion (SNF), and non-negative matrix factorization (NMF).
[0121] For example, SNF may be used to construct similarity networks for each omics dataset and iteratively fuse them into a single network that captures shared structures across data types. In some embodiments SNF may be used to cluster patient samples based on integrated multi-omics profiles, to uncover disease subtypes and / or biomarkers.
[0122] Moreover, NMF may be used to factorize high-dimensional genomic data into nonnegative matrices representing metagenes and metagene expression patterns to uncover latent molecular features and clustering genes or samples without supervision.
[0123] In certain embodiments a paired tumor-normal sample from the subject may be collected from the subject. The normal sample may serve as a subject specific germline reference against which to filter somatic variants. Such filtering may be useful for discriminating Clonal Hematopoiesis of Indeterminate Potential (CHIP) variants, determining true somatic variants, and / or selecting an MRD panel specific to the subject. For example, the normal sample can comprise white blood cells (WBC), peripheral blood mononuclear cells (PBMCs), plasma, serum, buccal swab, nasal swab, saliva, skin, fibroblasts, subcutaneous fat, and / or adipose tissue obtained from the subject.
[0124] Various examples of test data obtained from subjects are described throughout the specification. In additional embodiments, test samples may be collected prior, during, and / or after treatment from a subject diagnosed with cancer. Samples may be processed to obtain a plurality of molecules including DNA, RNA, and / or proteins. The samples may be further processed to capture molecular interactions for example, using formaldehyde to capture protein-DNA and / or DNA- DNA interactions. Following formaldehyde treatment, molecules are extracted and libraries prepared and sequenced. Molecular states and molecular interactions can then be determined bioinformatically. In some embodiments, when fed into the trained model, the test data is analyzed to determine a disease status, disease subtype, and / or disease reoccurrence. In some embodiments, when fed into the trained model, the test data is analyzed to determine treatment response and / or personalized treatment options for the subject. Moreover, when fed into the trained model, the test data may be analyzed to determine one or more epigenetic landscapes associated with disease in the subject.Neural Network; Integrating Inputs
[0125] The provided methods disclose a neural network that may be used for the integration of a plurality of datasets. Datasets may be generated through a variety of means, including assays performed on samples obtained from the subject. Such datasets may comprise sequencing datasets that comprise a plurality of epigenetic states. The integration of the datasets may allow the identification of disease specific epigenetic landscapes in a sample, such as a sample of cfDNA. Integration of datasets may elucidate molecular changes in a multilayer epigenome that can inform health and disease. Integration of datasets obtained by a variety of means may provide a more comprehensive insight into the disease state of the sample. For example, cancer is known to evolve wherein mutational signature changes over time. This often occurs in response to therapy or treatment, wherein the tumor adapts to the therapy or treatment it is exposed to. In such cases, tracking a singular or handful of mutations or epigenomic features in the sample may not be sufficient for generating a comprehensive overview of the disease state. Increasing the quantity information improves the ability to predict disease state. The provided neural network is able to integrate a plurality of datasets obtain using a variety of means to increase the predictability of disease state and monitoring.
[0126] In some embodiments, the integration is achieved using previously generated neural network. Methods for training such a neural network have been described in detail elsewhere. Insome embodiments, integration is achieved through a variety of artificial intelligence (Ai), machine learning, and deep learning methods.
[0127] Multiple layers of epigenome information may comprise information including, but not limited to, methylation patterns, transcription factor binding sites (TFBS), and histone modifications described earlier in the disclosure to predict regional features like the presence of a TFBS or chromatin interaction sites (e g., enhancer promoter interaction) at a specific genomic location. Moreover, datasets may comprise information such as mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation, methylation marks associated with poised enhancers, promoter regions and open chromatin. In some embodiments, dataset may comprise information such as a plurality of fragmentome biomarkers. In some embodiments, the fragmentome biomarkers comprise, fragment length score, fragment end-densities, the ratio of short (100-150 bp) to long (151-220 bp) fragments, and / or fragment length signatures.
[0128] In further embodiments, data inputted into neural network may comprise methylation data, histone modification data, chromatin conformation capture data, nucleosome positioning, histone variants for example, replacement of the canonical histone H2A with the variant H2A.Z, RNA methylation, chromatin accessibility, DNA hydromethylation, DNA phosphorylation, acetylation, transcription factor binding sites, and / or chromatin looping.
[0129] Datasets inputted into the neural network for integration can be obtained through a variety of means. Datasets may be obtained from a variety of assays that are performed on samples obtained from the subject, such as cfDNA samples. In some embodiments, datasets may be generated using a variety of sequence chemistries. These may include Illumina sequencing (MiSeq, HiSeq, NextSeq, NovaSeq, MiniSeq, iSeq 100), Oxford Nanopore sequencing (MinlON, GridlON, PromethlON, Flongle), Sanger sequencing (ABI 3730x1), Ion Torrent sequencing (Ion PGM, Ion S5, Ion GeneStudio S5), PacBio SMRT sequencing (Sequel, Sequel lie), Illumina NovaSeq 6000, PacBio HiFi sequencing (Sequel II and Sequel lie Systems using HiFi chemistry), Element Biosciences (AVITI System), Ultima Genomics (Ultima Sequencing), and Singular Genomics (G4 Sequencing System). In some embodiments, DNA methylation sequencing assays may be used. In some embodiments, single site methylation sequence assays may be used. In some embodiments, assays may be used to elucidate spatial organization. For example, methods such as Hi-C technique, as disclosed elsewhere herein, may be used to generate datasets for integration inthe neural network. In some embodiments, datasets can be generated using assays such as TAM- ChlP, CUT&RUN, CUT&Tag, ChlP-seq, and / or RNA-seq.
[0130] In some embodiments, the datasets that are inputted into the neural network may be obtained from a plurality of classes of molecules. In some embodiments, the plurality of molecules may comprise modified DNA, a DNA-protein complex, DNA-histone complex, and / or mRNA.
[0131] In some embodiments, datasets that are inputted into the neural network may be obtain at different time points. In some embodiments, datasets that are inputted into the neural network may be obtain at a plurality of time points. For example, datasets may be obtained from the subject at predetermined time points. In some embodiments, the subject may be undergoing treatment or therapy for a disease. In some embodiments, the datasets may be obtained at a plurality of time points during treatment or therapy for the disease. In some embodiments, the datasets may be obtained prior to treatment or therapy. In some embodiments, the datasets may be obtained subsequent to treatment or therapy. In some embodiments, the disease may be cancer.
[0132] In some embodiments, multiple datasets may be generated from the same sample. For example, a sample form the subject, such as a tissue sample, may be split and assays performed on the resulting partitions of the sample. Alternatively, assays may be performed in a sequential manner.
[0133] In some embodiments, chromatin structure and / or spatial chromatin organization may be extracted from a sample. From the same sample, genetic variants may be identified through sequencing assays. Identification of genetic variations may occur in parallel or sequential. Epigenetic variants may also be identified from the sample, using methods described herein. Identification of epigenetic variants in the sample can be performed in parallel or sequentially. The resulting datasets, comprising chromatin 3D structure information, genetic variant information, and / or epigenetic variant information, can be integrated using the trained neural network.
[0134] In some embodiments, DNA from the test subject may be used to obtain information that is inputted into the neural network. The DNA may be cfDNA. The DNA may be obtained from a tissue sample. Obtaining datasets to be inputted into the neural network may comprise capturing a plurality of sets of target regions from DNA from the subject. The plurality of sets of target regions may comprise sequence-variable target region set. The plurality of sets of target regions may comprise a sequence-variable target region set and an epigenetic target region set. The plurality of sets of target regions may comprise the sequence of at least one neoantigen. The sequence-variabletarget region set may comprise the sequence of at least one neoantigen. The capturing step may be performed according to any of the embodiments described elsewhere herein.Neural Network: Output
[0135] In some embodiments, a neural network generates an output including but not limited to a classification i.e., whether a specific epigenetic landscape is associated with the disease, a pattern recognition e.g., a specific methylation pattern or TFBS associated with the disease, and / or data clustering. In some embodiments, the neural network is a classifier.
[0136] In some embodiments, the neural network is designed to generate an output according to a specific prediction task. For example, the neural network may be a classifier and may output a binary decision based on the integration of the inputted data sets.
[0137] In some embodiments, the binary decision may be whether or not the subject has a disease. In some embodiments, the disease is cancer. In some embodiments, the disease is cancer comprising at least one of adenocarcinoma, basal cell carcinoma, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, cholangiocarcinoma, colorectal cancer, endometrial cancer, esophageal cancer, gallbladder cancer, gastric cancer, germ cell tumors, glioma, head and neck cancer, hepatocellular carcinoma, Kaposi sarcoma, kidney cancer, lip and oral cavity cancer, liver cancer, lung cancer, melanoma, mesothelioma, neuroendocrine tumors, ovarian cancer, pancreatic cancer, penile cancer, prostate cancer, sarcoma, skin cancer, small cell lung cancer, squamous cell carcinoma, stomach cancer, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, vulvar cancer.
[0138] In some embodiments, the binary decision may be whether the subject has tumor recurrence or cancer recurrence. In such examples, the neural network may be used to output a such a decision at multiple time points. In some embodiments these may be predetermined time points. For example, outputs may be generated during treatment or therapy, such as at predetermined time points during treatment or therapy. In further embodiments, outputs may be generated subsequent to treatment or therapy. In some embodiments, outputs may be generated subsequent to tumor resection.
[0139] In some embodiments, the output of the neural network may not be binary. For example, the output may determine a risk level based on the inputted data. In some embodiments, the risk level may be selected from a continuous scale. In some embodiments, the risk level may be selected from a discrete scale.RNA extraction and isolation
[0140] In some embodiments of the disclosed methods, the population of target nucleic acids comprises RNA and the method further comprises a cDNA synthesis step. RNA for use in the methods disclosed herein may be isolated from a blood sample or a sample comprising cells (such as a sample that includes immune and / or cancer-derived cells (e.g., a blood sample such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a peripheral blood PBMC sample)). General methods for RNA extraction and isolation (such as mRNA extraction and isolation) are known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al., Current Protocols of Molecular Biology, John Wiley and Sons (1997). Methods for RNA extraction from paraffin embedded tissues are disclosed, for example, in Rupp and Locker, Lab Invest. 56:A67 (1987), and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using a purification kit, buffer set, and protease(s) from commercial manufacturers, such as PreAnalytix GmbH or Qiagen, according to the manufacturer’s instructions. For example, RNA can be extracted from whole blood samples using the PAXgene® Blood RNA Kit (PreAnalytix GmbH). Other commercially available RNA isolation kits include MasterPure Complete DNA and RNA Purification Kit (EPICENTRE, Madison, WI), and Paraffin Block RNA Isolation Kit (Ambion, Inc.). Total RNA from tissue samples can be isolated using RNA Stat-60 (Tel -Test). RNA prepared from tumor tissue can be isolated, for example, by cesium chloride density gradient centrifugation. cDNA library preparation
[0141] Following RNA extraction from a sample (such as a blood sample), a cDNA library is typically prepared in preparation for sequencing, e.g., as in RNA-Seq. In some embodiments, the cDNAs in a library, such as an RNA-Seq library, can comprise a cDNA insert flanked by adapter sequences, such as adapter sequences used for amplification and sequencing on a particular platform. Exemplary cDNA library preparation methods are discussed below; however, cDNA library preparation methods can vary depending on the RNA species under investigation, which can differ in size, sequence, structural features and abundance. One of ordinary skill in the art will be able to select cDNA library preparation methods suitable for cDNA library preparation using an RNA species of interest. rRNA and / or globin mRNA depletion; polv(A) selection
[0142] Ribosomal RNAs (rRNAs) are the most abundant RNA species in most cells. Globin mRNA is also abundant in certain cell types found in the blood. Thus, some embodiments of the present disclosure comprise a step of ribosomal RNA (rRNA) depletion and / or a step of globinmRNA depletion. Such steps can be performed, e.g., following RNA extraction from a sample, and prior to a step of RNA fragmentation or cDNA fragmentation, prior to a step preparing cDNA from the RNA, prior to a step of ligating adapters to the cDNA, and prior to a sequencing step. In some embodiments, the methods include a step of rRNA depletion. In other embodiments, the methods include a step of globin mRNA depletion. In yet other embodiments, the methods disclosed herein include both a step of rRNA depletion and a step of globin mRNA depletion.
[0143] Any suitable rRNA depletion and / or globin mRNA depletion methods are of use in the present disclosure. One approach is to eliminate rRNAs uses sequence-specific probes that can hybridize to rRNAs (Hrdlickova et al., Wiley Interdiscip Rev RNA. 2017; 8(1): 10.1002 / wrna.1364). Unwanted rRNAs or their cDNAs are hybridized with biotinylated DNA or locked nucleic acid (LNA) probes, followed by depletion with streptavidin beads. Alternatively, rRNAs can be targeted by anti-sense DNA oligos and digested by RNase H, a method also known as probe-directed degradation (PDD). Another approach for rRNA reduction uses specific, not-so-random (NSR) primers that bind to the RNA molecules of interest during reverse transcription, thus avoiding reverse transcription of the rRNAs. For example, a method known as Ovation RNA-Seq (NuGen) uses hexamer or heptamer primers whose sequences are not present in rRNAs. In addition to sequence-based approaches, some methods take advantage of certain features of rRNAs for their elimination. The COT-hybridization method is based on heat denaturation, re-annealing, and selective degradation by a duplex-specific nuclease (DSN). Double-stranded cDNAs from abundant sequences are preferentially degraded because of their more rapid annealing kinetics compared to less abundant ones. Selective degradation has also been achieved using the enzyme terminator 5 ’-phosphate-dependent exonuclease (TEX), which recognizes RNA molecules with 5 ’-monophosphate, as with rRNAs and tRNAs. Further, commercial kits are available for rRNA and globin mRNA depletion, including, e.g., the Watchmaker Genomics RNA Library Prep Kit with Polaris Depletion.
[0144] Other embodiments of the present disclosure comprise a step of poly(A) selection. Such a step can be performed, e.g., following RNA extraction from a sample, and prior to a step of RNA fragmentation or cDNA fragmentation, prior to a step preparing cDNA from the RNA, prior to a step of ligating adapters to the cDNA, and prior to a sequencing step. In eukaryotic organisms, most protein coding RNAs (mRNAs) and many long noncoding RNAs (IncRNAs) (>200 nt) comprise a poly(A) tail (“polyadenylated RNAs”). The poly(A) tail may be used to enrich forpolyadenylated RNAs from total cellular RNA, in which polyadenylated RNAs may account for approximately 1-5% of total cellular RNA (Hrdlickova et al., Wiley Interdiscip Rev RNA. 2017;8(l): 10.1002 / wrna.1364). Exemplary poly(A) selection methods include, but are not limited to, use of magnetic or cellulose beads coated with oligo-dT molecules. Alternatively, polyadenylated RNAs can be selected using oligo-dT priming for reverse transcription (RT). Poly(A) selection may be combined with globin mRNA depletion.Fragmentation
[0145] In some embodiments, methods disclosed herein comprise fragmenting RNA isolated from a sample (such as RNA isolated from a sample comprising cells, such as a whole blood sample, a buffy coat sample, a leukapheresis sample, or a PBMC sample), such as following poly(A) selection or rRNA and / or globin mRNA depletion. RNA fragmentation methods can include physical fragmentation, chemical fragmentation, and / or enzymatic fragmentation. Physical fragmentation methods include, but are not limited to, acoustic or hydrodynamic shearing (such as sonication or point-sink shearing), needle shearing, and nebulization. Enzymatic fragmentation methods can include use of a ribonuclease (such as RNase III). RNA may also be fragmented using chemical shearing methods. Chemical fragmentation methods can include, but are not limited to, heat treatment of RNA in the presence of a divalent metal cation (such as magnesium or zinc). In some embodiments, the fragmenting provides RNA (such as mRNA) fragments of 25-400, 25- 300, 25-200, 50-400, 50-300, 50-250, 50-200, 100-400, 100-300, 100-200, 125-400, 125-300, 125- 200, 125-175, 150-400, 150-300, 200-400, 250-400, 300-400, 200-350, 200-300, 225-375, 250- 350, or 275-325 base pairs in length.
[0146] Alternatively, non-fragmented RNAs can be reverse transcribed, and the resultant cDNA can be fragmented. cDNA fragmentation methods can include physical fragmentation, chemical fragmentation, and / or enzymatic fragmentation. Physical fragmentation methods include, but are not limited to, acoustic or hydrodynamic shearing (such as sonication or point-sink shearing), needle shearing, and nebulization. Enzymatic fragmentation methods can include use of a restriction endonuclease (such as a 4-cutter or 5-cutter restriction endonuclease, e.g., Alul, Dpnl, Eco47I, Haelll, Hpall, Mbo I, Msel, MspI, PspGI, Rsal, Sse9I, or TaqI), a non-specific nuclease (e.g., micrococcal nuclease), or a transposase (for example, when insertion of an adapter into a fragmented double-stranded cDNA molecule is desired). cDNA may also be fragmented using chemical shearing methods. Chemical fragmentation methods can include, but are not limited to, heat digestion of cDNA in the presence of a divalent metal cation (such as magnesium or zinc). Insome embodiments, the fragmenting provides cDNA fragments of 25-400, 25-300, 25-200, 50- 400, 50-300, 50-250, 50-200, 100-400, 100-300, 100-200, 125-400, 125-300, 125-200, 125-175, 150-400, 150-300, 200-400, 250-400, 300-400, 200-350, 200-300, 225-375, 250-350, or 275-325 base pairs in length. cDNA preparation
[0147] Some embodiments of the disclosed methods comprise preparing cDNA from RNA (such as RNA extracted from a blood sample), such as by reverse transcription of the RNA template into cDNA. Reverse transcription is generally followed by exponential amplification of the cDNA, e.g., in a PCR reaction. Two commonly used reverse transcriptases are avian myeloblastosis vims reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the circumstances and the goal of expression profiling. For example, extracted RNA can be reverse transcribed using a Gene Amp RNA PCR kit (Perkin Elmer, Calif., USA), following the manufacturer's instructions. The derived cDNA can then be used as a template in the subsequent amplification (e.g., PCR) reaction. In some embodiments, RNA is converted to cDNA using random priming, followed by second strand synthesis, end repair, and optional A-tailing. Adapters comprising barcodes can then be ligated to the cDNA, which is then amplified.
[0148] Amplification is typically primed by primers that anneal or bind to primer binding sites in adapters flanking a cDNA molecule to be amplified. Amplification methods can involve cycles of denaturation, annealing and extension, resulting from thermocycling or can be isothermal as in transcription-mediated amplification. Other amplification methods include the ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and selfsustained sequence-based replication.
[0149] Although a PCR step can use a variety of thermostable DNA-dependent DNA polymerases, it typically employs the Taq DNA polymerase. TaqMan® PCR typically utilizes the 5'-nuclease activity of Taq or Tth polymerase to hydrolyze a hybridization probe bound to its target amplicon, but any enzyme with equivalent 5’ nuclease activity can be used. Two oligonucleotide primers are used to generate an amplicon typical of a PCR reaction. A third oligonucleotide, or probe, is designed to detect nucleotide sequence located between the two PCR primers. The probe is nonextendible by Taq DNA polymerase enzyme, and is labeled with a reporter fluorescent dye and a quencher fluorescent dye. Any laser-induced emission from the reporter dye is quenched by thequenching dye when the two dyes are located close together as they are on the probe. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a templatedependent manner. The resultant probe fragments disassociate in solution, and signal from the released reporter dye is free from the quenching effect of the second fluorophore. One molecule of reporter dye is liberated for each new molecule synthesized, and detection of the unquenched reporter dye provides the basis for quantitative interpretation of the data.
[0150] The primers used for the amplification are selected so as to amplify a unique segment of the gene of interest, such as RNA (such as mRNA) encoding a gene of a target gene set described herein. In some embodiments, expression of other genes is also detected, such as other known disease markers (such as known cancer markers) or housekeeping genes. Primers that can be used to amplify disease-related molecules are commercially available or can be designed and synthesized. In some examples, the primers specifically hybridize to a promoter or promoter region of a disease-related molecule. An alternative quantitative nucleic acid amplification procedure is described in U.S. Pat. No. 5,219,727. In this procedure, the amount of a target sequence in a sample is determined by simultaneously amplifying the target sequence and an internal standard nucleic acid segment. The amount of amplified cDNA from each segment is determined and compared to a standard curve to determine the amount of the target nucleic acid segment that was present in the sample prior to amplification. In some embodiments, the expression of a “housekeeping” gene or “internal control” can also be evaluated. These terms include any constitutively or globally expressed gene whose presence enables an assessment of mRNA levels provided herein. Such an assessment includes a determination of the overall constitutive level of gene transcription and a control for variations in RNA recovery. Exemplary housekeeping genes include tubulin, glyceraldehyde-3-phosphate-dehydrogenase (GAPDH), beta-actin, and 18S ribosomal RNA.Integrating molecular data with patient outcomes data, electronic health records and / or health insurance claims data
[0151] In various examples, the implementations described herein can integrate molecular data comprising transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, and / or H3S10ph with patientoutcomes data, electronic health records and / or health insurance claims data to: (1) Identify common methylation changes associated with cancer types or subtypes and stages. The goal of this analysis is to reveal potential methylation biomarkers for cancer diagnosis (2) Correlate methylation changes with clinical outcomes. The goal of this analysis is to understand the impact of these methylation changes on cancer progression, resistance, and treatment response. (3) Integrate the identified methylation changes with existing biological pathways and networks. The goal of this analysis is to uncover how methylation changes disrupt cellular processes and contribute to cancer development, resistance, and response. (4) Compare methylation data across different cancer types or subtypes to identify unique and shared mechanisms. The goal of this analysis is to understand cancer heterogeneity and similarities across different cancers. (5) Develop a predictive model using methylation data to predict treatment outcomes, recurrence, or drug resistance. The goal of this analysis is to help guide personalized treatment strategies. Such analysis can be achieved using a variety of statistical and machine learning models including Linear Regression and Logistic Regression: For continuous and binary outcomes, respectively, to model the relationship between methylation levels at specific sites and the presence or severity of cancer. Cox Proportional Hazards Model, can be useful for survival analysis to correlate methylation levels with the time to event data, such as time to cancer recurrence or progression. Mixed Models are helpful when the data comprises multiple measurements or hierarchical structures, mixed models can account for the correlation within subjects or groups. Multivariate Additionally, analysis like principal component analysis (PCA) or partial least squares regression (PLSR) can reduce dimensionality and identify patterns in methylation data that correlate with cancer types.
[0152] Other methods employed to explore relationships or correlations between methylation data and cancer type / subtype include machine learning models. For example, to handle complex, nonlinear relationships we employ Decision Trees and Random Forests these are particularly helpful in situations where the association between the methylation status of certain genes (or CpG sites) and cancer characteristics does not follow a straight-line pattern. For example, to model interaction effects i.e., methylation at one site might affect the impact of methylation at another site. Alone or in one or more combinations these models can identify specific methylation sites that are important for classifying cancer types. In some embodiments, Support Vector Machines (SVM) can be employed for classification tasks, including distinguishing between different types of cancer based on methylation patterns. In yet other embodiments, for example when the data is sufficiently large,deep learning approaches (e.g., convolutional neural networks for structured data like methylation arrays) can capture complex patterns including interactions in the data. Methods such as Gradient Boosting Machines (GBM) including models like XGBoost, LightGBM, and CatBoost can provide robust predictive models for cancer classification based on methylation data. These models include sequential addition of weak learners (e.g., decision trees) in such a way that each new tree corrects the errors made by the previous ones. GBMs can handle various types of data, including categorical and continuous variables. In additional embodiments, Cluster Analysis are employed to identify subgroups within cancer types that share similar methylation patterns. These include unsupervised learning models, including K-means clustering or hierarchical clustering.
[0153] In some embodiments, the cell-free DNA is from a subject having or suspected of having cancer and / or the cell-free DNA includes DNA from cancer cells. In some embodiments, the DNA is partitioned into a first subsample and a second subsample, wherein the first subsample comprises DNA with a nucleotide modification (e.g., a cytosine modification) in a greater proportion than the second subsample, and the first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, and the DNA is sequenced in a manner that distinguishes the first nucleobase from the second nucleobase in the DNA of the first subsample.Epigenetic therapeutic drugs
[0154] The intricate interplay of chemical modifications to DNA and histone proteins, known as the epigenome, plays a significant role in modulating gene expression in health and disease including cancer. Epigenetic changes in conjunction with genetic alterations contribute to the acquisition of cancer hallmarks such as sustaining proliferative signaling, evading growth suppressors, resisting cell death, enabling replicative immortality, inducing angiogenesis, and activating invasion and metastasi s(Hanahan, D. 2022). Given the reversible nature of epigenetic modifications, understanding these mechanisms offers promising avenues for therapeutic intervention, with several epigenetic drugs already approved or in clinical trials for the treatment of cancer.FDA approved epigenetic drugs
[0155] The FDA has approved several epigenetic drugs including DNA methyltransferase (DNMT) inhibitors, Histone deacetylase (HDAC) inhibitors, Lysine methyltransferase inhibitors, Lysine demethylase inhibitors, and Bromodomain inhibitors.
[0156] DNA Methyltransferase Inhibitors (DNMTi)- inhibit DNA methyltransferases, enzymes that add methyl groups to DNA, typically silencing gene expression. DNMTi drugs include Azacitidine which was approved for the treatment of myelodysplastic syndromes (MDS) and Decitabine also approved for MDS. Histone Deacetylase Inhibitors (HDACi) inhibit histone deacetylases, enzymes that remove acetyl groups from histone proteins, typically leading to a closed chromatin structure and gene silencing. Inhibiting these enzymes can reactivate silenced genes beneficial in cancer treatment. HDACi drugs include Vorinostat approved for the treatment of cutaneous T cell lymphoma (CTCL), Romidepsin approved for CTCL and peripheral T-cell lymphoma (PTCL), Belinostat approved for PTCL, and Panobinostat approved for multiple myeloma in combination with bortezomib and dexamethasone. EZH2 Inhibitors, EZH2 is a component of the polycomb repressive complex 2 (PRC2) that methylates histone H3 on lysine 27 (H3K27me3), leading to gene silencing. EZH2 Inhibitors include Tazemetostat approved for the treatment of epithelioid sarcoma and follicular lymphoma. Additional epigenetic drugs include Bromodomain Inhibitors that target bromodomains, which recognize acetylated lysine residues on histone tails, influencing chromatin structure and gene expression; however, currently there are no bromodomain inhibitors approved by the FDA.Partitioning the sample into a plurality of subsamples
[0157] In certain exemplary embodiments, described herein, a population of different forms of nucleic acids (e.g., hypermethylated and hypom ethylated DNA in a sample, such as a captured set of cfDNA as described herein) can be physically partitioned based on one or more characteristics of the nucleic acids prior to further analysis, e.g., differentially modifying or isolating a nucleobase, tagging, and / or sequencing. Additionally the population of different forms of nucleic acids may include transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated or contain a specific histone modification or functional element such as a TFBS. In some embodiments, hypermethylation variable epigenetic target regions are analyzed to determine whether they show hypermethylation characteristic of tumorcells and / or hypomethylation variable epigenetic target regions are analyzed to determine whether they show hypomethylation characteristic of tumor cells. Additionally, by partitioning a heterogeneous nucleic acid population, one may increase rare signals, e.g., by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, a genetic variation present in hyper-methylated DNA but less (or not) in hypomethylated DNA can be more easily detected by partitioning a sample into hyper-methylated and hypo- methylated nucleic acid molecules. By analyzing multiple fractions of a sample, a multidimensional analysis of a single locus of a genome or species of nucleic acid can be performed and hence, greater sensitivity can be achieved. In some embodiments of the disclosure each partition may comprise different forms of nucleic acids including transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated or contain a specific histone modification or functional element such as a TFBS.
[0158] In some instances, a heterogeneous nucleic acid sample is partitioned into two or more partitions. For instance, a minimum of three partitions, extending to four, five, six, seven, up to any number of partitions, where the total number can be any non-negative real number). In some embodiments, each partition is differentially tagged. Tagged partitions can then be pooled together for collective sample prep and / or sequencing. The partitioning-tagging-pooling steps can occur more than once, with each round of partitioning occurring based on a different characteristics (examples provided herein), and tagged using differential tags that are distinguished from other partitions and partitioning means.
[0159] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. Additional characteristics that can be used for partitioning include, transcription factor binding sites (TFBS), fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel,H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph, type of nucleic acid for example RNA, mRNA, cDNA, or DNA. Resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments and longer DNA fragments. In some embodiments, partitioning based on a cytosine modification (e.g., cytosine methylation) or methylation generally is performed and is optionally combined with at least one additional partitioning step, which may be based on any of the foregoing characteristics or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and without the one or more epigenetic modifications. Examples of epigenetic modifications include presence or absence of methylation; level of methylation; type of methylation (e.g., 5- methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and association and level of association with one or more proteins, such as histones. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules devoid of nucleosomes. Alternatively or additionally, a heterogeneous population of nucleic acids may be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, or additionally, a heterogeneous population of nucleic acids may be partitioned based on nucleic acid length (e.g., molecules of up to 160 bp and molecules having a length of greater than 160 bp).
[0160] In some instances, each partition (representative of a different nucleic acid form) is differentially labelled, and the partitions are pooled together prior to sequencing. In other instances, the different forms are separately sequenced.
[0161] In some embodiments, a population of different nucleic acids is partitioned into two or more different partitions. Each partition is representative of a different nucleic acid form, and a first partition (also referred to as a subsample) comprises DNA with a cytosine modification in a greater proportion than a second subsample. Each partition is distinctly tagged. The first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The tagged nucleic acids are pooled together prior to sequencing. Sequence reads are obtained and analyzed, including to distinguish the first nucleobase from the secondnucleobase in the DNA of the first subsample, in silico. Tags are used to sort reads from different partitions. Analysis to detect genetic variants, a variety of epigenetic marks or functional elements for example TFBS, CTCF binding sites can be performed on a partiti on-by-partition level, as well as whole nucleic acid population level. For example, analysis can include in silico analysis to determine genetic variants, such as CNV, SNV, indel, fusion in nucleic acids in each partition. Additionally, analysis can include in silico analysis to determine transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. In some instances, in silico analysis can include determining chromatin structure. For example, coverage of sequence reads can be used to determine nucleosome positioning in chromatin. Higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or nucleosome depleted region (NDR).
[0162] Samples can include nucleic acids varying in modifications including post-replication modifications to nucleotides and binding, usually noncovalently, to one or more proteins.
[0163] In an embodiment, the population of nucleic acids is one obtained from a serum, plasma or blood sample from a subject suspected of having neoplasia, a tumor, or cancer or previously diagnosed with neoplasia, a tumor, or cancer. The population of nucleic acids includes nucleic acids having varying levels of methylation. Methylation can occur from any one or more postreplication or transcriptional modifications. Post-replication modifications include modifications of the nucleotide cytosine, particularly at the 5-position of the nucleobase, e.g., 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine.
[0164] The affinity agents can be antibodies with the desired specificity, natural binding partners or variants thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or artificial peptides selected e.g., by phage display to have specificity to a given target.
[0165] Examples of capture moieties contemplated herein include methyl binding domain (MBDs) and methyl binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies preferentially binding to 5-methylcytosine.
[0166] Likewise, partitioning of different forms of nucleic acids can be performed using histone binding proteins which can separate nucleic acids bound to histones from free or unbound nucleic acids. Examples of histone binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48 and SANT domain peptides.
[0167] Although for some affinity agents and modifications, binding to the agent may occur in an essentially all or none manner depending on whether a nucleic acid bears a modification, the separation may be one of degree. In such instances, nucleic acids overrepresented in a modification bind to the agent at a greater extent that nucleic acids underrepresented in the modification. Alternatively, nucleic acids having modifications may bind in an all or nothing manner. But then, various levels of modifications may be sequentially eluted from the binding agent.
[0168] For example, in some embodiments, partitioning can be binary or based on degree / level of modifications. For example, all methylated fragments can be partitioned from unmethylated fragments using methyl-binding domain proteins (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequently, additional partitioning may involve eluting fragments having different levels of methylation by adjusting the salt concentration in a solution with the methyl-binding domain and bound fragments. As salt concentration increases, fragments having greater methylation levels are eluted.
[0169] In some instances, the final partitions are representative of nucleic acids having different extents of modifications (over representative or under representative of modifications). Overrepresentation and underrepresentation can be defined by the number of modifications bom by a nucleic acid relative to the median number of modifications per strand in a population. For example, if the median number of 5-methylcytosine residues in nucleic acid in a sample is 2, a nucleic acid including more than two 5-methylcytosine residues is overrepresented in this modification and a nucleic acid with 1 or zero 5-methylcytosine residues is underrepresented. The effect of the affinity separation is to enrich for nucleic acids overrepresented in a modification in a bound phase and for nucleic acids underrepresented in a modification in an unbound phase (i.e. in solution). The nucleic acids in the bound phase can be eluted before subsequent processing.
[0170] When using MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific) various levels of methylation can be partitioned using sequential elutions. For example, a hypomethylated partition (e.g., no methylation) can be separated from a methylated partition by contacting the nucleic acid population with the MBD from the kit, which is attached to magnetic beads. Similarly, histone marks, or TFBS can be can be separated using appropriate antibodiesspecific to the biochemical properties of each molecule, for example Anti-CTCF, Anti-Pol II (RNA Polymerase II), Anti-H3K4me3 (Histone H3 tri-methylated at lysine 4), Anti-H3K27me3 (Histone H3 tri-methylated at lysine 27), Anti-H3K9me3 (Histone H3 tri-methylated at lysine 9), Anti- H3K27ac (Histone H3 acetylated at lysine 27), Anti-H3K4mel (Histone H3 mono-methylated at lysine 4), Anti-H3K36me3 (Histone H3 tri-methylated at lysine 36). The beads are used to separate out the methylated nucleic acids from the non- methylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids having different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids are eluted, magnetic separation is once again used to separate higher level of methylated nucleic acids from those with lower level of methylation. The elution and magnetic separation steps can repeat themselves to create various partitions such as a hypom ethylated partition (representative of no methylation), a methylated partition (representative of low level of methylation), and a hyper methylated partition (representative of high level of methylation).
[0171] In some methods, nucleic acids bound to an agent used for affinity separation are subjected to a wash step. The wash step washes off nucleic acids weakly bound to the affinity agent. Such nucleic acids can be enriched in nucleic acids having the modification to an extent close to the mean or median (i.e., intermediate between nucleic acids remaining bound to the solid phase and nucleic acids not binding to the solid phase on initial contacting of the sample with the agent).
[0172] The affinity separation results in at least two, and sometimes three or more partitions of nucleic acids with different extents of a modification. While the partitions are still separate, the nucleic acids of at least one partition, and usually two or three (or more) partitions are linked to nucleic acid tags, usually provided as components of adapters, with the nucleic acids in different partitions receiving different tags that distinguish members of one partition from another. The tags linked to nucleic acid molecules of the same partition can be the same or different from one another. But if different from one another, the tags may have part of their code in common so as to identify the molecules to which they are attached as being of a particular partition.
[0173] For further details regarding portioning nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference.
[0174] In some embodiments, the nucleic acid molecules can be fractionated into different partitions based on the nucleic acid molecules that are bound to a specific protein or a fragment thereof and those that are not bound to that specific protein or fragment thereof.
[0175] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on a specific property of a protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation) or enzymatic activity. Examples of proteins which may bind to DNA and serve as a basis for fractionation may include, but are not limited to, protein A and protein G. Any suitable method can be used to fractionate the nucleic acid molecules based on protein bound regions. Examples of methods used to fractionate nucleic acid molecules based on protein bound regions include, but are not limited to, SDS-PAGE, chromatin-immuno-precipitation (ChIP), heparin chromatography, and asymmetrical field flow fractionation (AF4).
[0176] In certain exemplary embodiments, partitioning of the nucleic acids is performed by contacting the nucleic acids with a methylation binding domain (“MBD”) of a methylation binding protein (“MBP”). MBD binds to 5-methylcytosine (5mC). MBD is coupled to paramagnetic beads, such as Dynabeads® M-280 Streptavidin via a biotin linker. Partitioning into fractions with different extents of methylation can be performed by eluting fractions by increasing the NaCl concentration.
[0177] Examples of MBPs contemplated herein include, but are not limited to:(a) MeCP2 is a protein preferentially binding to 5-methyl-cytosine over unmodified cytosine.(b) RPL26, PRP8 and the DNA mismatch repair protein MHS6 preferentially bind to 5- hydroxymethyl-cytosine over unmodified cytosine.(c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferably bind to 5-formyl-cytosine over unmodified cytosine (lurlaro et al., Genome Biol. 14: R1 19 (2013)).(d) Antibodies specific to one or more methylated nucleotide bases.
[0178] In general, elution is a function of number of methylated sites per molecule, with molecules having more methylation eluting under increased salt concentrations. To elute the DNA into distinct populations based on the extent of methylation, one can use a series of elution buffers of increasing NaCl concentration. Salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process results in three (3) partitions. Molecules are contacted with a solution at a first salt concentration and comprising a molecule comprising a methyl bindingdomain, which molecule can be attached to a capture moiety, such as streptavidin. At the first salt concentration a population of molecules will bind to the MBD and a population will remain unbound. The unbound population can be separated as a “hypom ethylated” population. For example, a first partition representative of the hypomethylated form of DNA is that which remains unbound at a low salt concentration, e.g., 100 mM or 160 mM. A second partition representative of intermediate methylated DNA is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM concentration. This is also separated from the sample. A third partition representative of hypermethylated form of DNA is eluted using a high salt concentration, e.g., at least about 2000 mM.Adapter ligation or addition; tagging
[0179] In certain exemplary embodiments, the disclosed methods comprise adding adapters to DNA (such as cDNA, cell-free DNA, or fragmented genomic DNA). In some embodiments, adapters are added to the DNA before or after subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as after the subjecting. When adapters are added to the DNA before subjecting the DNA to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, they may comprise nucleotides that are resistant to the procedure. For example, where the procedure comprises contacting the DNA with a deaminase, the adapters may comprise deaminase-resistant cytosines such as 5-ghmC or 5-propynyl cytosine. Similarly, where the procedure comprises contacting the DNA with bisulfite, the adapters may comprise bi sulfite-resistant cytosines such as 5-mC or 5-hmC. In some embodiments, adapters may be added or to DNA concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer (where PCR is used, this can be referred to as library prep-PCR or LP-PCR), before or after an amplification step. In some embodiments, adapters are added by other approaches. In some such methods, first adapters are added to the nucleic acids by ligation to the 3’ ends thereof, which may include ligation to single-stranded DNA. The adapter can be used as a priming site for second-strand synthesis, e.g., using a universal primer and a DNA polymerase. A second adapter can then be ligated to at least the 3’ end of the second strand of the now double-stranded molecule. In some embodiments, the first adapter comprises an affinity tag, such as biotin, and nucleic acid ligated to the first adapter is bound to a solid support (e.g., bead), which may comprise a binding partner for the affinity tag such as streptavidin. For further discussion of a related procedure, see Gansauge et al., Nature Protocols 8:737-748 (2013). Commercial kits for sequencing library preparationcompatible with single-stranded nucleic acids are available, e.g., the Accel-NGS® Methyl-Seq DNA Library Kit from Swift Biosciences. In some embodiments, after adapter ligation, nucleic acids are amplified. In some embodiments, end repair of the DNA is performed prior to addition of adapters.
[0180] In some embodiments, the single-stranded DNA library preparation is performed in a one- step combined phosphorylation / ligation reaction, e g., as described in Troll et al., BMC Genomics, 20: 1023 (2019), available at https: / / doi.org / 10.1186 / sl2864-019-6355-0. This method, called Single Reaction Single-stranded LibrarY (“SRSLY,”) can be performed without end-polishing. SRSLY may be useful for converting short and fragmented DNA molecules, e.g., cfDNA fragments, into sequencing libraries while retaining native lengths and ends. The SRSLY method can create sequencing libraries (e.g., Illumina sequencing libraries) from fragmented or degraded template (input) DNA. In particular embodiments, template DNA is first heat denatured and then immediately cold shocked to render the template DNA molecules single-stranded. The DNA can be maintained as single-stranded throughout the ligation reaction by the inclusion of a thermostable single-stranded binding protein (SSB). Next, the template DNA, which at this point can be singlestranded and coated with SSB, is placed in a phosphorylation / ligation dual reaction with directional dsDNA NGS adapters that contain single-stranded overhangs. Both the forward and reverse sequencing adapters can share similar structures but differ in which termini is unblocked in order to facilitate proper ligations. Both sequencing adapters can comprise a dsDNA portion and a single-stranded splint overhang of random nucleotides that occurs on the 3 -prime terminus of the bottom strand of the forward adapter and the 5-prime terminus of the bottom strand of the reverse adapter. In this way, the forward adapter (e.g., (P5) Illumina adapter) can delivered to the 5-prime end of template molecules and the reverse adapter (e.g., (P7) Illumina adapter) is delivered to the 3-prime end of template molecules. Thus, the native polarity of input DNA molecules can be retained.
[0181] During the dual phosphorylation / ligation reaction, T4 Polynucleotide Kinase (PNK) can be used to prepare template DNA termini for ligation by phosphorylating 5-prime termini and dephosphorylating 3-prime termini. T4 PNK works on both ssDNA and dsDNA molecules and has no activity on the phosphorylation state of proteins. Simultaneously, the random nucleotides of the splint adapter can be annealed to the single- stranded template molecule. This creates a short, localized dsDNA molecule, enabling ligation of template to adapter with a ligase such as T4 DNA ligase, which has high ligation efficiency on dsDNA templates but low efficiency on ssDNA. Afterthe single phosphorylation / ligation reaction is complete, the library DNA can be, e.g., purified and placed directly into standard NGS indexing PCR, compatible with both traditional single or dual index primers.
[0182] In some embodiments, following attachment of adapters, the nucleic acids are subject to amplification. The amplification can use, e.g., universal primers that recognize primer binding sites in the adapters.
[0183] In some embodiments, the DNA is linked at both ends to Y-shaped adapters including primer binding sites and tags. In some such embodiments, the DNA is amplified.
[0184] In embodiments of the disclosed methods, a target nucleic acid comprises a 5’ adapter, a 3’ adapter, or both a 5’ adapter and a 3’ adapter. In such embodiments, the 5’ adapter, or both the 5’ adapter and the 3’ adapter comprise at least one sequence that is recognized by at least one restriction enzyme, such as a restriction enzyme described elsewhere herein. In particular embodiments, the 5’ adapter is downstream of an oligonucleotide probe binding site within the target nucleic acid.
[0185] Tagging DNA molecules is a procedure in which a tag is attached to or associated with the DNA molecules. Such tags can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. Tags can allow one to differentiate molecules from which sequence reads originated. For example, molecules can bear a sample tag (which distinguishes molecules in one sample from those in a different sample) or a molecular tag / molecular barcod e / barcode (which distinguishes different molecules from one another in both unique and non-unique tagging scenarios). For methods that involve a partitioning step, a partition tag (which distinguishes molecules in one partition from those in a different partition) may be included. In some embodiments, adapters added to DNA molecules comprise tags. In some such embodiments, the tag comprises one or a combination of barcodes. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or can have sequences having a certain hamming distance, as desired for the specific purpose. So, for example, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally or alternatively, for different partitions and / or samples, different sets of molecular barcodes, or molecular tags can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve toidentify the partition and / or sample to which they correspond based the set of which they are a member. Tags comprising barcodes can be incorporated into or otherwise joined to adapters. Tags can be incorporated by ligation, overlap extension PCR among other methods.
[0186] Tagging strategies can be divided into unique tagging and non-unique tagging strategies. In unique tagging, all or substantially all of the molecules in a sample bear a different tag, so that reads can be assigned to original molecules based on tag information alone. Tags used in such methods are sometimes referred to as “unique tags”. In non-unique tagging, different molecules in the same sample can bear the same tag, so that other information in addition to tag information is used to assign a sequence read to an original molecule. Such information may include start and stop coordinate, coordinate to which the molecule maps, start or stop coordinate alone, etc. Tags used in such methods are sometimes referred to as “non-unique tags”. Accordingly, it is not necessary to uniquely tag every molecule in a sample. It suffices to uniquely tag molecules falling within an identifiable class within a sample. Thus, molecules in different identifiable families can bear the same tag without loss of information about the identity of the tagged molecule.
[0187] In some embodiments, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites. In some embodiments, adapters include the same primer binding site.
[0188] In certain embodiments of non-unique tagging, the number of different tags used can be sufficient that there is a very high likelihood (e.g., at least 99%, at least 99.9%, at least 99.99% or at least 99.999% that all molecules of a particular group bear a different tag. In some embodiments comprising barcode attachment, e.g., randomly, to both ends of a molecule, the combination of barcodes, together, constitutes a tag. This number, in term, is a function of the number of molecules falling into the calls. For example, the class may be all molecules mapping to the same start-stop position on a reference genome. The class may be all molecules mapping across a particular genetic locus, e.g., a particular base or a particular region (e.g., up to 100 bases or a gene or an exon of a gene). In certain embodiments, the number of different tags used to uniquely identify a number of molecules, z, in a class can be between any of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11 *z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z or 100*z (e.g., lower limit) and any of 100,000*z, 10,000*z, 1000*z or 100*z (e.g., upper limit).
[0189] For example, in a sample of about 5 ng to 30 ng of DNA, one expects around 3000 molecules to map to a particular nucleotide coordinate, and between about 3 and 10 molecules having any start coordinate to share the same stop coordinate. Accordingly, about 50 to about 50,000 different tags (e.g., between about 6 and 220 barcode combinations) can suffice to uniquely tag all such molecules. To uniquely tag all 3000 molecules mapping across a nucleotide coordinate, about 1 million to about 20 million different tags would be required.
[0190] Generally, assignment of unique or non-unique tags barcodes in reactions follows methods and systems described by US patent applications 20010053519, 20030152490, 20110160078, and U.S. Pat. No. 6,582,908 and U.S. Pat. No. 7,537,898 and US Pat. No. 9,598,731. Tags can be linked to sample nucleic acids randomly or non-randomly.
[0191] In some embodiments, the tagged nucleic acids are sequenced after loading into a microwell plate. The microwell plate can have 96, 384, or 1536 microwells. In some cases, they are introduced at an expected ratio of unique tags to microwells. For example, the unique tags may be loaded so that more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genome sample. In some cases, the unique tags may be loaded so that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags are loaded per genome sample. In some cases, the average number of unique tags loaded per sample genome is less than, or greater than, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 unique tags per genome sample
[0192] In some embodiments, 20-50 different tags (e.g., barcodes) are ligated to both ends of target nucleic acids. For example, 35 different tags (e.g., barcodes) ligated to both ends of target molecules creating 35 x 35 permutations, which equals 1225 for 35 tags. Such numbers of tags are sufficient so that different molecules having the same start and stop points have a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) of receiving different combinations of tags. Other barcode combinations include any number between 10 and 500, e.g., about 15x15, about 35x35, about 75x75, about 100x100, about 250x250, about 500x500.
[0193] In some cases, unique tags may be predetermined or random or semi-random sequence oligonucleotides. In other cases, a plurality of barcodes may be used such that barcodes are not necessarily unique to one another in the plurality. In this example, barcodes may be ligated to individual molecules such that the combination of the barcode and the sequence it may be ligatedto creates a unique sequence that may be individually tracked. As described herein, detection of non-unique barcodes in combination with sequence data of beginning (start) and end (stop) portions of sequence reads may allow assignment of a unique identity to a particular molecule. The length or number of base pairs, of an individual sequence read may also be used to assign a unique identity to such a molecule. As described herein, fragments from a single strand of nucleic acid having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand.Tagging of partitions
[0194] In certain exemplary embodiments, two or more partitions, e.g., each partition, is / are differentially tagged. Tags or indexes can be molecules, such as nucleic acids, containing information that indicates a feature of the molecule with which the tag is associated. For example, molecules can bear a sample tag or sample index (which distinguishes molecules in one sample from those in a different sample), a partition tag (which distinguishes molecules in one partition from those in a different partition) and / or a molecular tag / molecular barcode / barcode (which distinguishes different molecules from one another (in both unique and non-unique tagging scenarios). In certain embodiments, a tag can comprise one or a combination of barcodes. As used herein, the term “barcode” refers to a nucleic acid molecule having a particular nucleotide sequence, or to the nucleotide sequence, itself, depending on context. A barcode can have, for example, between 10 and 100 nucleotides. A collection of barcodes can have degenerate sequences or can have sequences having a certain Hamming distance, as desired for the specific purpose. So, for example, a molecular barcode can be comprised of one barcode or a combination of two barcodes, each attached to different ends of a molecule. Additionally or alternatively, for different partitions and / or samples, different sets of molecular barcodes, molecular tags, or molecular indexes can be used such that the barcodes serve as a molecular tag through their individual sequences and also serve to identify the partition and / or sample to which they correspond based the set of which they are a member.
[0195] Tags can be used to label the individual polynucleotide population partitions so as to correlate the tag (or tags) with a specific partition. Alternatively, tags can be used in embodiments of the invention that do not employ a partitioning step. In some embodiments, a single tag can be used to label a specific partition. In some embodiments, multiple different tags can be used to label a specific partition. In embodiments employing multiple different tags to label a specific partition, the set of tags used to label one partition can be readily differentiated for the set of tags used tolabel other partitions. In some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as unique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations, for example as in Kinde et al., Proc Nat’l Acad Sci USA 108: 9530-9535 (2011), Kou et al., PLoS 0NE,W. e0146638 (2016)) or used as non-unique molecule identifiers, for example as described in US Pat. No. 9,598,731. Similarly, in some embodiments, the tags may have additional functions, for example the tags can be used to index sample sources or used as nonunique molecular identifiers (which can be used to improve the quality of sequencing data by differentiating sequencing errors from mutations).
[0196] In one embodiment, partition tagging comprises tagging molecules in each partition with a partition tag. After re-combining partitions (e.g., to reduce the number of sequencing runs needed and avoid unnecessary cost) and sequencing molecules, the partition tags identify the source partition. In another embodiment, different partitions are tagged with different sets of molecular tags, e g., comprised of a pair of barcodes. In this way, each molecular barcode indicates the source partition as well as being useful to distinguish molecules within a partition. For example, a first set of 35 barcodes can be used to tag molecules in a first partition, while a second set of 35 barcodes can be used tag molecules in a second partition.
[0197] In some embodiments, after partitioning and tagging with partition tags, the molecules may be pooled for sequencing in a single run. In some embodiments, a sample tag is added to the molecules, e.g., in a step subsequent to addition of partition tags and pooling. Sample tags can facilitate pooling material generated from multiple samples for sequencing in a single sequencing run.
[0198] Alternatively, in some embodiments, partition tags may be correlated to the sample as well as the partition. As a simple example, a first tag can indicate a first partition of a first sample; a second tag can indicate a second partition of the first sample; a third tag can indicate a first partition of a second sample; and a fourth tag can indicate a second partition of the second sample.
[0199] While tags may be attached to molecules already partitioned based on one or more characteristics for example transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3 Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3,H3S10ph, the final tagged molecules in the library may no longer possess that characteristic. In additional examples, while single stranded DNA molecules may be partitioned and tagged, the final tagged molecules in the library are likely to be double stranded. Similarly, while DNA may be subject to partition based on different levels of methylation, in the final library, tagged molecules derived from these molecules are likely to be unmethylated. Accordingly, the tag attached to molecule in the library typically indicates the characteristic of the “parent molecule” from which the ultimate tagged molecule is derived, not necessarily to characteristic of the tagged molecule, itself.
[0200] As an example, barcodes 1, 2, 3, 4. . .n, etc. are used to tag and label molecules in the first partition; barcodes A, B, C, D, etc. are used to tag and label molecules in the second partition; and barcodes a, b, c, d, etc. are used to tag and label molecules in the third partition. Differentially tagged partitions can be pooled prior to sequencing. Differentially tagged partitions can be separately sequenced or sequenced together concurrently, e.g., in the same flow cell of an Illumina sequencer.
[0201] After sequencing, analysis of reads to detect genetic variants can be performed on a partition-by-partition level, as well as a whole nucleic acid population level. Tags are used to sort reads from different partitions. Analysis can include in silico analysis to determine genetic and epigenetic variation (one or more of methylation, chromatin structure, etc.) using sequence information, genomic coordinates length, coverage, and / or copy number. In some embodiments, higher coverage can correlate with higher nucleosome occupancy in genomic region while lower coverage can correlate with lower nucleosome occupancy or a nucleosome depleted region (NDR).Alternative Methods of Modified Nucleic Acid Analysis
[0202] In some embodiments the adapters are added to the nucleic acids after partitioning the nucleic acids, in other embodiments the adapters may be added to the nucleic acids prior to partitioning the nucleic acids based on molecular characteristics for example transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. In some such methods, a population of nucleic acids bearing the modification to different extents (e.g., 0, 1, 2, 3, 4, or more methyl groups per nucleic acid molecule) is contacted with adapters before fractionation of thepopulation depending on the extent of the modification. Adapters attach to either one end or both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. Adapters, whether bearing the same or different tags, can include the same or different primer binding sites, but preferably adapters include the same primer binding site. Following attachment of adapters, the nucleic acids are contacted with an agent that preferentially binds to nucleic acids bearing the modification (such as the previously described such agents). The nucleic acids are partitioned into at least two subsamples differing in the extent to which the nucleic acids bear the modification from binding to the agents. For example, if the agent has affinity for nucleic acids bearing the modification, nucleic acids overrepresented in the modification (compared with median representation in the population) preferentially bind to the agent, whereas nucleic acids underrepresented for the modification do not bind or are more easily eluted from the agent. Following partitioning, the first subsample is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. The nucleic acids are then amplified from primers binding to the primer binding sites within the adapters. Following amplification, the different partitions can then be subject to further processing steps, which typically include further (e.g., clonal) amplification, and sequence analysis, in parallel but separately. Sequence data from the different partitions can then be compared.
[0203] In another embodiment, a partitioning scheme can be performed using the following exemplary procedure. Nucleic acids are linked at both ends to Y-shaped adapters including primer binding sites and tags. The molecules are amplified. The amplified molecules are then fractionated by contact with an antibody preferentially binding to 5-methylcytosine to produce two partitions. One partition includes original molecules lacking methylation and amplification copies having lost methylation. The other partition includes original DNA molecules with methylation. The partition including original DNA molecules with methylation is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobaseand the second nucleobase have the same base pairing specificity. The two partitions are then processed and sequenced separately with further amplification of the methylated partition. The sequence data of the two partitions can then be compared. In this example, tags are not used to distinguish between methylated and unmethylated DNA but rather to distinguish between different molecules within these partitions so that one can determine whether reads with the same start and stop points are based on the same or different molecules.
[0204] The disclosure provides further methods for analyzing a population of nucleic acids in which at least some of the nucleic acids include one or more modified cytosine residues, such as 5-methylcytosine and any of the other modifications described previously. In these methods, after partitioning, the subsamples of nucleic acids are contacted with adapters including one or more cytosine residues modified at the 5C position, such as 5-methylcytosine. Preferably all cytosine residues in such adapters are also modified, or all such cytosines in a primer binding region of the adapters are modified. Adapters attach to both ends of nucleic acid molecules in the population. Preferably, the adapters include different tags of sufficient numbers that the number of combinations of tags results in a low probability e.g., 95, 99 or 99.9% of two nucleic acids with the same start and stop points receiving the same combination of tags. The primer binding sites in such adapters can be the same or different, but are preferably the same. After attachment of adapters, the nucleic acids are amplified from primers binding to the primer binding sites of the adapters. The amplified nucleic acids are split into first and second aliquots. The first aliquot is assayed for sequence data with or without further processing. The sequence data on molecules in the first aliquot is thus determined irrespective of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase comprises a cytosine modified at the 5 position, and the second nucleobase comprises unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. The nucleic acids subjected to the procedure are then amplified with primers to the original primer binding sites of the adapters linked to nucleic acid. Only the nucleic acid molecules originally linked to adapters (as distinct from amplification products thereof) are now amplifiable because these nucleic acids retain cytosines in the primer binding sites of the adapters, whereas amplification products have lost the methylation of these cytosine residues, which have undergone conversion to uracils in the bisulfite treatment. Thus, only original molecules in the populations, at least some of which are methylated, undergoamplification. After amplification, these nucleic acids are subject to sequence analysis. Comparison of sequences determined from the first and second aliquots can indicate among other things, which cytosines in the nucleic acid population were subject to methylation.
[0205] Such an analysis can be performed using the following exemplary procedure. After partitioning, methylated DNA is linked to Y -shaped adapters at both ends including primer binding sites and tags. The cytosines in the adapters are modified at the 5 position (e.g., 5-methylated). The modification of the adapters serves to protect the primer binding sites in a subsequent conversion step (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but affects unmodified cytosine). After attachment of adapters, the DNA molecules are amplified. The amplification product is split into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequence analysis with or without further processing. The other aliquot is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase comprises a cytosine modified at the 5 position, and the second nucleobase comprises unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosines to uracils. Only primer binding sites protected by modification of cytosines can support amplification when contacted with primers specific for original primer binding sites. Thus, only original molecules and not copies from the first amplification are subjected to further amplification. The further amplified molecules are then subjected to sequence analysis. Sequences can then be compared from the two aliquots. As in the separation scheme discussed above, nucleic acid tags in adapters are not used to distinguish between methylated and unmethylated DNA but to distinguish nucleic acid molecules within the same partition.Subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample
[0206] In some embodiments, methods disclosed herein comprise a step of subjecting DNA, or a subsample thereof, to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the procedure chemically converts the first or second nucleobase such that the base pairing specificity of the converted nucleobase is altered. In someembodiments, DNA is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA before library preparation using the DNA, before a first amplification of the DNA, before dividing the DNA into a plurality of subsamples, or any combination thereof. In certain embodiments, the DNA is subjected to the procedure before or after contacting the DNA with a methylation-sensitive nuclease.
[0207] In some embodiments, the procedure that affects a first nucleobase of the DNA differently from a second nucleobase of the DNA is performed prior to the sequencing and / or (a) prior to or after the selectively depleting the target nucleic acid comprising the wild-type sequence, the target nucleic acid comprising the converted nucleotide, or the target nucleic acid that does not comprise the converted nucleotide; (b) prior to the amplifying the selectively digested population of target nucleic acids; (c) prior to or after the partitioning the population of target nucleic acids into a plurality of subsamples; and / or (d) prior to or after a step of enriching for one or more sets of target regions of DNA.
[0208] In some embodiments, if the first nucleobase is a modified or unmodified adenine, then the second nucleobase is a modified or unmodified adenine; if the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine; if the first nucleobase is a modified or unmodified guanine, then the second nucleobase is a modified or unmodified guanine; and if the first nucleobase is a modified or unmodified thymine, then the second nucleobase is a modified or unmodified thymine (where modified and unmodified uracil are encompassed within modified thymine for the purpose of this step).
[0209] In some embodiments, the first nucleobase is a modified or unmodified cytosine, then the second nucleobase is a modified or unmodified cytosine. For example, first nucleobase may comprise unmodified cytosine (C) and the second nucleobase may comprise one or more of 5- methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.
[0210] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g. 5-formyl cytosine (fC) or 5 -carboxylcytosine (caC)) to uracil whereas other modified cytosines (e.g., 5- methylcytosine, 5-hydroxylmethylcystosine) are not converted. Thus, where bisulfite conversionis used, the first nucleobase comprises one or more of unmodified cytosine, 5-formyl cytosine, 5- carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of bisulfite- treated DNA identifies positions that are read as cytosine as being mC or hmC positions. Meanwhile, positions that are read as T are identified as being T or a bisulfite-susceptible form of C, such as unmodified cytosine, 5-formyl cytosine, or 5-carboxylcytosine. Performing bisulfite conversion, such as on a DNA sample as described herein, facilitates identifying positions containing mC or hmC using the sequence reads obtained from the exemplary sample. For an exemplary description of bisulfite conversion, see, e.g., Moss et al., Nat Commun. 2018; 9: 5068.
[0211] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises oxidative bisulfite (Ox-BS) conversion. This procedure first converts hmC to fC, which is bisulfite susceptible, followed by bisulfite conversion. Thus, when oxidative bisulfite conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, hmC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises mC. Sequencing of Ox-BS converted DNA identifies positions that are read as cytosine as being mC positions. Meanwhile, positions that are read as T are identified as being T, hmC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or hmC. Performing Ox-BS conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing mC using the sequence reads obtained from the sample. For an exemplary description of oxidative bisulfite conversion, see, e.g., Booth et al., Science 2012; 336: 934-937.
[0212] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted bisulfite (TAB) conversion. In TAB conversion, hmC is protected from conversion and mC is oxidized in advance of bisulfite treatment, so that positions originally occupied by mC are converted to U while positions originally occupied by hmC remain as a protected form of cytosine. For example, as described in Yu et al., Cell 2012; 149: 1368-80, -glucosyl transferase can be used to protect hmC (forming 5- glucosylhydroxymethylcytosine (ghmC)), then a TET protein such as mTetl can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while ghmC remains unaffected.
[0213] Alternatively, a carbamoyltransferase enzyme, such as 5-hydroxymethylcytosine carbamoyltransferase as described in Yang et al., Bio-protocol, 2023; 12(17): e4496, can be used to protect hmC (by converting hmC to 5-carbamoyloxymethylcytosine (5cmC)), then a TETprotein such as mTetl or a TET2 comprising a T1372S mutation, can be used to convert mC to caC, and then bisulfite treatment can be used to convert C and caC to U while 5cmC remains unaffected. Thus, when TAB conversion is used, the first nucleobase comprises one or more of unmodified cytosine, fC, caC, mC, or other cytosine forms affected by bisulfite, and the second nucleobase comprises hmC. Sequencing of TAB -converted DNA identifies positions that are read as cytosine as being hmC positions. Meanwhile, positions that are read as T are identified as being T, mC, or a bisulfite-susceptible form of C, such as unmodified cytosine, fC, or caC. Performing TAB conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing hmC using the sequence reads obtained from the sample.
[0214] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises Tet-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In Tet-assisted pic-borane conversion with a substituted borane reducing agent conversion, a TET protein is used to convert mC and hmC to caC, without affecting unmodified C. caC, and fC if present, are then converted to dihydrouracil (DHU) by treatment with 2-picoline borane (pic-borane) or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting unmodified C. See, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429 (e.g., at Supplementary Fig. 1 and Supplementary Note 7). Thus, when this type of conversion is used, the first nucleobase comprises one or more of 5mC, 5fC, 5caC, or 5hmC, and the second nucleobase comprises unmodified cytosine. DHU is read as a T in sequencing. Thus, when this type of conversion is used, the first nucleobase comprises one or more of mC, f'C, caC, or hmC, and the second nucleobase comprises unmodified cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T, mC, fC, caC, or hmC. Performing TAP conversion, such as on a DNA sample as described herein, thus facilitates identifying positions containing unmodified C using the sequence reads obtained from the sample. This procedure encompasses Tet-assisted pyridine borane sequencing (TAPS), described in further detail in Liu et al. 2019, supra.
[0215] Alternatively, protection of hmC (e.g., using 0GT or 5-hydroxymethylcytosine carbamoyltransferase) can be combined with Tet-assisted conversion with a substituted borane reducing agent, e.g. as described above. In this method (TAPS-0), 5hmC can be protected fromconversion, for example through glucosylation using P-glucosyl transferase (PGT), forming 5- glucosylhydroxymethylcytosine (5ghmC), or through carbamoylation using 5- hydroxymethylcytosine carbamoyltransferase, forming 5cmC. This is described in Yu et al., Cell 2012; 149: 1368-80. Treatment with a TET protein, such as mTetl ora TET2 comprising a T1372S mutation, then converts mC to caC but does not convert C, 5ghmC, or 5cmC. 5caC is then converted to DHU by treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane, also without affecting ghmC, 5cmC, or unmodified C. Thus, when Tet-assisted conversion with a substituted borane reducing agent is used, the first nucleobase comprises mC, and the second nucleobase comprises one or more of unmodified cytosine or hmC, such as unmodified cytosine and optionally hmC, fC, and / or caC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either hmC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or mC. Performing TAPSp conversion, such as on a DNA sample as described herein, thus facilitates distinguishing positions containing unmodified C or hmC on the one hand from positions containing mC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424- 429. 5-hydroxymethylcytosine carbamoyltransferase is described in Yang et al., Bio-protocol, 2023; 12(17): e4496.
[0216] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises chemical-assisted conversion with a substituted borane reducing agent, optionally wherein the substituted borane reducing agent is 2-picoline borane, borane pyridine, tert-butylamine borane, or ammonia borane. In chemical-assisted conversion with a substituted borane reducing agent, an oxidizing agent such as potassium perruthenate (KRuO4) (also suitable for use in ox-BS conversion) is used to specifically oxidize hmC to fC. Treatment with pic-borane or another substituted borane reducing agent such as borane pyridine, tert-butylamine borane, or ammonia borane converts fC and caC to DHU but does not affect mC or unmodified C. Thus, when this type of conversion is used, the first nucleobase comprises one or more of hmC, fC, and caC, and the second nucleobase comprises one or more of unmodified cytosine or mC, such as unmodified cytosine and optionally mC. Sequencing of the converted DNA identifies positions that are read as cytosine as being either mC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T, fC, caC, or hmC. Performing this type of conversion, such as on a DNA sample as described herein, thus facilitatesdistinguishing positions containing unmodified C or mC on the one hand from positions containing hmC using the sequence reads obtained from the sample. For an exemplary description of this type of conversion, see, e.g., Liu et al., Nature Biotechnology 2019; 37:424-429.
[0217] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises APOBEC-coupled epigenetic (ACE) conversion. In ACE conversion, an AID / APOBEC family DNA deaminase enzyme such as APOBEC3A (A3 A) is used to deaminate unmodified cytosine and mC without deaminating hmC, fC, or caC. Thus, when ACE conversion is used, the first nucleobase comprises unmodified C and / or mC (e.g., unmodified C and optionally mC), and the second nucleobase comprises hmC. Sequencing of ACE-converted DNA identifies positions that are read as cytosine as being hmC, fC, or caC positions. Meanwhile, positions that are read as T are identified as being T, unmodified C, or mC. Performing ACE conversion on a DNA sample as described herein thus facilitates distinguishing positions containing hmC from positions containing mC or unmodified C using the sequence reads obtained from the sample. For an exemplary description of ACE conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0218] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase, e.g., as in EM-Seq. See, e.g., Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692vl. For example, TET2 and T4- GT or 5-hydroxymethylcytosine carbamoyltransferase (described in Yang et al., Bio-protocol, 2023; 12(17): e4496) can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by a deaminase (e.g., APOBEC3A), and then a deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosines converting them to uracils.
[0219] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using a non-specific, modification-sensitive double-stranded DNA deaminase, e.g., as in SEM- seq. See, e.g., Vaisvila et al. (2023) Discovery of novel DNA cytosine deaminase activities enables a nondestructive single-enzyme methylation sequencing method for base resolution high-coverage methylome mapping of cell-free and ultra-low input DNA. bioRxiv; DOI: 10.1101 / 2023.06.29.547047, available at https: / / www.biorxiv.org / content / 10.1101 / 2023.06.29.547047vl. SEM-Seq employs a non-specific, modification-sensitive double-stranded DNA deaminase (MsddA) in a nondestructive single-enzyme 5 -methyl ctyosine sequencing (SEM-seq) method that deaminates unmodified cytosines. Accordingly, SEM-seq does not require the TET2 and T4-PGT or 5- hydroxymethylcytosine carbamoyltransferase protection and denaturing steps that are of use, e.g., in APOEC3A-based protocols. Additionally, MsddA does not deaminate 5-formylated cytosines (5fC) or 5-carboxylated cytosines (5caC). In SEM-seq, unmodified cytosines in the DNA are deaminated to uracil and is read as “T” during sequencing. Modified cytosines (e.g., 5mC) are not converted and are read as “C” during sequencing. Cytosines that are read as thymines are identified as unmodified (e.g., unmethylated) cytosines or as thymines in the DNA. Performing SEM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained. In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises enzymatic conversion of the first nucleobase using MsddA.
[0220] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample converts a modified nucleoside. In some embodiments, the conversion procedure which converts a modified nucleosides comprises enzymatic conversion, such as DM-seq, for example, as described in WO2023 / 288222A1. In DM- seq, unmodified cytosines in the DNA are enzymatically protected from a subsequent deamination step wherein 5mC in 5mCpG is converted to T. The enzymatically protected unmodified (e.g., unmethylated) cytosines are not converted and are read as “C” during sequencing. Cytosines that are read as thymines (in a CpG context) are identified as methylated cytosines in the DNA. Thus, when this type of conversion is used, the first nucleobase comprises unmodified (such as unmethylated) cytosine, and the second nucleobase comprises modified (such as methylated) cytosine. Sequencing of the converted DNA identifies positions that are read as cytosine as being unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion thus facilitates identifying positions containing 5mC using the sequence reads obtained.
[0221] Exemplary cytosine deaminases for use herein include APOBEC enzymes, for example, APOBEC3A. Generally, AID / APOBEC family DNA deaminase enzymes such as APOBEC3A (A3A) are used to deaminate (unprotected) unmodified cytosine and 5mC. For an exemplary description of APOBEC conversion, see, e.g., Schutsky et al., Nature Biotechnology 2018; 36: 1083-1090.
[0222] The enzymatic protection of unmodified cytosines in the DNA comprises addition of a protective group to the unmodified cytosines. Such protective groups can comprise an alkyl group, an alkyne group, a carboxyl group, a carboxyalkyl group, an amino group, a hydroxymethyl group, a glucosyl group, a glucosylhydroxymethyl group, an isopropyl group, or a dye. For example, DNA can be treated with a methyltransferase, such as a CpG-specific methyltransferase, which adds the protective group to unmodified cytosines. The term methyltransferase is used broadly herein to refer to enzymes capable of transferring a methyl or substituted methyl (e.g., carboxymethyl) to a substrate (e.g., a cytosine in a nucleic acid). In some embodiments, the DNA is contacted with a CpG-specific DNA methyltransferase (MTase), such as a CpG-specific carboxymethyltransferase (CxMTase), and a substituted methyl donor, such as a carboxymethyl donor (e.g., carboxymethyl-S-adenosyl-L-methionine). See, e.g., WO2021 / 236778A2. In particular embodiments, the CxMTase can facilitate the addition of a protective carboxymethyl group to an unmethylated cytosine. In some embodiments, the unmethylated cytosine is unmodified cytosine. The carboxymethyl group can prevent deamination of the cytosine during a deamination step (such as a deamination step using an APOBEC enzyme, such as A3A). Substituted methyl or carboxymethyl donors useful in the disclosed methods include but are not limited to, S-adenosyl-L-methionine (SAM) analogs, optionally wherein the SAM analog is carboxy-S-adenosyl-L-methionine (CxSAM). SAM analogs are described, for example, in WO2022 / 197593A1. The MTase may be, for example, a CpG methyltransferase from Spiroplasma sp. strain MQ1 (M.SssI), DNA-methyltransferase 1 (DNMT1), DNA-methyltransferase 3 alpha (DNMT3A), DNA-methyltransferase 3 beta (DNMT3B), or DNA adenine methyltransferase (Dam). The CxMTase may be a CpG methyltransferase from Mycoplasma penetrans (M.Mpel). In a particular embodiment, the methyltransferase enzyme is a variant of M.Mpel having SEQ ID NO: 1 or SEQ ID NO: 2, or a sequence at least 90%, at least 92%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% identical thereto, optionally wherein the amino acid corresponding to position 374 is R or K.
[0223] In one embodiment, the methyltransferase enzyme is a variant of M.Mpel having an N374R substitution or an N374K substitution. The methyltransferase of SEQ ID NO: 1 or SEQ ID NO: 2 can further comprise one or more amino acid substitutions selected from a) substitution of one or both residues T300 and E305 with S, A, G, Q, D, or N; b) substitution of one or more residues A323, N306, and Y299 with a positively charged amino acid selected from K, R or H; and / or c) substitution of S323 with A, G, K, R or H, which may enhance the activity of the enzyme.
[0224] Optionally, the conversion procedure further includes enzymatic protection of 5hmCs, such as by glucosylation of the 5hmCs (e.g., using PGT) or by carbamoylation of the 5hmCs (e.g., using 5-hydroxymethylcytosine carbamoyltransferase), in the DNA prior to the deamination of unprotected modified cytosines. In this method, 5hmC can be protected from conversion, for example through glucosylation using P-glucosyl transferase (PGT), forming (5- glucosylhydroxymethyl cytosine) 5ghmC, or through carbamoylation using 5- hydroxymethylcytosine carbamoyltransferase, forming 5cmC. This is described, for example, in Yu et al., Cell 2012; 149: 1368-80, and in Yang et al., Bio-protocol, 2023; 12(17): e4496. Glucosylation or carbamoylation of 5hmC can reduce or eliminate deamination of 5hmC by a deaminase such as APOBEC3A. Treatment with an MTase or CxMTase then adds a protecting group to unmodified (unmethylated) cytosines in the DNA. 5mC (but not protected, unmodified cytosine and not 5ghmC or 5cmC) is then deaminated (converted to T in the case of 5mC) by treatment with a deaminase, for example, an APOBEC enzyme (such as APOBEC3 A). Sequencing of the converted DNA identifies positions that are read as cytosine as being either 5hmC or unmodified C positions. Meanwhile, positions that are read as T are identified as being T or 5mC. Performing DM-seq conversion with glucosylation of 5hmC on a sample as described herein thus facilitates distinguishing positions containing unmodified C or 5hmC on the one hand from positions containing 5mC using the sequence reads obtained.
[0225] Also provided herein are methods in which alternative base conversion schemes are used. For example, unmethylated cytosines can be left intact while methylated cytosines and hydroxymethylcytosines are converted to a base read as a thymine (e.g., uracil, thymine, or dihydrouracil).
[0226] In some embodiments, methylating a cytosine in at least one first complementary strand or second complementary strand comprises contacting the cytosine with a methyltransferase such as DNMT1 or DNMT5. In such embodiments, the step of oxidizing a 5-hydroxymethylated cytosine to 5-formylcytosine (such as by contacting the 5 -hydroxymethyl cytosine in a first strand and a second strand with KRuO4) can be optional.
[0227] In some embodiments, converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine comprises oxidizing a hydroxymethyl cytosine, e.g., the hydroxymethyl cytosine is oxidized to formylcytosine. In some embodiments, oxidizing the hydroxymethyl cytosine to formylcytosine comprises contacting the hydroxymethyl cytosine with a ruthenate, such as potassium ruthenate (KRuO4).
[0228] In some embodiments, the modified cytosine is converted to thymine, uracil, or dihydrouracil. In any such embodiments, amplification methods may comprise uracil- and / or dihydrouracil-tolerant amplification methods, such as PCR using a uracil- and / or dihydrouracil - tolerant DNA polymerase.
[0229] In some embodiments, the method comprises converting a formylcytosine and / or a methylcytosine to carboxylcytosine as part of converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine. For example, converting the formylcytosine and / or the methylcytosine to carboxylcytosine can comprise contacting the formylcytosine and / or the methylcytosine with a TET enzyme, such as TET1, TET2, TET3, or a TET2 comprising a T1372S mutation. In some embodiments, the method comprises reducing the carboxylcytosine as part of converting the modified cytosine in at least one first or second strand to a thymine or a base read as thymine, and / or the carboxylcytosine is reduced to dihydrouracil. In some embodiments, reducing the carboxylcytosine comprises contacting the carboxylcytosine with a borane or borohydride reducing agent.
[0230] In some embodiments, the borane or borohydride reducing agent comprises pyridine borane, 2-picoline borane, borane, tert-butylamine borane, ammonia borane, sodium borohydride, sodium cyanoborohydride (NaBH3CN), lithium borohydride (LiBH4), ethylenediamine borane, dimethylamine borane, sodium triacetoxyborohydride, morpholine borane, 4-methylmorpholine borane, trimethylamine borane, dicyclohexylamine borane, or a salt thereof. In other embodiments, the reducing agent comprises lithium aluminum hydride, sodium amalgam, amalgam, sulfur dioxide, dithionate, thiosulfate, iodide, hydrogen peroxide, hydrazine, diisobutylaluminum hydride, oxalic acid, carbon monoxide, cyanide, ascorbic acid, formic acid, dithiothreitol, betamercaptoethanol, or any combination thereof.
[0231] Various TET enzymes may be used in the disclosed methods as appropriate. In some embodiments, the one or more TET enzymes comprise TETv. TETv is described in US Patent 10,260,088 and its sequence is SEQ ID NO: 1 therein (SEQ ID NO: 3 in the present application). In some embodiments, the one or more TET enzymes comprise TETcd. TETcd is described in US Patent 10,260,088 and its sequence is SEQ ID NO: 3 therein (SEQ ID NO: 4 in the present application). In some embodiments, the one or more TET enzymes comprise TET1. In some embodiments, the one or more TET enzymes comprise TET2. TET2 may be expressed and used as a fragment comprising TET2 residues 1129-1480 joined to TET2 residues 1844-1936 by a linker (SEQ ID NO: 5 of the present application) as described, e.g., in US Patent 10,961,525. In someembodiments, the one or more TET enzymes comprise TET1 and TET2. In some embodiments, the one or more TET enzymes comprise a VI 900 TET mutant, such as a V1900A, V1900C, V1900G, VI 9001, or V1900P TET mutant. In some embodiments, the one or more TET enzymes comprise a VI 900 TET2 mutant, such as a VI 900 A, V1900C, V1900G, VI 9001, or V1900P TET2 mutant. Examples of V1900A, V1900C, V1900G, V1900I, and V1900P TET2 mutants are provided as SEQ ID NOs: 6-10. In some embodiments, the V1900 TET mutant has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 6, 7, 8, 9, or 10. Position 1900 of the wild-type TET2 sequence corresponds to position 438 in each of SEQ ID NOs: 5-10. It can be beneficial to use a TET enzyme that maximizes formation of 5- carboxylcytosine (5-caC) relative to less oxidized modified cytosines, particularly 5- formylcytosine, because 5-caC is not a substrate for enzymatic deamination, e.g., by APOBEC enzymes such as APOBEC3A. Maximizing formation of 5-caC thus reduces the risk of false calls in which a base is identified as unmethylated because it underwent deamination even though it was methylated (or hydroxymethyl ated) in the original sample. Accordingly, in some embodiments, the TET enzyme comprises a mutation that increases formation of 5-caC. Exemplary mutations are set forth above. “A mutation that increases formation of 5-caC” means that the TET enzyme having the mutation produces more 5-caC than a TET enzyme that lacks the mutation but is otherwise identical. 5-caC production can be measured as described, e.g., in Liu et al., Nat Chem Biol 13: 181-187 (2017) (see Online Methods section, TET reactions in vitro subsection, “driving” conditions). Any variants and / or mutants described in Liu et al. (2017) can be used in the disclosed methods as appropriate.
[0232] In some embodiments, the one or more TET enzymes comprise a TET2 enzyme comprising a T1372S mutation, such as TET2-CS-T1372S and TET2-CD-T1372S. Examples of TET2-CS- T1372S and TET2-CD-T1372S are provided as SEQ ID NOs: 11 and 12. A TET2 comprising a T1372S mutation is described in US Patent 10,961,525 and may be expressed and used as a fragment comprising TET2 residues 1129-1480 joined to TET2 residues 1844-1936 by a linker. Position 1372 of TET2 corresponds to position 258 of SEQ ID NO: 21 (wild type TET2 catalytic domain) of US Patent 10,961,525. Thus, the sequence of a T1372S TET2 catalytic domain may be obtained by changing the threonine at position 258 of SEQ ID NO: 21 of US Patent 10,961,525 to serine. TET2 comprising a T1372S mutation is also described in Liu et al., Nat Chem Biol. 2017 February; 13(2): 181-187. As demonstrated in Liu et al., TET2 comprising a T1372S mutation can more efficiently oxidize 5mC to produce 5-carboxylcytosine (5caC) than other versions ofTET2 such as TET2 lacking a T1372S mutation. In some embodiments, the TET2 enzyme comprises SEQ ID NO: 14 or optionally a variant of SEQ ID NO: 14 in which at least 5, 6, 7, or 8 positions match SEQ ID NO: 14 including position 5 of SEQ ID NO: 14. In some embodiments, the TET2 enzyme is a human TET2 enzyme comprising a T1372S mutation. In some embodiments, the TET2 enzyme comprises the sequence of SEQ ID NO: 11. In some embodiments, the TET2 enzyme comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 11. In some embodiments, the TET2 enzyme comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 12. In some embodiments, the TET2 enzyme comprises the sequence of SEQ ID NO: 12. The sequences of SEQ ID NOs: 11 and 12 are shown in the Table of Sequences herein.
[0233] Provided herein is a method comprising contacting DNA contacting DNA with a TET2 enzyme comprising a T1372S mutation to oxidize 5-methylcytosine (5mC) and / or 5- hydroxymethylcytosine (5hmC) present in the DNA to 5-carboxycytosine (5caC), subsequently contacting at least a portion of the DNA with a substituted borane reducing agent, thereby converting 5-caC in the DNA to dihydrouracil (DHU), thereby producing treated DNA, and sequencing at least a portion of the treated DNA.
[0234] In some embodiments, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA comprises separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase. In some such embodiments, the first nucleobase is hmC. DNA originally comprising the first nucleobase may be separated from other DNA using a labeling procedure comprising biotinylating positions that originally comprised the first nucleobase. In some embodiments, the first nucleobase is first derivatized with an azide-containing moiety, such as a glucosyl-azide containing moiety. The azide-containing moiety then may serve as a reagent for attaching biotin, e.g., through Huisgen cycloaddition chemistry. Then, the DNA originally comprising the first nucleobase, now biotinylated, can be separated from DNA not originally comprising the first nucleobase using a biotin-binding agent, such as avidin, neutravidin (deglycosylated avidin with an isoelectric point of about 6.3), or streptavidin. An example of a procedure for separating DNA originally comprising the first nucleobase from DNA not originally comprising the first nucleobase is hmC-seal, which labels hmC to form P-6-azide-glucosyl-5-hydroxymethylcytosine and then attaches a biotin moiety through Huisgen cycloaddition, followed by separation of the biotinylated DNA from other DNA using a biotin-binding agent. For an exemplary description of hmC-seal, see, e.g., Han et al., Mol.Cell 2016; 63: 711-719. This approach is useful for identifying fragments that include one or more hmC nucleobases.
[0235] In some embodiments, following such a separation, the method further comprises differentially tagging each of the DNA originally comprising the first nucleobase, the DNA not originally comprising the first nucleobase. The method may further comprise pooling the DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase following differential tagging. The DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase may then be used in downstream analyses. For example, the pooled DNA originally comprising the first nucleobase and the DNA not originally comprising the first nucleobase may be sequenced in the same sequencing cell (such as after being subjected to further treatments, such as those described herein) while retaining the ability to resolve whether a given read came from a molecule of DNA originally comprising the first nucleobase or DNA not originally comprising the first nucleobase using the differential tags.
[0236] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).
[0237] Techniques comprising partitioning based on methylation status or methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases such as mC, mA, caC (which may be generated by oxidation of mC or hmC with Tet2, e.g., before enzymatic conversion of unmodified C to U, e.g., using a deaminase such as APOBEC3A), or dihydrouracil from other DNA. See, e.g., Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. An antibody specific for mA is described in Sun et al., Bioessays 2015; 37: 1155-62. Antibodies for various modified nucleobases, such as mC, caC, and forms of thymine / uracil including dihydrouracil or halogenated forms such as 5-bromouracil, are commercially available. Various modified bases can also be detected based on alterations in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read in sequencing as a G. See, e.g., US Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, N.Y., 2002, chapter 14, “Mutation, Repair, and Recombination.”Enriching / Capturing step; amplification; adaptors; barcodes
[0238] In some embodiments, methods disclosed herein comprise a step of capturing one or more sets of target regions of DNA, such as cfDNA. Capture may be performed using any suitable approach known in the art.
[0239] In some embodiments, capturing comprises contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes may have any of the features described herein for sets of target-specific probes, including but not limited to in the embodiments set forth above and the sections relating to probes below. Capturing may be performed on one or more subsamples prepared during methods disclosed herein. In some embodiments, DNA is captured from at least the first subsample or the second subsample, e.g., at least the first subsample and the second subsample. Where the first subsample undergoes a separation step (e.g., separating DNA originally comprising the first nucleobase (e.g., hmC) from DNA not originally comprising the first nucleobase, such as hmC-seal), capturing may be performed on any, any two, or all of the DNA originally comprising the first nucleobase (e.g., hmC), the DNA not originally comprising the first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled before undergoing capture.
[0240] The capturing step may be performed using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on features of the probes such as length, base composition, etc. Those skilled in the art will be familiar with appropriate conditions given general knowledge in the art regarding nucleic acid hybridization. In some embodiments, complexes of target-specific probes and DNA are formed.
[0241] In some embodiments, a method described herein comprises capturing cfDNA obtained from a test subject for a plurality of sets of target regions. The target regions comprise epigenetic target regions, which may show differences in methylation levels and / or fragmentation patterns depending on whether they originated from a tumor or from healthy cells. The target regions also comprise sequence-variable target regions, which may show differences in sequence depending on whether they originated from a tumor or from healthy cells. The capturing step produces a captured set of cfDNA molecules, and the cfDNA molecules corresponding to the sequence-variable target region set are captured at a greater capture yield in the captured set of cfDNA molecules than cfDNA molecules corresponding to the epigenetic target region set. For additional discussion of capturing steps, capture yields, and related aspects, see W02020 / 160414, which is incorporated herein by reference for all purposes.
[0242] In some embodiments, a method described herein comprises contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set.
[0243] It can be beneficial to capture cfDNA corresponding to the sequence-variable target region set at a greater capture yield than cfDNA corresponding to the epigenetic target region set because a greater depth of sequencing may be necessary to analyze the sequence-variable target regions with sufficient confidence or accuracy than may be necessary to analyze the epigenetic target regions. The volume of data needed to determine fragmentation patterns (e.g., to test fsor perturbation of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated partitions) is generally less than the volume of data needed to determine the presence or absence of cancer-related sequence mutations. Capturing the target region sets at different yields can facilitate sequencing the target regions to different depths of sequencing in the same sequencing run (e.g., using a pooled mixture and / or in the same sequencing cell).
[0244] In various embodiments, the methods further comprise sequencing the captured cfDNA, e.g., to different degrees of sequencing depth for the epigenetic and sequence-variable target region sets, consistent with the discussion herein.
[0245] In some embodiments, complexes of target-specific probes and DNA are separated from DNA not bound to target-specific probes. For example, where target-specific probes are bound covalently or noncovalently to a solid support, a washing or aspiration step can be used to separate unbound material. Alternatively, where the complexes have chromatographic properties distinct from unbound material (e.g., where the probes comprise a ligand that binds a chromatographic resin), chromatography can be used.
[0246] As discussed in detail elsewhere herein, the set of target-specific probes may comprise a plurality of sets such as probes for a sequence-variable target region set and probes for an epigenetic target region set. In some such embodiments, the capturing step is performed with the probes for the sequence-variable target region set and the probes for the epigenetic target region set in the same vessel at the same time, e.g., the probes for the sequence-variable and epigenetic target region sets are in the same composition. This approach provides a relatively streamlined workflow. In some embodiments, the concentration of the probes for the sequence-variable target region set is greater that the concentration of the probes for the epigenetic target region set.
[0247] Alternatively, the capturing step is performed with the sequence-variable target region probe set in a first vessel and with the epigenetic target region probe set in a second vessel, or the contacting step is performed with the sequence-variable target region probe set at a first time and a first vessel and the epigenetic target region probe set at a second time before or after the first time. This approach allows for preparation of separate first and second compositions comprising captured DNA corresponding to the sequence-variable target region set and captured DNA corresponding to the epigenetic target region set. The compositions can be processed separately as desired (e.g., to fractionate based on methylation as described elsewhere herein) and recombined in appropriate proportions to provide material for further processing and analysis such as sequencing.
[0248] In some embodiments, the DNA is amplified. In some embodiments, amplification is performed before the capturing step. In some embodiments, amplification is performed after the capturing step.
[0249] In some embodiments, adapters are included in the DNA. This may be done concurrently with an amplification procedure, e.g., by providing the adapters in a 5’ portion of a primer, e.g., as described above. Alternatively, adapters can be added by other approaches, such as ligation.
[0250] In some embodiments, tags, which may be or include barcodes, are included in the DNA. Tags can facilitate identification of the origin of a nucleic acid. For example, barcodes can be used to allow the origin (e.g., subject) whence the DNA came to be identified following pooling of a plurality of samples for parallel sequencing. This may be done concurrently with an amplification procedure, e.g., by providing the barcodes in a 5’ portion of a primer, e.g., as described above. In some embodiments, adapters and tags / barcodes are provided by the same primer or primer set. For example, the barcode may be located 3’ of the adapter and 5’ of the target-hybridizing portion of the primer. Alternatively, barcodes can be added by other approaches, such as ligation, optionally together with adapters in the same ligation substrate.
[0251] Additional details regarding amplification, tags, and barcodes are discussed in the “General Features of the Methods” section below, which can be combined to the extent practicable with any of the foregoing embodiments and the embodiments set forth in the introduction and summary section.Captured set
[0252] In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. With respect to the disclosed methods, the captured set of DNA may be provided, e g., by performing a capturing step after a partitioning step as described herein. The captured set may comprise DNA corresponding to a sequence-variable target region set, an epigenetic target region set, or a combination thereof. In some embodiments the quantity of captured sequence-variable target region DNA is greater than the quantity of the captured epigenetic target region DNA, when normalized for the difference in the size of the targeted regions (footprint size).
[0253] Alternatively, first and second captured sets may be provided, comprising, respectively, DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set. The first and second captured sets may be combined to provide a combined captured set.In some embodiments in which a captured set comprising DNA corresponding to the sequencevariable target region set and the epigenetic target region set includes a combined captured set as discussed above, the DNA corresponding to the sequence-variable target region set may be present at a greater concentration than the DNA corresponding to the epigenetic target region set, e.g., a 1.1 to 1.2-fold greater concentration, a 1.2- to 1.4-fold greater concentration, a 1.4- to 1.6-fold greater concentration, a 1.6- to 1.8-fold greater concentration, a 1.8- to 2.0-fold greater concentration, a 2.0- to 2.2-fold greater concentration, a 2.2- to 2.4-fold greater concentration a2.4- to 2.6-fold greater concentration, a 2.6- to 2.8-fold greater concentration, a 2.8- to 3.0-fold greater concentration, a 3.0- to 3.5-fold greater concentration, a 3.5- to 4.0, a 4.0- to 4.5-fold greater concentration, a 4.5- to 5.0-fold greater concentration, a 5.0- to 5.5-fold greater concentration, a 5.5- to 6.0-fold greater concentration, a 6.0- to 6.5-fold greater concentration, a6.5- to 7.0-fold greater, a 7.0- to 7.5-fold greater concentration, a 7.5- to 8.0-fold greater concentration, an 8.0- to 8.5-fold greater concentration, an 8.5- to 9.0-fold greater concentration, a 9.0- to 9.5-fold greater concentration, 9.5- to 10.0-fold greater concentration, a 10- to 11-fold greater concentration, an 11- to 12-fold greater concentration a 12- to 13 -fold greater concentration, a 13- to 14-fold greater concentration, a 14- to 15-fold greater concentration, a 15- to 16-fold greater concentration, a 16- to 17-fold greater concentration, a 17- to 18-fold greater concentration, an 18- to 19-fold greater concentration, a 19- to 20-fold greater concentration, a 20- to 30-fold greater concentration, a 30- to 40-fold greater concentration, a 40- to 50-fold greater concentration, a 50- to 60-fold greater concentration, a 60- to 70-fold greater concentration, a 70-to 80-fold greater concentration, a 80- to 90-fold greater concentration, a 90- to 100-fold greater concentration, a 10- to 20-fold greater concentration, a 10- to 40-fold greater concentration, a 10- to 50-fold greater concentration, a 10- to 70-fold greater concentration, or a 10- to 100-fold greater concentration. The degree of difference in concentrations accounts for normalization for the footprint sizes of the target regions, as discussed in the definition section.Epigenetic target region set
[0254] The epigenetic target region set may comprise one or more types of target regions likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells and from healthy cells, e.g., non- neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. The epigenetic target region set may also comprise one or more control regions, e.g., as described herein. In additional embodiments the epigenetic target region set may comprise different forms of nucleic acids may include transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.
[0255] In some embodiments, the epigenetic target region set has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has a footprint in the range of 100-1000 kb, e.g., 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1,000 kb.Hypermethylation variable target regions
[0256] In some embodiments, the epigenetic target region set comprises one or more hypermethylation variable target regions. In general, hypermethylation variable target regions refer to regions where an increase in the level of observed methylation, e.g., in a cfDNA sample, indicates an increased likelihood that a sample (e.g., of cfDNA) contains DNA produced by neoplastic cells, such as tumor or cancer cells. For example, hypermethylation of promoters of tumor suppressor genes has been observed repeatedly. See, e.g., Kang et al., Genome Biol. 18:53 (2017) and references cited therein. In an example, hypermethylation variable target regions can include regions that do not necessarily differ in methylation in cancerous tissue relative to DNAfrom healthy tissue of the same type, but do differ in methylation (e.g., have more methylation) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such hypermethylation variable target regions. In some embodiments, hypermethylation variable target regions include one or more genomic regions, where the cfDNA molecules in those regions do not differ in methylation state in cancer subjects relative to cfDNA from healthy subjects, but the presence / increased quantity of hypermethylated cfDNA in those regions is indicative of a particular tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g. tumor shedding) into circulation.
[0257] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866: 106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4 and NDRG4. An exemplary set of hypermethylation variable target regions based on colorectal cancer (CRC) studies is provided in Table 2. Many of these genes likely have relevance to cancers beyond colorectal cancer; for example, TP53 is widely recognized as a critically important tumor suppressor and hypermethylation-based inactivation of this gene may be a common oncogenic mechanism.Table 2. Exemplary Hypermethylation Target Regions based on CRC studies.
[0258] In some embodiments, the hypermethylation variable target regions comprise a plurality of loci listed in Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. For example, for each locus included as a target region, there may be one or more probes with a hybridization site that binds between the transcription start site and the stop codon (the last stop codon for genes that are alternatively spliced) of the gene, or in the promoter region of the gene. In some embodiments, the one or more probes bind within 300 bp of the transcription start site of a gene in Table 2, e.g., within 200 or 100 bp.
[0259] Methylation variable target regions in various types of lung cancer are discussed in detail, e.g., in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453- 74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017); Shi et al., BMC Genomics 18:901 (2017); Schneider et al., BMC Cancer. 11 : 102 (2011); Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016); Skvortsova et al., Br. J. Cancer. 94(10): 1492-1495 (2006); Kim et al., Cancer Res. 61 :3419-3424 (2001); Furonaka et al., Pathology International 55:303-309 (2005); Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014); Kim et al., Oncogene. 20: 1765-70 (2001); Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003); Kikuchi et al., Clin. Cancer Res. 11 :2954-61 (2005); Heller et al., Oncogene 25:959-968 (2006); Licchesi et al., Carcinogenesis. 29:895-904 (2008); Guo et al., Clin. Cancer Res. 10:7917-24 (2004); Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61 :4556-4560, (2001).
[0260] An exemplary set of hypermethylation variable target regions based on lung cancer studies is provided in Table 3. Many of these genes likely have relevance to cancers beyond lung cancer; for example, Casp8 (Caspase 8) is a key enzyme in programmed cell death and hypermethylationbased inactivation of this gene may be a common oncogenic mechanism not limited to lung cancer. Additionally, a number of genes appear in both Tables 1 and 2, indicating generality.Table 3. Exemplary Hypermethylation Target Regions based on Lung Cancer studies
[0261] Any of the foregoing embodiments concerning target regions identified in Table 3 may be combined with any of the embodiments described above concerning target regions identified in Table 2. In some embodiments, the hypermethylation variable target regions comprise a plurality of loci listed in Table 2 or Table 3, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2 or Table 3.
[0262] Additional hypermethylation target regions may be obtained, e.g., from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017), describe construction of a probabilisticmethod called CancerLocator using hypermethylation target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylation target regions can be specific to one or more types of cancer. Accordingly, in some embodiments, the hypermethylation target regions include one, two, three, four, or five subsets of hypermethylation target regions that collectively show hypermethylation in one, two, three, four, or five of breast, colon, kidney, liver, and lung cancers.
[0263] In yet additional embodiments the, epigenetic target region set comprises transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated or contain a specific histone modification or functional element such as a TFBS.Hypomethylation variable target regions
[0264] Global hypomethylation is a commonly observed phenomenon in various cancers. See, e.g., Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1 :239-259 (2009) (review article noting observations of hypomethylation in colon, ovarian, prostate, leukemia, hepatocellular, and cervical cancers). For example, regions such as repeated elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and intergenic regions that are ordinarily methylated in healthy cells may show reduced methylation in tumor cells. Accordingly, in some embodiments, the epigenetic target region set includes hypomethylation variable target regions, where a decrease in the level of observed methylation indicates an increased likelihood that a sample (e.g., of cfDNA) contains DNA produced by neoplastic cells, such as tumor or cancer cells. In an example, hypomethylation variable target regions can include regions that do not necessarily differ in methylation state in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ in methylation (e.g., are less methylated) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such hypomethylation variable target regions. In some embodiments, hypomethylation variable target regions include one or more genomic regions, where the cfDNA molecules in those regions do notdiffer in methylation state in cancer subjects relative to cfDNA from healthy subjects, but the presence / increased quantity of hypom ethylated cfDNA in those regions is indicative of a particular tissue type (e.g., cancer origin) and is presented as cfDNA with increased apoptosis (e.g. tumor shedding) into circulation.
[0265] In some embodiments, hypomethylation variable target regions include repeated elements and / or intergenic regions. In some embodiments, repeated elements include one, two, three, four, or five of LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.
[0266] Exemplary specific genomic regions that show cancer-associated hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1. In some embodiments, the hypomethylation variable target regions overlap or comprise one or both of these regions.CTCF binding sites
[0267] CTCF is a DNA-binding protein that contributes to chromatin organization and often colocalizes with cohesin. Perturbation of CTCF binding sites has been reported in a variety of different cancers. See, e.g., Katainen et al., Nature Genetics, doi: 10.1038 / ng.3335, published online 8 June 2015; Guo et al., Nat. Commun. 9: 1520 (2018). CTCF binding results in recognizable patterns in cfDNA that can be detected by sequencing, e.g., through fragment length analysis. Details regarding sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016); WO 2018 / 009723; and US20170211143A1, each of which are incorporated herein by reference.
[0268] Thus, perturbations of CTCF binding result in variation in the fragmentation patterns of cfDNA. As such, CTCF binding sites represent a type of fragmentation variable target regions.
[0269] There are many known CTCF binding sites. See, e.g., the CTCFBSDB (CTCF Binding Site Database), available on the Internet at insulatordb.uthsc.edu / ; Cuddapah et al., Genome Res. 19:24-32 (2009); Martin et al., Nat. Struct. Mol. Biol. 18:708-14 (2011); Rhee et al., Cell. 147: 1408-19 (2011), each of which are incorporated by reference. Exemplary CTCF binding sites are at nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13.
[0270] Accordingly, in some embodiments, the epigenetic target region set includes CTCF binding regions. In some embodiments, the CTCF binding regions comprise at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCFbinding regions, e.g., such as CTCF binding regions described above or in one or more of CTCFBSDB or the Cuddapah et al., Martin et al., or Rhee et al. articles cited above.
[0271] In some embodiments, at least some of the CTCF sites can be methylated or unmethylated, wherein the methylation state is correlated with the whether or not the cell is a cancer cell. In some embodiments, the epigenetic target region set comprises at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream regions of the CTCF binding sites.Transcriptional start sites
[0272] Transcriptional start sites may also show perturbations in neoplastic cells. For example, nucleosome organization at various transcription start sites in healthy cells of the hematopoietic lineage — which contributes substantially to cfDNA in healthy individuals — may differ from nucleosome organization at those transcription start sites in neoplastic cells. This results in different cfDNA patterns that can be detected by sequencing, as discussed generally in Snyder et al., Cell 164:57-68 (2016); WO 2018 / 009723; and US20170211143A1. In another example, transcription start sites that do not necessarily differ epigenetically in cancerous tissue relative to DNA from healthy tissue of the same type, but do differ epigenetically (e.g., with respect to nucleosome organization) relative to cfDNA that is typical in healthy subjects. Where, for example, the presence of a cancer results in increased cell death such as apoptosis of cells of the tissue type corresponding to the cancer, such a cancer can be detected at least in part using such transcription start sites.
[0273] Thus, perturbations of transcription start sites also result in variation in the fragmentation patterns of cfDNA. As such, transcription start sites also represent a type of fragmentation variable target regions.
[0274] Human transcriptional start sites are available from DBTSS (DataBase of Human Transcription Start Sites), available on the Internet at dbtss.hgc.jp and described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), which is incorporated herein by reference.
[0275] Accordingly, in some embodiments, the epigenetic target region set includes transcriptional start sites. In some embodiments, the transcriptional start sites comprise at least 10, 20, 50, 100, 200, or 500 transcriptional start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcriptional start sites, e.g., such as transcriptional start sites listed in DBTSS. In some embodiments, at least some of the transcription start sites can be methylated or unmethylated,wherein the methylation state is correlated with whether or not the cell is a cancer cell. In some embodiments, the epigenetic target region set comprises at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, at least 1000 bp upstream and downstream regions of the transcription start sites.Focal amplifications
[0276] Although focal amplifications are somatic mutations, they can be detected by sequencing based on read frequency in a manner analogous to approaches for detecting certain epigenetic changes such as changes in methylation. As such, regions that may show focal amplifications in cancer can be included in the epigenetic target region set and may comprise one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAFI. For example, in some embodiments, the epigenetic target region set comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the foregoing targets.Methylation control regions
[0277] It can be useful to include control regions to facilitate data validation. In some embodiments, the epigenetic target region set includes control regions that are expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is derived from a cancer cell or a normal cell. In some embodiments, the epigenetic target region set includes control hypomethylated regions that are expected to be hypomethylated in essentially all samples. In some embodiments, the epigenetic target region set includes control hypermethylated regions that are expected to be hypermethylated in essentially all samples.Sequence-variable target region set
[0278] In some embodiments, the sequence-variable target region set comprises a plurality of regions known to undergo somatic mutations in cancer.
[0279] In some aspects, the sequence-variable target region set targets a plurality of different genes or genomic regions (“panel”) selected such that a determined proportion of subjects having a cancer exhibits a genetic variant or tumor marker in one or more different genes or genomic regions in the panel. The panel may be selected to limit a region for sequencing to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA, e.g., by adjusting the affinity and / or amount of the probes as described elsewhere herein. The panel may be further selected to achieve a desired sequence read depth. The panel may be selected to achieve a desiredsequence read depth or sequence read coverage for an amount of sequenced base pairs. The panel may be selected to achieve a theoretical sensitivity, a theoretical specificity, and / or a theoretical accuracy for detecting one or more genetic variants in a sample.
[0280] Probes for detecting the panel of regions can include those for detecting genomic regions of interest (hotspot regions) as well as nucleosome-aware probes (e.g., KRAS codons 12 and 13) and may be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation impacted by nucleosome binding patterns and GC sequence composition. Regions used herein can also include non-hotspot regions optimized based on nucleosome positions and GC models.
[0281] Examples of listings of genomic locations of interest may be found in Table 4 and Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprise at least a portion of at least 1, at least 2, or 3 of the indels of Table 4. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes of Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs of Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 5. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17,or 18 of the indels of Table 5. Each of these genomic locations of interest may be identified as a backbone region or hot-spot region for a given panel. An example of a listing of hot-spot genomic locations of interest may be found in Table 6. In some embodiments, a sequence-variable target region set used in the methods of the present disclosure comprises at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes of Table 6. Each hot-spot genomic region is listed with several characteristics, including the associated gene, chromosome on which it resides, the start and stop position of the genome representing the gene’s locus, the length of the gene’s locus in base pairs, the exons covered by the gene, and the critical feature (e.g., type of mutation) that a given genomic region of interest may seek to capture.Table 4Table 5Table 6
[0282] Additionally or alternatively, suitable target region sets are available from the literature. For example, Gale et al., PLoS One 13: e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of a sequence-variable target region set. These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESRI, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED 12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0283] In some embodiments, the sequence-variable target region set comprises target regions from at least 10, 20, 30, or 35 cancer-related genes, such as the cancer-related genes listed above.Subjects
[0284] In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject having a cancer. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject suspected of having a cancer. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject having a tumor. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA or genomic DNA from tissue) is obtained from a subject having neoplasia. In some embodiments, the DNA (e.g., cfDNA, or genomic DNA from tissue) is obtained from a subject suspected of having neoplasia. In some embodiments, the DNA (e.g., cfDNA, or genomic DNA from tissue) is obtained from a subject in remission from a tumor, cancer, or neoplasia (e.g., following chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the lung. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the breast. In someembodiments, the cancer, tumor, or neoplasia or suspected cancer, tumor, or neoplasia is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.Sequencing
[0285] In general, sample nucleic acids flanked by adapters with or without prior amplification can be subject to sequencing. Sequencing methods include, for example, Sanger sequencing, high- throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by- hybridization, Digital Gene Expression (Helicos), Next generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, and sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which may multiple lanes, multiple channels, multiple wells, or other mean of processing multiple sample sets substantially simultaneously. Sample processing unit can also include multiple sample chambers to enable processing of multiple runs simultaneously. Additionally, sequencing chemistries can include Illumina sequencing (MiSeq, HiSeq, NextSeq, NovaSeq, MiniSeq, iSeq 100), Oxford Nanopore sequencing (MinlON, GridlON, PromethlON, Flongle), Sanger sequencing (ABI 3730x1), Ion Torrent sequencing (Ion PGM, Ion S5, Ion GeneStudio S5), PacBio SMRT sequencing (Sequel, Sequel lie), Illumina NovaSeq 6000, PacBio HiFi sequencing (Sequel II and Sequel lie Systems using HiFi chemistry). Element Biosciences (AVITI System). Ultima Genomics (Ultima Sequencing), and Singular Genomics (G4 Sequencing System).
[0286] The sequencing reactions can be performed on one or more forms of nucleic acids at least one of which is known to contain markers of cancer or of other disease. The sequencing reactions can also be performed on any nucleic acid fragments present in the sample. In some embodiments, sequence coverage of the genome may be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%. In some embodiments, the sequence reactions may provide for sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, or 80% of the genome. Sequence coverage can performed on at least 5, 10, 20, 70, 100, 200 or 500 different genes, or at most 5000, 2500, 1000, 500 or 100 different genes.
[0287] Simultaneous sequencing reactions may be performed using multiplex sequencing. In some cases, cell-free nucleic acids may be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases cell-free nucleic acids may be sequenced with less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. Sequencing reactions may be performed sequentially or simultaneously. Subsequent data analysis may be performed on all or part of the sequencing reactions. In some cases, data analysis may be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, data analysis may be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. An exemplary read depth is 1000-50000 reads per locus (base).Differential depth of sequencing
[0288] In some embodiments, nucleic acids corresponding to the sequence-variable target region set are sequenced to a greater depth of sequencing than nucleic acids corresponding to the epigenetic target region set. For example, the depth of sequencing for nucleic acids corresponding to the sequence variant target region set may be at least 1.25-, 1.5-, 1.75-, 2-, 2.25-, 2.5-, 2.75-, 3-,3.5-, 4-, 4.5-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, or 15-fold greater, or 1.25- to 1.5-, 1.5- to 1.75-, 1.75- to 2-, 2- to 2.25-, 2.25- to 2.5-, 2.5- to 2.75-, 2.75- to 3-, 3- to 3.5-, 3.5- to 4-, 4- to4.5-, 4.5- to 5-, 5- to 5.5-, 5.5- to 6-, 6- to 7-, 7- to 8-, 8- to 9-, 9- to 10-, 10- to 11-, 11- to 12-, 13- to 14-, 14- to 15-fold, or 15- to 100-fold greater, than the depth of sequencing for nucleic acids corresponding to the epigenetic target region set. In some embodiments, said depth of sequencing is at least 2-fold greater. In some embodiments, said depth of sequencing is at least 5-fold greater. In some embodiments, said depth of sequencing is at least 10-fold greater. In some embodiments, said depth of sequencing is 4- to 10-fold greater. In some embodiments, said depth of sequencing is 4- to 100-fold greater. Each of these embodiments refer to the extent to which nucleic acids corresponding to the sequence-variable target region set are sequenced to a greater depth of sequencing than nucleic acids corresponding to the epigenetic target region set.
[0289] In some embodiments, the captured cfDNA corresponding to the sequence-variable target region set and the captured cfDNA corresponding to the epigenetic target region set are sequenced concurrently, e.g., in the same sequencing cell (such as the flow cell of an Illumina sequencer) and / or in the same composition, which may be a pooled composition resulting from recombining separately captured sets or a composition obtained by capturing the cfDNA corresponding to thesequence-variable target region set and the captured cfDNA corresponding to the epigenetic target region set in the same vessel.Analysis
[0290] In some embodiments, a method described herein comprises identifying the presence or absence of DNA produced by a tumor (or neoplastic cells, or cancer cells).
[0291] The present methods can be used to diagnose presence or absence of conditions, particularly cancer, in a subject, to characterize conditions (e.g., staging cancer or determining heterogeneity of a cancer), monitor response to treatment of a condition, effect prognosis risk of developing a condition or subsequent course of a condition. The present disclosure can also be useful in determining the efficacy of a particular treatment option. Successful treatment options may increase the amount of copy number variation or rare mutations detected in subject's blood if the treatment is successful as more cancers may die and shed DNA. In other examples, this may not occur. In another example, perhaps certain treatment options may be correlated with genetic profiles of cancers over time. This correlation may be useful in selecting a therapy.
[0292] Additionally, if a cancer is observed to be in remission after treatment, the present methods can be used to monitor residual disease or recurrence of disease.
[0293] The types and number of cancers that may be detected may include blood cancers, brain cancers, lung cancers, skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, bowel cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, mouth cancers, stomach cancers, solid state tumors, heterogeneous tumors, homogenous tumors and the like. Type and / or stage of cancer can be detected from genetic variations including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5 -methylcytosine.
[0294] Genetic data can also be used for characterizing a specific form of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data may allow characterization of specific sub-types of cancer that may be important in the diagnosis or treatment of that specific sub-type. This information may also provide a subject or practitioner cluesregarding the prognosis of a specific type of cancer and allow either a subject or practitioner to adapt treatment options in accord with the progress of the disease. Some cancers can progress to become more aggressive and genetically unstable. Other cancers may remain benign, inactive or dormant. The system and methods of this disclosure may be useful in determining disease progression.
[0295] Further, the methods of the disclosure may be used to characterize the heterogeneity of an abnormal condition in a subject. Such methods can include, e.g., generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile comprises a plurality of data resulting from copy number variation and rare mutation analyses. In some embodiments, an abnormal condition is cancer. In some embodiments, the abnormal condition may be one resulting in a heterogeneous genomic population. In the example of cancer, some tumors are known to comprise tumor cells in different stages of the cancer. In other examples, heterogeneity may comprise multiple foci of disease. Again, in the example of cancer, there may be multiple tumor foci, perhaps where one or more foci are the result of metastases that have spread from a primary site.
[0296] The present methods can be used to generate or profile, fingerprint or set of data that is a summation of genetic information derived from different cells in a heterogeneous disease. This set of data may comprise copy number variation, epigenetic variation, and mutation analyses alone or in combination.
[0297] The present methods can be used to diagnose, prognose, monitor or observe cancers, or other diseases. In some embodiments, the methods herein do not involve the diagnosing, prognosing or monitoring a fetus and as such are not directed to non-invasive prenatal testing. In other embodiments, these methodologies may be employed in a pregnant subject to diagnose, prognose, monitor or observe cancers or other diseases in an unborn subject whose DNA and other polynucleotides may co- circulate with maternal molecules.
[0298] An exemplary method for molecular tag identification of MBD-bead partitioned libraries through NGS which includes a step of subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample is as follows:1. Physical partitioning of an extracted DNA sample (e.g., extracted blood plasma DNA from a human sample, which has optionally been subjected to target capture as described herein)using a methyl-binding domain protein-bead purification kit, saving all elutions from process for downstream processing.2. Parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, the hypermethylated, residual methylation ('wash'), and hypomethylated partitions are ligated with NGS- adapters with molecular tags.3. Subj ect hypermethylated partition to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein.4. Re-combining all molecular tagged partitions, and subsequent amplification using adapterspecific DNA primer sequences.5. Capture / hybridization of re-combined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).6. Re-amplification of the captured DNA library, appending a sample tag. Different samples are pooled, and assayed in multiplex on an NGS instrument.7. Bioinformatics analysis of NGS data, with the molecular tags being used to identify unique molecules, as well deconvolution of the sample into molecules that were differentially MBD- partitioned. This analysis can yield information on relative 5-methylcytosine for genomic regions, concurrent with standard genetic sequencing / variant detection.
[0299] In some embodiments of methods described herein, including but not limited to the method shown above, the molecular tags consist of nucleotides that are not altered by the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein (e.g., mC along with A, T, and G where the procedure is bisulfite conversion or any other conversion that does not affect mC; hmC along with A, T, and G where the procedure is a conversion that does not affect hmC; etc.). In some embodiments of methods described herein, including but not limited to the method shown above, the molecular tags do not comprise nucleotides that are altered by the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein (e.g., the tags do not comprise unmodified C where the procedure is bisulfite conversion or any other conversion that affects C; the tags do not comprise mC where the procedure is a conversion that affects mC; the tags do not comprise hmC where the procedure is a conversion that affects hmC; etc.).
[0300] In general, the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA may instead be performed before the step of parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, this may be done where the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA is a separation, such as hmC-seal, and in such a case the separated populations may themselves be differentially tagged relative to each other. Such an exemplary method is as follows:1. Physical partitioning of an extracted DNA sample (e.g., extracted blood plasma DNA from a human sample, which has optionally been subjected to target capture as described herein) using a methyl-binding domain protein-bead purification kit, saving all elutions from process for downstream processing.2. Subject hypermethylated partition to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein.3. Parallel application of differential molecular tags and NGS-enabling adapter sequences to each partition. For example, the hypermethylated partition (or where applicable, two or more sub-partitions of the hypermethylated partition), residual methylation ('wash') partition, and hypomethylated partition are ligated with NGS- adapters with molecular tags.4. Re-combining all molecular tagged partitions, and subsequent amplification using adapterspecific DNA primer sequences.5. Capture / hybridization of re-combined and amplified total library, targeting genomic regions of interest (e.g., cancer-specific genetic variants and differentially methylated regions).6. Re-amplification of the captured DNA library, appending a sample tag. Different samples are pooled, and assayed in multiplex on an NGS instrument.7. Bioinformatics analysis of NGS data, with the molecular tags being used to identify unique molecules, as well deconvolution of the sample into molecules that were differentially MBD- partitioned. This analysis can yield information on relative 5-methylcytosine for genomic regions, concurrent with standard genetic sequencing / variant detection.Exemplary partitioning workflows
[0301] Exemplary workflows for partitioning and library preparation are provided herein. In some embodiments, some or all features of the partitioning and library preparation workflows may be used in combination.Partitioning
[0302] In some embodiments, sample DNA (e.g., between 5 and 200 ng) is mixed with methyl binding domain (MBD) buffer and magnetic beads conjugated with MBD proteins and incubated overnight. Methylated DNA (hypermethylated DNA) binds the MBD protein on the magnetic beads during this incubation. Non-methylated (hypomethylated DNA) or less methylated DNA (intermediately methylated) is washed away from the beads with buffers containing increasing concentrations of salt. For example, one, two, or more fractions containing non-methylated, hypomethylated, and / or intermediately methylated DNA may be obtained from such washes. Finally, a high salt buffer is used to elute the heavily methylated DNA (hypermethylated DNA) from the MBD protein. In some embodiments, these washes result in three partitions (hypomethylated partition, intermediately methylated fraction and hypermethylated partition) of DNA having increasing levels of methylation.
[0303] In some embodiments, the three partitions of DNA are desalted and concentrated in preparation for the enzymatic steps of library preparation.Library preparation
[0304] In some embodiments (e.g., after concentrating the DNA in the partitions), the partitioned DNA is made ligatable, e.g., by extending the end overhangs of the DNA molecules are extended, and adding adenosine residues to the 3’ ends of fragments and phosphorylating the 5’ end of each DNA fragment. DNA ligase and adapters are added to ligate each partitioned DNA molecule with an adapter on each end. These adapters contain partition tags (e.g., non-random, non-unique barcodes) that are distinguishable from the partition tags in the adapters used in the other partitions. Either before or after making the portioned DNA ligatable and performing the ligation, the hypermethylated partition is subjected to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA, such as any of those described herein. Where the procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA further partitions the hypermethylated partition, the ligation of adapters should be performed after the procedure so that the sub-partitions of the hypermethylated partition can bedifferentially tagged. Then, the three (or more) partitions are pooled together and are amplified (e.g., by PCR, such as with primers specific for the adapters).
[0305] Following PCR, amplified DNA may be cleaned and concentrated prior to enrichment. The amplified DNA is contacted with a collection of probes described herein (which may be, e.g., biotinylated RNA probes) that target specific regions of interest. The mixture is incubated, e.g., overnight, e g., in a salt buffer. The probes are captured (e.g., using streptavidin magnetic beads) and separated from the amplified DNA that was not captured, such as by a series of salt washes, thereby enriching the sample. After the enrichment, the enriched sample is amplified by PCR. In some embodiments, the PCR primers contain a sample tag, thereby incorporating the sample tag into the DNA molecules. In some embodiments, DNA from different samples is pooled together and then multiplex sequenced, e.g., using an Illumina NovaSeq sequencer.Additional features of certain disclosed methodsSamples
[0306] A sample can be any biological sample isolated from a subject. A sample can be a bodily sample. Samples can include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, cerebrospinal fluid synovial fluid, lymphatic fluid, ascites fluid, interstitial or extracellular fluid, the fluid in spaces between cells, including gingival crevicular fluid, bone marrow, pleural effusions, cerebrospinal fluid, saliva, mucous, sputum, semen, sweat, urine. Samples are preferably body fluids, particularly blood and fractions thereof, and urine. A sample can be in the form originally isolated from a subject or can have been subjected to further processing to remove or add components, such as cells, or enrich for one component relative to another. Thus, a preferred body fluid for analysis is plasma or serum containing cell-free nucleic acids. A sample can be isolated or obtained from a subject and transported to a site of sample analysis. The sample may be preserved and shipped at a desirable temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. A sample can be isolated or obtained from a subject at the site of the sample analysis. The subject can be a human, a mammal, an animal, a companion animal, a service animal, or a pet. The subject may have a cancer. The subject may not have cancer or a detectable cancer symptom. The subject may have been treated with one or more cancer therapy, e.g., any one or more of chemotherapies, antibodies, vaccines or biologies. The subject may be in remission. The subject may or may not be diagnosed of being susceptible to cancer orany cancer-associated genetic mutations / disorders. In some embodiments, the sample is a polynucleotides sample obtained from a tumor tissue biopsy.Tissue sample
[0307] In additional embodiments a sample can comprise a tissue sample from epithelial tissue (skin, lining of the digestive tract), connective tissue (bones, tendons, fat and other soft padding tissue), muscle tissue (heart, muscles of the limbs), and / or nervous tissue (brain, spinal cord, nerves). Other tissues include liver tissue including hepatocytes (liver cells), liver sinusoidal endothelial cells, heart for example myocardial cells, endocardial cells, brain including neurons, glial cells (astrocytes, microglia), neural stem cells, kidneys including renal tubular cells, glomerular cells, mesangial cells, lungs including alveolar cells, bronchial epithelial cells, pancreas including islets of Langerhans cells, acinar cells, skin including keratinocytes, melanocytes, fibroblasts, and blood including red blood cells, white blood cells (lymphocytes, monocytes, neutrophils), platelets. In some embodiments, the tissue samples may undergo preservation methods after extraction these include flash freezing, formaldehyde fixation (Formalin-fixed), paraffin embedding, cryopreservation, ethanol fixation, methanol fixation, RNAlater stabilization, Bouin’s solution fixation, zinc fixation, and / or freezing in optimal cutting temperature (OCT) compound.
[0308] The volume of plasma can depend on the desired read depth for sequenced regions. Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For examples, the volume can be 0.5 mL, 1 mL, 5 mL 10 mL, 20 mL, 30 mL, or 40 mL. A volume of sampled plasma may be 5 to 20 mL.
[0309] A sample can comprise various amount of nucleic acid that contains genome equivalents. For example, a sample of about 30 ng DNA can contain about 10,000 (104) haploid human genome equivalents and, in the case of cfDNA, about 200 billion (2xlOn) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents and, in the case of cfDNA, about 600 billion individual molecules.
[0310] A sample can comprise nucleic acids from different sources, e.g., from cells and cell-free of the same subject, from cells and cell-free of different subjects. A sample can comprise nucleic acids carrying mutations. For example, a sample can comprise DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations existing in germline DNA of a subject. Somatic mutations refer to mutations originating in somatic cells of a subject, e.g., cancer cells. A sample can comprise DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations). A sample can comprise an epigenetic variant (i.e. a chemical or proteinmodification), wherein the epigenetic variant associated with the presence of a genetic variant such as a cancer-associated mutation. In some embodiments, the sample comprises an epigenetic variant associated with the presence of a genetic variant, wherein the sample does not comprise the genetic variant.
[0311] Exemplary amounts of cell-free nucleic acids in a sample before amplification range from about 1 fg to about 1 pg, e.g., 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can comprise obtaining 1 femtogram (fg) to 200 ng cell-free nucleic acid molecules from samples.
[0312] Cell-free nucleic acids are nucleic acids not contained within or otherwise bound to a cell or in other words nucleic acids remaining in a sample after removing intact cells. Cell- free nucleic acids include DNA, RNA, and hybrids thereof, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi- interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or a hybrid thereof. A cell-free nucleic acid can be released into bodily fluid through secretion or cell death processes, e.g., cellular necrosis and apoptosis. Some cell-free nucleic acids are released into bodily fluid from cancer cells e.g., circulating tumor DNA, (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell free nucleic acids are produced by tumor cells. In some embodiments, cell free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.
[0313] Cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of molecules, with a mode of about 168 nucleotides and a second minor peak in a range between 240 to 440 nucleotides.
[0314] Cell-free nucleic acids can be isolated from bodily fluids through a fractionation or partitioning step in which cell-free nucleic acids, as found in solution, are separated from intact cells and other non-soluble components of the bodily fluid. Partitioning may include techniquessuch as centrifugation or filtration. Alternatively, cells in bodily fluids can be lysed and cell-free and cellular nucleic acids processed together. Generally, after addition of buffers and wash steps, nucleic acids can be precipitated with an alcohol. Further clean up steps may be used such as silica based columns to remove contaminants or salts. Non-specific bulk carrier nucleic acids, such as C 1 DNA, DNA or protein for bisulfite sequencing, hybridization, and / or ligation, may be added throughout the reaction to optimize certain aspects of the procedure such as yield.
[0315] After such processing, samples can include various forms of nucleic acid including double stranded DNA, single stranded DNA and single stranded RNA. In some embodiments, single stranded DNA and RNA can be converted to double stranded forms so they are included in subsequent processing and analysis steps.
[0316] Double-stranded DNA molecules in a sample and single stranded nucleic acid molecules converted to double stranded DNA molecules can be linked to adapters at either one end or both ends. Typically, double stranded molecules are blunt ended by treatment with a polymerase with a 5'-3' polymerase and a 3 '-5' exonuclease (or proof reading function), in the presence of all four standard nucleotides. Klenow large fragment and T4 polymerase are examples of suitable polymerase. The blunt ended DNA molecules can be ligated with at least partially double stranded adapter (e.g., a Y shaped or bell -shaped adapter). Alternatively, complementary nucleotides can be added to blunt ends of sample nucleic acids and adapters to facilitate ligation. Contemplated herein are both blunt end ligation and sticky end ligation. In blunt end ligation, both the nucleic acid molecules and the adapter tags have blunt ends. In sticky-end ligation, typically, the nucleic acid molecules bear an “A” overhang and the adapters bear a “T” overhang.Amplification
[0317] Sample nucleic acids flanked by adapters can be amplified by PCR and other amplification methods. Amplification is typically primed by primers binding to primer binding sites in adapters flanking a DNA molecule to be amplified. Amplification methods can involve cycles of denaturation, annealing and extension, resulting from thermocycling or can be isothermal as in transcription-mediated amplification. Other amplification methods include the ligase chain reaction, strand displacement amplification, nucleic acid sequence based amplification, and selfsustained sequence based replication.
[0318] In some embodiments, the present methods perform dsDNA ligations with T-tailed and C- tailed adapters, which result in amplification of at least 50, 60, 70 or 80% of double strandednucleic acids. Preferably the present methods increase the amount or number of amplified molecules relative to control methods performed with T-tailed adapters alone by at least 10, 15 or 20%.Tags
[0319] Tags comprising barcodes can be incorporated into or otherwise joined to adapters. Tags can be incorporated by ligation, overlap extension PCR among other methods.Molecular tagging strategies
[0320] In some embodiments, the nucleic acid molecules (from the sample of polynucleotides) may be tagged with sample indexes and / or molecular barcodes (referred to generally as “tags”). Tags may be incorporated into or otherwise joined to adapters by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR), among other methods. Such adapters may be ultimately joined to the target nucleic acid molecule. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are generally applied to introduce sample indexes to a nucleic acid molecule using conventional nucleic acid amplification methods. The amplifications may be conducted in one or more reaction mixtures (e.g., a plurality of microwells in an array). Molecular barcodes and / or sample indexes may be introduced simultaneously, or in any sequential order. In some embodiments, molecular barcodes and / or sample indexes are introduced prior to and / or after sequence capturing steps are performed. In some embodiments, only the molecular barcodes are introduced prior to probe capturing and the sample indexes are introduced after sequence capturing steps are performed. In some embodiments, both the molecular barcodes and the sample indexes are introduced prior to performing probe-based capturing steps. In some embodiments, the sample indexes are introduced after sequence capturing steps are performed. In some embodiments, molecular barcodes are incorporated to the nucleic acid molecules (e.g. cfDNA molecules) in a sample through adapters via ligation (e g., blunt-end ligation or sticky-end ligation). In some embodiments, sample indexes are incorporated to the nucleic acid molecules (e.g. cfDNA molecules) in a sample through overlap extension polymerase chain reaction (PCR). Typically, sequence capturing protocols involve introducing a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, e.g., a coding sequence of a genomic region and mutation of such region is associated with a cancer type.
[0321] In some embodiments, the tags may be located at one end or at both ends of the sample nucleic acid molecule. In some embodiments, tags are predetermined or random or semi-random sequence oligonucleotides. In some embodiments, the tags may be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotides in length. The tags may be linked to sample nucleic acids randomly or non-randomly.
[0322] In some embodiments, each sample is uniquely tagged with a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or sub-sample is uniquely tagged with a molecular barcode or a combination of molecular barcodes. In other embodiments, a plurality of molecular barcodes may be used such that molecular barcodes are not necessarily unique to one another in the plurality (e g., non-unique molecular barcodes). In these embodiments, molecular barcodes are generally attached (e.g., by ligation) to individual molecules such that the combination of the molecular barcode and the sequence it may be attached to creates a unique sequence that may be individually tracked. Detection of non-unique molecular barcodes in combination with endogenous sequence information (e g., the beginning (start) and / or end (stop) genomic location / position corresponding to the sequence of the original nucleic acid molecule in the sample, start and stop genomic positions corresponding to the sequence of the original nucleic acid molecule in the sample, the beginning (start) and / or end (stop) genomic location / position of the sequence read that is mapped to the reference sequence, start and stop genomic positions of the sequence read that is mapped to the reference sequence, sub-sequences of sequence reads at one or both ends, length of sequence reads, and / or length of the original nucleic acid molecule in the sample) typically allows for the assignment of a unique identity to a particular molecule. In some embodiments, beginning region comprises the first 1, first 2, the first 5, the first 10, the first 15, the first 20, the first 25, the first 30 or at least the first 30 base positions at the 5' end of the sequencing read that align to the reference sequence. In some embodiments, the end region comprises the last 1, last 2, the last 5, the last 10, the last 15, the last 20, the last 25, the last 30 or at least the last 30 base positions at the 3' end of the sequencing read that align to the reference sequence. The length, or number of base pairs, of an individual sequence read are also optionally used to assign a unique identity to a given molecule. As described herein, fragments from a single strand of nucleic acid having been assigned a unique identity, may thereby permit subsequent identification of fragments from the parent strand, and / or a complementary strand.
[0323] In certain embodiments, the number of different tags used to uniquely identify a number of molecules, z, in a class can be between any of 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11*z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z or 100*z (e.g., lower limit) and any of 100,000*z, 10,000*z, 1000*z or 100*z (e.g., upper limit). In some embodiments, molecular barcodes are introduced at an expected ratio of a set of identifiers (e.g., a combination of unique or non-unique molecular barcodes) to molecules in a sample. One example format uses from about 2 to about 1 ,000,000 different molecular barcode sequences, or from about 5 to about 150 different molecular barcode sequences, or from about 20 to about 50 different molecular barcode sequences, ligated to both ends of a target molecule. Alternatively, from about 25 to about 1,000,000 different molecular barcode sequences may be used. For example, 20-50 x 20-50 molecular barcode sequences (i.e., one of the 20-50 different molecular barcode sequences can be attached to each end of the target molecule) can be used. Such numbers of identifiers are typically sufficient for different molecules having the same start and stop points to have a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of molecules have the same combinations of molecular barcodes.
[0324] In some embodiments, the assignment of unique or non-unique molecular barcodes in reactions is performed using methods and systems described in, for example, U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U. S . Patent Nos. 6, 582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, different nucleic acid molecules of a sample may be identified using only endogenous sequence information (e.g., start and / or stop positions, subsequences of one or both ends of a sequence, and / or lengths).Bait sets; Capture moieties
[0325] As discussed above, nucleic acids in a sample can be subject to a capture step, in which molecules having target sequences are captured for subsequent analysis. Target capture can involve use of a bait set comprising oligonucleotide baits labeled with a capture moiety, such as biotin or the other examples noted below. The probes can have sequences selected to tile across a panel of regions, such as genes or different forms of nucleic acids including transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac,H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph. In some embodiments, a bait set can have higher and lower capture yields for sets of target regions such as those of the sequencevariable target region set and the epigenetic target region set, respectively, as discussed elsewhere herein. Such bait sets are combined with a sample under conditions that allow hybridization of the target molecules with the baits. Then, captured molecules are isolated using the capture moiety. For example, a biotin capture moiety by bead-based streptavidin. Such methods are further described in, for example, U.S. patent 9,850,523, issuing December 26, 2017, which is incorporated herein by reference.
[0326] Capture moieties include, without limitation, biotin, avidin, streptavidin, a nucleic acid comprising a particular nucleotide sequence, a hapten recognized by an antibody, and magnetically attractable particles. The extraction moiety can be a member of a binding pair, such as biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety that is attached to an analyte is captured by its binding pair which is attached to an isolatable moiety, such as a magnetically attractable particle or a large particle that can be sedimented through centrifugation. The capture moiety can be any type of molecule that allows affinity separation of nucleic acids bearing the capture moiety from nucleic acids lacking the capture moiety. Exemplary capture moieties are biotin which allows affinity separation by binding to streptavidin linked or linkable to a solid phase or an oligonucleotide, which allows affinity separation through binding to a complementary oligonucleotide linked or linkable to a solid phase.Collections of target-specific probes
[0327] In some embodiments, a collection of target-specific probes is used in methods described herein. In some embodiments, the collection of target-specific probes comprises target-binding probes specific for a sequence-variable target region set and target-binding probes specific for an epigenetic target region set. In some embodiments, the capture yield of the target-binding probes specific for the sequence-variable target region set is higher (e.g., at least 2-fold higher) than the capture yield of the target-binding probes specific for the epigenetic target region set. In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the sequence-variable target region set higher (e.g., at least 2-fold higher) than its capture yield specific for the epigenetic target region set.
[0328] In some embodiments, the capture yield of the target-binding probes specific for the sequence-variable target region set is at least 1.25-, 1.5-, 1.75-, 2-, 2.25-, 2.5-, 2.75-, 3-, 3.5-, 4-, 4.5-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, or 15 -fold higher than the capture yield of the target-binding probes specific for the epigenetic target region set. In some embodiments, the capture yield of the target-binding probes specific for the sequence-variable target region set is 1.25- to 1.5-,1.5- to 1.75-, 1.75- to 2-, 2- to 2.25-, 2.25- to 2.5-, 2.5- to 2.75-, 2.75- to 3-, 3- to 3.5-, 3.5- to 4-, 4- to 4.5-, 4.5- to 5-, 5- to 5.5-, 5.5- to 6-, 6- to 7-, 7- to 8-, 8- to 9-, 9- to 10-, 10- to 11-, 11- to 12-, 13- to 14-, or 14- to 15-fold higher than the capture yield of the target-binding probes specific for the epigenetic target region set.
[0329] In some embodiments, the collection of target-specific probes is configured to have a capture yield specific for the sequence-variable target region set at least 1.25-, 1.5-, 1.75-, 2-, 2.25-,2.5-, 2.75-, 3-, 3.5-, 4-, 4.5-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, or 15-fold higher than its capture yield for the epigenetic target region set. In some embodiments, the collection of targetspecific probes is configured to have a capture yield specific for the sequence-variable target region set is 1.25- to 1.5-, 1.5- to 1.75-, 1.75- to 2-, 2- to 2.25-, 2.25- to 2.5-, 2.5- to 2.75-, 2.75- to 3-, 3- to 3.5-, 3.5- to 4-, 4- to 4.5-, 4.5- to 5-, 5- to 5.5-, 5.5- to 6-, 6- to 7-, 7- to 8-, 8- to 9-, 9- to 10-, 10- to 11 -, 11- to 12-, 13- to 14-, or 14- to 15-fold higher than its capture yield specific for the epigenetic target region set.
[0330] The collection of probes can be configured to provide higher capture yields for the sequence-variable target region set in various ways, including concentration, different lengths and / or chemistries (e.g., that affect affinity), and combinations thereof. Affinity can be modulated by adjusting probe length and / or including nucleotide modifications as discussed below.
[0331] In some embodiments, the target-specific probes specific for the sequence-variable target region set are present at a higher concentration than the target-specific probes specific for the epigenetic target region set. In some embodiments, concentration of the target-binding probes specific for the sequence-variable target region set is at least 1.25-, 1.5-, 1.75-, 2-, 2.25-, 2.5-,2.75-, 3-, 3.5-, 4-, 4.5-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, or 15-fold higher than the concentration of the target-binding probes specific for the epigenetic target region set. In some embodiments, the concentration of the target-binding probes specific for the sequence-variable target region set is 1.25- to 1.5-, 1.5- to 1.75-, 1.75- to 2-, 2- to 2.25-, 2.25- to 2.5-, 2.5- to 2.75-,2.75- to 3-, 3- to 3.5-, 3.5- to 4-, 4- to 4.5-, 4.5- to 5-, 5- to 5.5-, 5.5- to 6-, 6- to 7-, 7- to 8-, 8- to 9-, 9- to 10-, 10- to 11-, 11- to 12-, 13- to 14-, or 14- to 15-fold higher than the concentration of the target-binding probes specific for the epigenetic target region set. In such embodiments, concentration may refer to the average mass per volume concentration of individual probes in each set.
[0332] In some embodiments, the target-specific probes specific for the sequence-variable target region set have a higher affinity for their targets than the target-specific probes specific for the epigenetic target region set. Affinity can be modulated in any way known to those skilled in the art, including by using different probe chemistries. For example, certain nucleotide modifications, such as cytosine 5-methylation (in certain sequence contexts), modifications that provide a heteroatom at the 2’ sugar position, and LNA nucleotides, can increase stability of double-stranded nucleic acids, indicating that oligonucleotides with such modifications have relatively higher affinity for their complementary sequences. See, e.g., Severin et al., Nucleic Acids Res. 39: 8740- 8751 (2011); Freier et al., Nucleic Acids Res. 25: 4429-4443 (1997); US Patent No. 9,738,894. Also, longer sequence lengths will generally provide increased affinity. Other nucleotide modifications, such as the substitution of the nucleobase hypoxanthine for guanine, reduce affinity by reducing the amount of hydrogen bonding between the oligonucleotide and its complementary sequence. In some embodiments, the target-specific probes specific for the sequence-variable target region set have modifications that increase their affinity for their targets. In some embodiments, alternatively or additionally, the target-specific probes specific for the epigenetic target region set have modifications that decrease their affinity for their targets. In some embodiments, the target-specific probes specific for the sequence-variable target region set have longer average lengths and / or higher average melting temperatures than the target-specific probes specific for the epigenetic target region set. These embodiments may be combined with each other and / or with differences in concentration as discussed above to achieve a desired fold difference in capture yield, such as any fold difference or range thereof described above.
[0333] In some embodiments, the target-specific probes comprise a capture moiety. The capture moiety may be any of the capture moieties described herein, e.g., biotin. In some embodiments, the target-specific probes are linked to a solid support, e g., covalently or non-covalently such as through the interaction of a binding pair of capture moieties. In some embodiments, the solid support is a bead, such as a magnetic bead.
[0334] In some embodiments, the target-specific probes specific for the sequence-variable target region set and / or the target-specific probes specific for the epigenetic target region set are a bait set as discussed above, e g., probes comprising capture moieties and sequences selected to tile across a panel of regions, such as genes.
[0335] In some embodiments, the target-specific probes are provided in a single composition. The single composition may be a solution (liquid or frozen). Alternatively, it may be a lyophilizate.
[0336] Alternatively, the target-specific probes may be provided as a plurality of compositions, e.g., comprising a first composition comprising probes specific for the epigenetic target region set and a second composition comprising probes specific for the sequence-variable target region set. These probes may be mixed in appropriate proportions to provide a combined probe composition with any of the foregoing fold differences in concentration and / or capture yield. Alternatively, they may be used in separate capture procedures (e.g., with aliquots of a sample or sequentially with the same sample) to provide first and second compositions comprising captured epigenetic target regions and sequence-variable target regions, respectively.Probes specific for epigenetic target regions
[0337] The probes for the epigenetic target region set may comprise probes specific for one or more types of target regions likely to differentiate DNA from neoplastic (e.g., tumor or cancer) cells from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein, e.g., in the sections above concerning captured sets. The probes for the epigenetic target region set may also comprise probes for one or more control regions, e.g., as described herein. Examples of epigenetic target regions can include different forms of nucleic acids including transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.
[0338] In some embodiments, the probes for the epigenetic target region probe set have a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the probes for the epigenetic target region set have a footprint in the range of 100-1000 kb, e.g., 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1,000 kb. In some embodiments, the probes for the epigenetic target region probe set have a footprint of less than 5 kb, at least 5 kb, e.g., at least 10, 20, or 50 kb.Hypermethylation variable target regions
[0339] In some embodiments, the probes for the epigenetic target region set comprise probes specific for one or more hypermethylation variable target regions. The hypermethylation variable target regions may be any of those set forth above. For example, in some embodiments, the probesspecific for hypermethylation variable target regions comprise probes specific for a plurality of loci listed in Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. In some embodiments, the probes specific for hypermethylation variable target regions comprise probes specific for a plurality of loci listed in Table 3, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 3. In some embodiments, the probes specific for hypermethylation variable target regions comprise probes specific for a plurality of loci listed in Table 2 or Table 3, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2 or Table 3. In some embodiments, for each locus included as a target region, there may be one or more probes with a hybridization site that binds between the transcription start site and the stop codon (the last stop codon for genes that are alternatively spliced) of the gene. In some embodiments, the one or more probes bind within 300 bp of the listed position, e.g., within 200 or 100 bp. In some embodiments, a probe has a hybridization site overlapping the position listed above. In some embodiments, the probes specific for the hypermethylation target regions include probes specific for one, two, three, four, or five subsets of hypermethylation target regions that collectively show hypermethylation in one, two, three, four, or five of breast, colon, kidney, liver, and lung cancers.Hypomethylation variable target regions
[0340] In some embodiments, the probes for the epigenetic target region set comprise probes specific for one or more hypomethylation variable target regions. The hypomethylation variable target regions may be any of those set forth above. For example, the probes specific for one or more hypomethylation variable target regions may include probes for regions such as repeated elements, e.g., LINE1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and satellite DNA, and intergenic regions that are ordinarily methylated in healthy cells may show reduced methylation in tumor cells.
[0341] In some embodiments, probes specific for hypomethylation variable target regions include probes specific for repeated elements and / or intergenic regions. In some embodiments, probes specific for repeated elements include probes specific for one, two, three, four, or five of LINE 1 elements, Alu elements, centromeric tandem repeats, pericentromeric tandem repeats, and / or satellite DNA.
[0342] Exemplary probes specific for genomic regions that show cancer-associated hypomethylation include probes specific for nucleotides 8403565-8953708 and / or 151104701- 151106035 of human chromosome 1. In some embodiments, the probes specific forhypomethylation variable target regions include probes specific for regions overlapping or comprising nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1.CTCF binding regions
[0343] In some embodiments, the probes for the epigenetic target region set include probes specific for CTCF binding regions. In some embodiments, the probes specific for CTCF binding regions comprise probes specific for at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, e g., such as CTCF binding regions described above or in one or more of CTCFBSDB or the Cuddapah et al., Martin et al., or Rhee et al. articles cited above. In some embodiments, the probes for the epigenetic target region set comprise at least 100 bp, at least 200 bp at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream regions of the CTCF binding sites.Epigenetic target regions
[0344] In some embodiments, the probes for the epigenetic target region set include different forms of nucleic acids including transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel , H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, H3S10ph.Transcriptional start sites
[0345] In some embodiments, the probes for the epigenetic target region set include probes specific for transcriptional start sites. In some embodiments, the probes specific for transcriptional start sites comprise probes specific for at least 10, 20, 50, 100, 200, or 500 transcriptional start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcriptional start sites, e.g., such as transcriptional start sites listed in DBTSS. In some embodiments, the probes for the epigenetic target region set comprise probes for sequences at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and downstream of the transcriptional start sites.Focal amplifications
[0346] As noted above, although focal amplifications are somatic mutations, they can be detected by sequencing based on read frequency in a manner analogous to approaches for detecting certain epigenetic changes such as changes in methylation. As such, regions that may show focal amplifications in cancer can be included in the epigenetic target region set, as discussed above. In some embodiments, the probes specific for the epigenetic target region set include probes specific for focal amplifications. In some embodiments, the probes specific for focal amplifications include probes specific for one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAFI . For example, in some embodiments, the probes specific for focal amplifications include probes specific for one or more of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the foregoing targets.Control regions
[0347] It can be useful to include control regions to facilitate data validation. In some embodiments, the probes specific for the epigenetic target region set include probes specific for control methylated regions that are expected to be methylated in essentially all samples. In some embodiments, the probes specific for the epigenetic target region set include probes specific for control hypomethylated regions that are expected to be hypomethylated in essentially all samples.Probes specific for sequence-variable target regions
[0348] The probes for the sequence-variable target region set may comprise probes specific for a plurality of regions known to undergo somatic mutations in cancer. The probes may be specific for any sequence-variable target region set described herein. Exemplary sequence-variable target region sets are discussed in detail herein, e.g., in the sections above concerning captured sets.
[0349] In some embodiments, the sequence-variable target region probe set has a footprint of at least 0.5 kb, e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb. In some embodiments, the epigenetic target region probe set has a footprint in the range of 0.5-100 kb, e.g., 0.5-2 kb, 2-10 kb, 10-20 kb, 20-30 kb, 30-40 kb, 40-50 kb, 50-60 kb, 60-70 kb, 70-80 kb, 80-90 kb, and 90-100 kb.
[0350] In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or at 70 of the genes of Table 4. In some embodiments, probes specific for the sequence-variable targetregion set comprise probes specific for the at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs of Table 4. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 4. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least a portion of at least 1, at least 2, or 3 of the indels of Table 4. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes of Table 5. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs of Table 5. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions of Table 5. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels of Table5. In some embodiments, probes specific for the sequence-variable target region set comprise probes specific for at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes of Table 6.
[0351] In some embodiments, the probes specific for the sequence-variable target region set comprise probes specific for target regions from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1 , ALK, BRAF, CCND1 , CDK2A, CTNNB 1 , EGFR, ERBB2, ESRI , FGFR1 , FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK1 1, TP53, and U2AFl .Compositions comprising captured DNA
[0352] Provided herein is a combination comprising first and second populations of captured DNA. The first population may comprise or be derived from DNA with a cytosine modification in a greater proportion than the second population. The first population may comprise a form of aI l lfirst nucleobase originally present in the DNA with altered base pairing specificity and a second nucleobase without altered base pairing specificity, wherein the form of the first nucleobase originally present in the DNA prior to alteration of base pairing specificity is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the form of the first nucleobase originally present in the DNA prior to alteration of base pairing specificity and the second nucleobase have the same base pairing specificity. The second population does not comprise the form of the first nucleobase originally present in the DNA with altered base pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleobase is a modified or unmodified cytosine and the second nucleobase is a modified or unmodified cytosine. The first and second nucleobase may be any of those discussed herein in the Summary or with respect to subjecting the first subsample to a procedure that affects a first nucleobase in the DNA differently from a second nucleobase in the DNA of the first subsample.
[0353] In some embodiments, the first population comprises a sequence tag selected from a first set of one or more sequence tags and the second population comprises a sequence tag selected from a second set of one or more sequence tags, and the second set of sequence tags is different from the first set of sequence tags. The sequence tags may comprise barcodes.
[0354] In some embodiments, the first population comprises protected hmC, such as glucosylated hmC.
[0355] In some embodiments, the first population was subjected to any of the conversion procedures discussed herein, such as bisulfite conversion, Ox-BS conversion, TAB conversion, ACE conversion, TAP conversion, TAPSP conversion, or CAP conversion. In some embodiments, the first population was subjected to protection of hmC followed by deamination of mC and / or C.
[0356] In some embodiments of the combination, the first population comprises or was derived from DNA with a cytosine modification in a greater proportion than the second population and the first population comprises first and second subpopulations, and the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first nucleobase and the second nucleobase have the same base pairing specificity. In some embodiments, the second population does not comprise the first nucleobase. In some embodiments, the first nucleobase is a modified or unmodified cytosine, and the second nucleobase is a modified or unmodified cytosine, optionally wherein the modified cytosine is mC or hmC. In some embodiments, the first nucleobase is amodified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine, optionally wherein the modified adenine is mA.
[0357] In some embodiments, the first ...
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method configured to use a previously generated neural network, trained to determine cancer recurrence in a patient, comprising:(a) extracting a plurality of molecules from a tumor sample obtained from the patient to acquire chromatin conformation data from the tumor sample;(b) extracting a plurality of cell-free DNA (cfDNA) molecules from a blood sample obtained from the patient, wherein the blood sample is collected from the same patient at a later timepoint following a curative intent treatment;(c) sequencing a portion of the cfDNA molecules extracted from the blood sample, to acquire sequencing reads, wherein the plurality of sequencing reads comprise genetic sequence information from the cfDNA molecules;(d) inputting a plurality of features to the previously trained neural network, wherein the plurality of features comprises at least a portion of the chromatin conformation data from the tumor sample and at least a portion of the genetic sequence information from the cfDNA molecules; and(e) outputting from the classifier a determination of whether or not the cancer reoccurred in the cancer patient.
2. The method claim 1, wherein the tumor sample from the subject is genetically assayed to further obtain tumor specific somatic genetic variation.
3. The method of claims 1 or 2, wherein the input features for the neural network comprise at least a portion of the somatic genetic variation obtained from the tumor.
4. The method of any one of the preceding claims, wherein the tumor sample from the subject is epigenetically assayed to further obtain tumor specific methylation patterns.
5. The method of claim 4, wherein the input features for the neural network comprise at least a portion of the information of tumor specific methylation patterns.
6. The method of claim 5, wherein the input features for the neural network comprise at least a portion of the tumor specific somatic genetic variation and at least a portion of the tumor specific methylation patterns.
7. The method of any one of the preceding claims, wherein prior to sequencing, the cfDNA molecules are enriched for genomic regions that comprise tumor specific somatic genetic variation obtained in the genetic assay.
8. A neural network for determining the reoccurrence of a tumor in a subject previously treated to remove the tumor, wherein the neural network has been trained with at least one of a plurality of feature vectors, each vector comprising at least one of: i) chromatin conformation data obtained from the tumor prior to resecting the tumor from the subject, ii) tumor specific somatic genetic variation data obtained from the tumor prior to resecting the tumor from the subject, iii) DNA sequence information from a cell-free DNA (cfDNA) sample obtained at a later time point after resecting the tumor from the subject; and iv) the tumor recurrence status of the patient at the time the cfDNA sample was obtained.
9. A method for determining a plurality of epigenetic landscapes in a sample from a subject having a disease, comprising:(a) determining a plurality of classes of molecules in the sample using a plurality of assays, wherein the sample comprises cell -free DNA (cfDNA);(b) generating a set of target regions by capturing each of the plurality of classes of molecules isolated from the sample, wherein each target region of the set of target regions spans an epigenetic landscape from the plurality of epigenetic landscapes known to be associated with the disease in the subject; and(c) for each of the plurality of epigenetic landscapes, determining the sequence of at least a segment of each target region for each of the plurality of classes of molecules, wherein the segment comprises an epigenetic biomarker or a portion thereof, thereby determining the plurality of epigenetic landscapes in the sample from the subject having the disease.
10. The method of claim 9, further comprising, extracting the plurality of molecules from a tumor tissue sample obtained from the same patient to acquire chromatin conformation data from the tumor tissue sample.
11. The method of any one of claims 9 or 10, wherein the epigenetic landscapes further comprise information layers comprising biochemical states of each of the plurality of classes of molecules.
12. The method of claim 11, wherein the biochemical states comprise cytosine methylation, transcription factor binding sites (TFBS), mRNA expression, fragmentomic patterns, fragmentomic levels, fragment end point densities, histone acetylation or methylation marks associated with poised enhancers including H3K4mel, H3K27ac, H3K27me3, promoter regions including H3K4me3, H3 / H4ac, H3K4mel, H3K27me3, H3K9me3 and / or H3.3, open chromatin including H3Ac and H4Ac, H3K4mel, H3K4me2, H3K4me3, H2BK120ub, H3.3, and H3S10ph.
13. The method of any one of claims 9 to 12, wherein each of the epigenetic landscapes in the plurality of epigenetic landscapes comprises higher-order chromatin states associated with the disease in the subject.
14. The method of any one of claims 9 to 13, wherein the sample comprises cell-free DNA molecules isolated from blood.
15. The method of any one of claims 9 to 14, wherein the disease is cancer comprising at least one of adenocarcinoma, basal cell carcinoma, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, cholangiocarcinoma, colorectal cancer, endometrial cancer, esophageal cancer, gallbladder cancer, gastric cancer, germ cell tumors, glioma, head and neck cancer, hepatocellular carcinoma, Kaposi sarcoma, kidney cancer, lip and oral cavity cancer, liver cancer, lung cancer, melanoma, mesothelioma, neuroendocrine tumors, ovarian cancer, pancreatic cancer, penile cancer, prostate cancer, sarcoma, skin cancer, small cell lung cancer, squamous cell carcinoma, stomach cancer, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, and vulvar cancer.
16. The method of any one of claims 9 to 15, wherein the plurality of classes of molecules comprises modified DNA, a DNA-protein complex, DNA-histone complex, and / or mRNA.
17. The method of any one of claims 9 to 15, wherein the plurality of assays comprise TAM- ChlP, CUT&RUN, CUT&Tag, ChlP-seq, and / or RNA-seq.
18. The method of any one of claims 9 to 17, wherein the determining the sequence of at least a segment of each target region for each of the plurality of classes of molecules comprises sequencing the DNA or mRNA molecules to generate sequencing reads.
19. The method of claim 9, wherein the epigenetic biomarker further comprises a plurality of fragmentomic biomarkers.
20. The method of claim 19, wherein the fragmentomic biomarkers comprise, fragment length score, fragment end-densities, the ratio of short (100-150 bp) to long (151-220 bp) fragments, and / or fragment length signatures.
21. The method of any one of the preceding claims, wherein the plurality of assays are performed at a plurality of time points.
22. The method of any one of the preceding claims, wherein the subject has not undergone surgery to remove disease.
23. The method of any one of the preceding claims, wherein the subject has undergone surgery to remove disease.
24. The method of any one of the preceding claims, further comprising administering a therapy to the individual, where the therapy is effective in treating the disease having at least one of the epigenetic landscapes from the plurality of epigenetic landscapes determined in the sample from the subject.
25. The method of claim 24, wherein the therapy comprises an epigenetic therapy, targeting DNA methylation and histone modification mechanisms.
26. The method of claim 25, wherein the epigenetic therapy comprises azacitidine, decitabine, vorinostat, and / or romidepsin.
27. The method of claim 24, wherein the therapy comprises a BET protein inhibitors, a histone methyltransferase inhibitors, a lysine-specific demethylase inhibitors, and / or Bromodomain inhibitors.
28. The method of any one of the preceding claims, wherein the capturing comprises hybrid capture using a set of capture probes, multiplexed PCR amplification, or multiplexed qPCR using sets of primers.
29. A method for determining disease reoccurrence in a patient, comprising:(a) obtaining a tissue sample from a tumor in a patient, wherein the tissue sample is obtained prior curative intent treatment;(b) determining a first plurality of molecular classes in the tissue sample through a plurality of assays, wherein said molecular classes include at least one of cross-linked DNA- protein complexes, DNA-DNA complexes, DNA-histone complexes, and RNA;(c) generating a first set of target regions by capturing at least one of the molecular classes isolated from the tissue sample of the patient;(d) selecting a subset of the target regions based on one or more functional properties and / or features, wherein the subset of target regions comprise a patient specific cell-free DNA (cfDNA) assay;(e) obtaining a cell free DNA sample from the patient; and(f) determining the subset of patient specific target regions in the cfDNA sample obtained from the patient by: i) determining a second plurality of molecular classes in a cfDNA sample using a second plurality of assays, ii) generating a second set of target regions by capturing each of the second plurality of molecular classes isolated from the cfDNA, thereby determining the plurality of patient specific target regions in the cfDNA sample from the patient.
30. A method to remove noise in epigenetic datasets using chromatin interaction data comprising:(a) providing a first sample from a subject comprising cellular genomic DNA and processing the first sample to determine a set of chromatin interaction partners in the cellular genomic DNA of the subject;(b) providing a second sample from the subject comprising cell-free DNA (cfDNA) molecules processing the second sample to determine a set of features in the cfDNA molecules of the subject, wherein the features comprise genetic, epigenetic, and / or fragmentomic markers; and(c) applying a machine learning model trained on characteristics and relationships between interaction partners to identify and correct discordant features in the second sample comprising cfDNA, thereby remove noise in epigenetic datasets using chromatin interaction data.
31. The method of claim 30, further comprising determining that the chromatin interaction partners exhibit similar epigenetic state, and using the similarity information to error correct epigenetic states between interaction partners.
32. The method of claim 30, wherein the first and second samples are taken from the subject at different time points.
33. The method of claim 30, wherein the first and second samples are taken from the subject at the same time points.
34. The method of any one of the preceding claims, further comprising annotating chromatin interaction partners with functional genomic data to enrich functionally relevant genomic loci.
35. The method of any one of the preceding claims, wherein the first dataset comprises promoter capture Hi-C.
36. The method of any one of the preceding claims, wherein the first sample comprises a tissue sample and the second sample comprises cell-free DNA (cfDNA).
37. A method for enriching disease associated signals comprising:(a) providing a first sample from a subject comprising cellular genomic DNA and processing the first sample to determine a set of chromatin interaction partners in the cellular genomic DNA of the subject;(b) providing a second sample from the subject comprising cfDNA molecules processing the second sample to determine a set of features in the cfDNA molecules of the subject, wherein the features comprise genetic, epigenetic, and / or fragmentomic markers; and(c) annotating the set of chromatin interaction partners in the first sample with one or more features in the cfDNA molecules in the second sample, thereby integrating information from the first sample comprising cellular genomic DNA and the second sample comprising cfDNA molecules from the subject to enrich disease associated signals in the subject based on the combined genetic, epigenetic and fragmentomic landscape of the interaction partners, wherein the landscape is indicative of a biological function or disfunction.
38. A method for determining disease associated signals comprising:(a) providing a first sample from a subject comprising cellular genomic DNA and processing the sample by: i) linking the cellular genomic DNA to produce linked genomic DNA; ii) contacting the linked genomic DNA with one or more reagents that fragment and attach at least one label to the genomic DNA to produce labeled nucleic acid molecules;iii) contacting the labeled nucleic acid molecules with one or more reagents that proximity ligates the labeled nucleic acid molecules to generate ligated nucleic acid molecules; iv) shearing the ligated nucleic acid molecules to generate sheared nucleic acid molecules; v) sequencing the sheared nucleic acid molecules to produce a first dataset comprising a first plurality of sequencing reads;(b) providing a second sample from the subject comprising cell-free DNA (cfDNA) molecules and processing the sample by: i) partitioning the cfDNA molecules on the basis of at least one feature thereby generating a plurality of subsamples comprising at least a first and a second subsample; ii) attaching a set of adapters to the cfDNA molecules in each of the sub samples to generate adapter-ligated cfDNA molecules in the first and the second subsamples, wherein the molecular barcodes attached to the cfDNA molecules in the first subsample differ from the molecular barcodes attached to the cfDNA molecules in the second subsample, wherein the adapters are attached to both ends of the cfDNA molecules in the first and second subsamples; iii) sequencing at least one subsample to generate a second dataset comprising a second plurality of sequencing reads;(c) processing the first dataset comprising a first plurality of sequencing reads and the second dataset comprising a second plurality of sequencing reads to determine overlapping genomic regions between the first and second datasets; and(d) determining from the overlapping genomic regions disease associated signals in the subject.
39. A method for enriching disease associated signals comprising:(a) providing a first sample from a subject comprising cellular genomic DNA and processing the sample by: i) cross-linking the cellular genomic DNA to produced crosslinked nucleic acid molecules;ii) digesting the cross-linked nucleic acid molecules to produce digested, cross-linked nucleic acid molecules comprising at least one 3' overhang sequence and at least one 5' overhang sequence; iii) repairing the at least one 3’ and 5’ overhang sequences; iv) incorporating a capture label on the least 3’ and 5’ overhang sequences thereby creating labeled overhang sequences; v) ligating the labeled overhang sequences to generate a junction marker between the at least one 3’ end overhang sequence and at least one 5’ overhang sequence in the nucleic acid molecules; vi) fragmenting the nucleic acid molecules in (v) generate, fragmented nucleic acid molecules comprising a junction marker and fragmented nucleic acid molecules lacking a junction maker; vii) separating the fragmented nucleic acid molecules by contacting the molecules in (iv) with a capture molecule with affinity for the capture label to generate fragmented nucleic acid molecules comprising a junction marker, wherein the junction marker comprises a capture label bound to the capture molecule and fragmented nucleic acid molecules lacking a junction marker; and viii) sequencing the fragmented nucleic acid molecules comprising a junction marker to generate a first set of sequencing reads comprising chromatin interaction partners;(b) providing a second sample comprising cell-free DNA (cfDNA) molecules and processing the sample by: i) partitioning the cfDNA molecules on the basis of at least one genetic, epigenetic and / or fragmentomic biomarkers thereby generating a plurality of subsamples comprising at least a first and a second subsample; ii) ligating a first set of adapters to the plurality of the subsamples, wherein the molecular barcodes for the first subsample differ from the molecular barcodes of the second subsample to generate adapter-ligated DNA molecules in the first and the second subsamples, wherein the adapters are ligated at both ends of the doublestranded DNA molecules; iii) pooling the adapter ligated molecules from the first subsample and the adapter ligated molecules from the second subsample to generate a mixture comprising adapter-ligated molecules from the first subsample and adapter ligated molecules from the second subsample; iv) amplifying the mixture in (iii) to generate amplified molecules; v) enriching the amplified molecules in (iv) to generate a plurality of enriched molecules; vi) amplifying the enriched molecules in (v) to generate amplified enrich molecules; vii) sequencing the amplified enrich molecules to generate a second set of sequencing reads.(c) processing the second set of sequencing reads to generate a second dataset comprising a plurality of genetic, epigenetic, and / or fragmentomic markers;(d) determining for the chromatin interaction partners in the first dataset one or more genetic, epigenetic, and / or fragmentomic states using the second dataset thereby generating annotated chromatin interaction partners, thereby integrating information from the first dataset and the second datasets; and(e) determining from the annotated chromatin interaction partners in (d) a plurality of cancer associated signals thereby enriching disease associated signals.
40. A computer-implemented method to enrich disease associated signals in a plurality of samples from a subject, comprising:(a) providing, in computer memory, a first dataset from a first sample comprising a plurality of chromatin interaction partners and a second dataset from a second sample comprising epigenetic data;(b) determining from the second dataset an epigenetic state for the plurality of chromatin interaction partners in the first dataset; and(c) determining a plurality of disease associated signals from the epigenetic state of the plurality of interaction partners, thereby enriching disease associated signals in the plurality of samples from the subject.
Citation Information
Patent Citations
Compositions and methods for analyzing modified nucleotides
US10260088B2
Hyperactive AID / APOBEC and hmC dominant TET enzymes
US10961525B2
Oligonucleotides
US20010053519A1
Method and apparatus for imaging a sample on a device
US20030152490A1
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1