Detection of promoter methylation
Simultaneous genomic and epigenomic analysis addresses the precision gap in cancer treatment by integrating promoter methylation data, improving treatment efficacy through comprehensive biomarker integration.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2024-04-12
- Publication Date
- 2026-05-01
AI Technical Summary
Current cancer treatment methods lack precision due to the separation and incompatibility of genomic and epigenomic analysis, leading to incomplete diagnostic testing and ineffective treatment decisions.
A method for simultaneously detecting genomic and epigenomic characteristics, particularly promoter methylation, using a combination of genomic and epigenomic components to improve treatment prediction accuracy.
Enhances treatment prediction accuracy by incorporating comprehensive biomarker information, enabling personalized cancer treatment decisions based on both genomic and epigenomic data.
Smart Images

Figure 2026514005000001 
Figure 2026514005000002 
Figure 2026514005000003
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims the interests of U.S. Provisional Patent Application No. 63 / 509,917, filed June 23, 2023, and U.S. Provisional Patent Application No. 63 / 495,688, filed April 12, 2023, both of which are incorporated herein by reference in their entirety.
[0002] Field of Invention This specification describes diagnostic methods for selecting therapies for personalized cancer treatment by simultaneously detecting genomic and epigenomic characteristics from a single patient sample, including the quantification of promoter methylation. [Background technology]
[0003] background Treatment choices for cancer patients are not always precise. Often, DNA, RNA, and proteins derived from patient samples are analyzed for patterns that can predict responses to specific treatments. These biomarkers can range from single genes (e.g., EGFR, via real-time PCR) and proteins (e.g., HER2, via immunohistochemistry) to complex genomic signatures (e.g., tumor mutational burden, via next-generation sequencing). Testing workflows for multiple types of analytes are generally separate and cannot be combined due to incompatibility in separation, chemistry, and quantification processes. As a result, diagnostic testing cannot investigate the full range of biomarkers with significant informational value, and multiple independent tests must be performed when multi-omics results are desired. In most cases, multiple tests are not performed for several reasons, including a lack of sufficient patient samples, and clinical decisions are made based on incomplete information.
[0004] There is a great need for improvement in personalized medicine within this technological field. By simultaneously investigating the genomic and epigenomic states of patient samples, more accurate predictions of treatment effectiveness can be achieved.
[0005] This specification describes the simultaneous testing and incorporation of information derived from both genomic and epigenomic components of patient samples, allowing such diagnostic testing to take into account additional information that would otherwise be unavailable. Of particular interest is the methylation status, including methylation patterns related to epigenetic allele status, especially in promoter regions, which can explain cases where variability in the effectiveness of treatments in patients, including PARP inhibitors (PARPi), cannot be explained by genomic modification. [Overview of the project] [Means for solving the problem]
[0006] Summary of the Invention A method is described herein that includes the steps of detecting methylation in one or more promoter regions of at least one of several genes, and generating multiple methylation calls to quantify the methylation of one or more promoter regions. In other embodiments, the method includes the step of obtaining a sample. In other embodiments, the method includes the step of characterizing a sample by processing the quantity of methylation in one or more promoter regions. In other embodiments, the step of characterizing a sample, as included in the method, includes HRD, cancer-derived promoter methylation, familial form of colorectal cancer, or Lynch syndrome tumor type. In other embodiments, the promoter includes a 5kb region upstream of the transcription start site (TSS), where the 5kb region is further refined using one or more of the exclusion of a costume panel region, methylation peaks found in clinical samples, and peaks found in normal samples. In other embodiments, the TSS is defined at the transcript level. In other embodiments, the TSS is defined at the gene level. In other embodiments, the method includes the step of determining the ratio of the number of molecules overlapping the target region, normalized by a total positive control molecule. In other embodiments, a step in the method to determine a ratio includes filtering molecules based on at least the number of overlapping CpGs. In other embodiments, a step in the method to quantify methylation of one or more promoter regions is based on the number of methylated CpGs. In other embodiments, the method includes a step to refine one or more promoter regions based on at least literature annotations, common methylation peak locations, and / or public datasets. In other embodiments, the genes include tumor suppressor genes, HRR genes, and IO genes. In other embodiments, the HRR genes include at least BRCA1 and BRCA2.
[0007] In other embodiments, the method includes the step of comparing with a minimum methylation threshold derived from a population of training samples. In other embodiments, the training samples include cancer-free samples. In other embodiments, the minimum methylation threshold for Cole is a minimum molecular count of 1 to 100, and the minimum methylation score per gene is the 95th quantile in normality + 8 × 10⁻⁶. 5 The method includes at least one of the following: or being the maximum value among median + 5 * median absolute deviation. In other embodiments, a step of quantifying methylation of one or more promoter regions, included in the method, predicts the therapeutic response. In other embodiments, a step of quantifying methylation of one or more promoter regions, included in the method, is combined with the MSI-H status.
[0008] In other embodiments, the treatment comprises one or more of the following: an immune checkpoint inhibitor, a poly(ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor. In other embodiments, the immune checkpoint inhibitor is pembrolizumab. In other embodiments, the method comprises olaparib or talazoparib, which are poly(ADP-ribose) polymerase (PARP) inhibitors. In other embodiments, the treatment is a combination of a PI3K and mTOR inhibitor and a poly(ADP-ribose) polymerase (PARP) inhibitor. In other embodiments, the PI3K and mTOR inhibitors are jedatricib, and the poly(ADP-ribose) polymerase (PARP) inhibitor is talazoparib.
[0009] A method is described herein that includes the steps of determining at least one promoter region from multiple genes obtained from multiple samples, determining a methylation score for the promoter region to generate multiple methylation calls and / or quantifications of promoter methylation, and processing the multiple methylation calls to generate a prediction that the test sample represents the genomic state.
[0010] A method is described herein that includes the steps of determining promoter regions of multiple genes obtained from multiple samples, determining methylation scores for the promoter regions to generate multiple methylation calls and / or quantifications of promoter methylation, and processing the multiple methylation calls to generate a prediction that the test sample exhibits a genomic state. In other embodiments, the genomic state includes HRD, cancer-derived promoter methylation, familial form of colorectal cancer, or Lynch syndrome tumor type. In other embodiments, the promoter includes a 5kb region upstream of the TSS, and the 5kb region is further refined using one or more of the exclusion of a custom panel region, methylation peaks found in clinical samples, and peaks found in normal samples. In other embodiments, the TSS is defined at the transcript level. In other embodiments, the TSS is defined at the gene level. In other embodiments, the methylation score is determined as the ratio of the number of molecules overlapping the target region, normalized by the total positive control molecules. In other embodiments, molecules supporting the methylation score are filtered based on at least the number of overlapping CpGs. In other embodiments, the promoter region is refined based on at least literature annotations, common methylation peak locations, and / or public datasets. In other embodiments, the genes include tumor suppressor genes, HRR genes, and IO genes. In other embodiments, the HRR genes include at least BRCA1 and BRCA2. In other embodiments, the call includes deriving a minimum methylation threshold from a population of training samples. In other embodiments, the training samples include cancer-free samples. In other embodiments, the minimum methylation threshold for the call is,
[0011] Minimum molecular count of 1 to 100, and / or minimum methylation score per gene is the 95th quantile in normality + 8 × 10⁻⁶. 5Alternatively, it includes being the maximum of the median + 5 * median absolute deviation. In other embodiments, a combination of promoter methylation call and MSI-H status predicts the therapeutic response. In other embodiments, the treatment includes one or more of the following: an immune checkpoint inhibitor, a poly(ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor. In other embodiments, the immune checkpoint inhibitor is pembrolizumab. In other embodiments, the poly(ADP-ribose) polymerase (PARP) inhibitor is olaparib or talazoparib. In other embodiments, the treatment is a combination of a PI3K and mTOR inhibitor and a poly(ADP-ribose) polymerase (PARP) inhibitor. In other embodiments, the PI3K and mTOR inhibitor is jedatricib, and the poly(ADP-ribose) polymerase (PARP) inhibitor is talazoparib.
[0012] This specification describes a method comprising the steps of determining promoter regions of BRCA1 and BRCA2 obtained from multiple samples, determining methylation scores for the promoter regions to generate multiple methylation calls, and processing the multiple methylation calls to generate a prediction that the patient exhibits loss of both BRCA1 and BRCA2 alleles.
[0013] This specification describes a method comprising the steps of determining promoter regions of BRCA1 and BRCA2 obtained from multiple samples, determining methylation scores for the promoter regions to generate multiple methylation calls, processing the multiple methylation calls to generate a prediction that the patient exhibits loss of both BRCA1 and BRCA2 alleles, and determining that the patient is a candidate for treatment using PARPi.
[0014] A method is described herein that includes the steps of determining the promoter regions of BRCA1 and BRCA2, each obtained from multiple samples; determining methylation scores for the promoter regions to generate multiple methylation calls; processing the multiple methylation calls to generate a prediction that the patient exhibits loss of both BRCA1 and BRCA2 alleles; and determining that the patient is a candidate for treatment with jedatricib and talazoparib. In other embodiments, the method includes sensitizing advanced TNBC or BRCA1 / 2 mutant breast cancer to PARP inhibition with talazoparib using jedatricib.
[0015] A method is described herein that includes the steps of determining the promoter region of MLH1 obtained from multiple samples, determining a methylation score for the promoter region to generate multiple promoter methylation calls, and determining from genomic data that a patient is BRAF V600E positive, wherein detection of promoter methylation in a BRAF V600E-positive patient identifies the patient as a patient at risk of hereditary / familial forms of colorectal cancer or Lynch syndrome-related tumor types.
[0016] A method is described herein that includes the steps of obtaining sequencing reads derived from a sample of interest using a computing system having one or more hardware processors and memory; determining one or more classification regions corresponding to a plurality of genes contained in the sample; and determining the methylation level of one or more classification regions by generating a quantitative scale derived from the sequencing reads in the sample of interest. In other embodiments, the method includes the step of obtaining a sample. In other embodiments, the method includes the step of characterizing a sample by processing the methylation level of one or more classification regions. In other embodiments, the method includes the step of characterizing a sample, which includes determining cancer-related HRD status, promoter methylation. In other embodiments, the quantitative scale includes determining the ratio of the number of molecules overlapping with the classification region, normalized by a total positive control molecule, where the molecule represents a threshold amount of methylated cytosine. In other embodiments, the quantitative scale is compared to a predetermined threshold to call the methylation status of one or more classification regions. In other embodiments, the step of determining the ratio includes filtering molecules based on at least a threshold amount of methylated cytosine. In other embodiments, determining the methylation level of one or more classification regions is based on the number of methylated CpGs. In other embodiments, the classification region includes a promoter region. In other embodiments, one or more classification regions individually correspond to genomic regions where the cytosine methylation rate in the genomic region of nucleic acid derived from cells obtained from a subject with cancer differs from the cytosine methylation rate in the genomic region of nucleic acid derived from cells obtained from a subject without cancer. In other embodiments, multiple samples and additional samples include cell-free nucleic acids. In other embodiments, the method includes the step of generating a model by having a computing system perform a training process using training data, wherein the training process includes the computing system determining one or more additional weights for individual samples included in the training data based on cancer indicators for individual samples within a threshold confidence level.In other embodiments, the cancer index for individual samples is outside the threshold confidence level, and the method includes the step of applying a penalty to the weights of individual samples during the training process by a computing system. Any of the methods of the foregoing claims includes the steps of the computing system performing one or more first iterations of the training process on a model using one or more machine learning algorithms with respect to a portion of the training data, and the computing system generating first output data for the model based on one or more first iterations of the training process, wherein the first output data corresponds to one or more first additional indexes indicating the presence of cancer in a first individual subject of a plurality of subjects, and the first individual subject corresponds to a portion of the training data. In other embodiments, the method includes the steps of: having a computing system combine first output data and training data to produce additional training data; having the computing system perform one or more second iterations of the training process on a model using a portion of the additional training data; and having the computing system generate second output data for the model based on one or more second iterations of the training process, wherein the second output data indicates one or more second additional indicators of cancer presence in a second individual subject among a plurality of subjects, and the second individual subject corresponds to a portion of the additional training data. In other embodiments, the weights for individual classification regions of a plurality of classification regions are determined based on the first and second output data. In other embodiments, the method includes the steps of: having a computing system determine that several indicators of cancer presence, determined during one or more iterations of the training process, are thresholds for at least one or more samples included in the training data; and having a computing system determine that any modification to one or more weights of the model is either not modified or modified by a minimal amount.In other embodiments, the method comprises determining, by a computing system, that the number of additional indicators of the presence of cancer determined during one or more iterations of a training process is less than a threshold for one or more additional samples included in training data, and determining, by the computing system, that a modification to one or more additional weights of the model is modified more than a minimum amount. In other embodiments, the method comprises combining a plurality of nucleic acids derived from at least one of a subject's blood and tissue with a solution containing a certain amount of methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution, and performing a plurality of washes of the nucleic acid-MBD protein solution with a salt solution to produce a plurality of nucleic acid fractions, wherein each nucleic acid fraction has a threshold number of methylated cytosines within a region of a plurality of nucleic acids having at least a threshold cytosine-guanine content. In other embodiments, one of the plurality of washes is performed using a solution having a certain concentration of sodium chloride (NaCl) to produce one nucleic acid fraction of a plurality of nucleic acid fractions having a series of binding strengths to the MBD protein. In other embodiments, the method comprises determining that a first nucleic acid fraction is associated with a first distribution of a plurality of distributions of nucleic acids, the first distribution corresponding to a first range of binding strengths to the MBD protein, and binding a first molecular barcode to the nucleic acids of the first nucleic acid fraction, the first molecular barcode being included within a first set of molecular barcodes associated with the first distribution, and determining that a second nucleic acid fraction is associated with a second distribution of a plurality of distributions of nucleic acids, the second distribution corresponding to a second range of binding energies to the MBD protein that is different from the first range of binding strengths to the MBD protein, and binding a second molecular barcode to the nucleic acids of the second nucleic acid fraction, the second molecular barcode being included within a second set of molecular barcodes associated with the second distribution.
[0017] In other embodiments, the method comprises producing at least a portion of a plurality of samples for use in generating sequencing reads by combining at least a portion of some nucleic acid fractions with a restriction enzyme that cleaves a molecule having a certain amount of one or more unmethylated cytosines, wherein a threshold amount of methylated cytosine corresponds to a minimum frequency of methylated cytosines within a region having at least a threshold cytosine-guanine content.
[0018] A method is described herein that includes obtaining sequencing reads from a sample of interest by a computing system having one or more hardware processors and memory, determining one or more classification regions corresponding to a plurality of genes included in the sample, determining the methylation level of the one or more classification regions by generating a quantitative measure that includes a ratio of the number of molecules overlapping the classification region normalized by a total positive control molecule, wherein the molecule exhibits a threshold amount of methylated cytosine, and comparing the quantitative measure to a predetermined threshold to call the methylation status of the one or more classification regions.
[0019] In various embodiments, determining the quantitative measure can include combining a plurality of nucleic acids from at least one of a blood or tissue sample of interest with a solution containing a certain amount of a methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution, and performing a plurality of washes of the nucleic acid-MBD protein solution with a salt solution to produce some nucleic acid fractions. In some cases, an individual nucleic acid fraction has a threshold number of methylated cytosines within a region of a plurality of nucleic acids having at least a threshold cytosine-guanine content. Then, one of the plurality of washes is performed using a solution having one concentration of sodium chloride (NaCl) to produce one nucleic acid fraction of some nucleic acid fractions having a series of binding strengths to the MBD protein.
[0020] It can be determined that a first nucleic acid fraction is associated with a first compartment of multiple compartments of nucleic acid, and a first molecular barcode can be attached to the nucleic acid of the first nucleic acid fraction; then, it can be determined that a second nucleic acid fraction is associated with a second compartment of multiple compartments of nucleic acid, and a second molecular barcode can be attached to the nucleic acid of the second nucleic acid fraction, where the first compartment corresponds to a first range of binding strength to MBD protein, and the first molecular barcode is contained within a first set of molecular barcodes associated with the first compartment; and the second compartment corresponds to a second range of binding energy to MBD protein, which is different from the first range of binding strength to MBD protein, and the second molecular barcode is contained within a second set of molecular barcodes associated with the second compartment.
[0021] In some cases, at least a portion of several nucleic acid fractions can be combined with restriction enzymes that cleave molecules having one or more unmethylated cytosines in a certain amount to produce at least a portion of multiple samples used to produce sequencing reads, where the threshold amount of methylated cytosine corresponds to the lowest frequency of methylated cytosine within a region having at least a threshold cytosine-guanine content.
[0022] Furthermore, at least a portion of several nucleic acid fractions can be combined with restriction enzymes that cleave molecules having one or more methylated cytosines in a certain amount to produce at least a portion of multiple samples used to produce sequencing reads, where the threshold amount of unmethylated cytosine corresponds to the highest frequency of uncleaved methylated cytosines within a region having at least a threshold cytosine-guanine content. [Brief explanation of the drawing]
[0023] [Figure 1] The BRCA1 promoter region. The 11 CpG sites (circles) with core promoter activity, which have been shown to be hypermethylated in breast cancer (pink circles), are covered by the panel BRCA1 promoter definition. The numbers indicate the nucleotide position relative to the transcription start of BRCA1.
[0024] [Figure 2] The 95% limit of detection (LoD) for BRCA1 is 0.6%, determined by titration of the well-characterized breast cancer cell line HCC-38. Previously, HCC-38 was confirmed to be epigenetically silenced by promoter methylation at the BRCA1 locus by bisulfite sequencing and RT-PCR (Stefansson 2012, Xu 2010). In comparison, our method detected no BRCA1 promoter methylation in the 80 cancer-free donors tested, demonstrating 100% specificity.
[0025] [Figure 3] The prevalence of BRCA1 promoter methylation across cancer types in the selected patient cohort. Note that differences in methylation frequencies may be due to the unselected, non-randomized patient subtype composition in the GuardantInfinity cohort, as well as cancer stage (methylation may be lost in patients during the course of treatment), and may not be directly comparable to the TCGA patient cohort. Abbreviations: Ovarian (OVCA), Breast (BRCA), Bladder (BLCA), Lung Adenocarcinoma (LUAD), Colorectal Adenocarcinoma (COAD), Lung Squamous Cell Carcinoma (LUSC), Melanoma (SKCM).
[0026] [Figure 4]Oncoprint analysis of epigenetic and genomic alterations in the HRR gene. Pathogenicity was defined as any nonsense, frameshift, rearrangement, or pathogenic ClinVar missense mutation in the HRR gene described above. Somatic truncated mutations in ATM and CHEK2 were excluded from this analysis due to their potential interference with clonal hematopoiesis. Promoter methylation is highlighted in pink—note that these alterations are primarily mutually exclusive with other pathogenic alterations in other HRR genes.
[0027] [Figure 5] Characteristics of promoter coverage in the sample panel. This figure shows the minimum of 10 CpGs in a 200 bp sliding window.
[0028] [Figure 6] Promoter methylation region definition removal: Sex chromosomes + normal noisy regions. Aggregate the upstream 5kb of each TSS for each gene. If there are multiple TSSs, aggregate at the gene level. Definition approach: Separate for each TSS and report at the transcript level; refine promoter regions by literature and other data (e.g., MBD compartment peaks, RNA / methylation associations). For almost all of the 16,000 genes, at least two probes within the promoter region are covered in the panel.
[0029] [Figure 7] Exemplary analysis and verification: Detection limit.
[0030] [Figure 8] Epigenetic MLH1 vs. MSI-H association, MSI promoter definition.
[0031] [Figure 9] Regional patterns of MSI-H versus MSS / MSI-L for MLH1+.
[0032] [Figure 10]BRCA1 - Clinical samples and cell lines.
[0033] [Figure 11] Promoter methylation: Partial methylation versus complete methylation. In some cases, gene inactivation can be caused only by complete methylation. Promoter methylation is often present on one allele, while the other allele is inactivated by other events (e.g., coexistence of BRCA LoH / promoter in HRD+). Functional methylation changes may include the distinction between partial and complete methylation.
[0034] [Figure 12] EM-seq Overview: Panel Design. To demonstrate the capabilities of the detection scheme, an orthogonal approach was used to design the EM-seq panel for pan-cancer methylation enrichment. 13,090 probes were used, targeting 1.54 Mb (125,080 CpGs) at a depth of 15,000 ×. Epigenome probes covered 1.00 Mb and 90,949 CpGs (65% of sequences, 73% of CpGs); of these, 876 kb and 70,493 CpGs overlapped with refseq promoter regions, indicating MLH1 and BRCA1.
[0035] [Figure 13] Consistency of EM-seq data with publicly available array data. The accuracy of orthogonal EM-seq results using neat cell lines is shown. Variant-level (left) and probe-level (right) beta are consistent between KM12 EM-seq (x axis) and Illumina 450K array data (NCI, y axis). Probe beta is the average of all CpG betas within each EM-seq region (both datasets).
[0036] [Figure 14]TF from the perspective of the epigenome detection region versus TF from the perspective of the EM-Seq region. Here, the accuracy of the positive predictive accuracy (PPA) in samples with an epigenome level of death (LOD) or higher (rough estimate >0.3% TF = red square in the plot on the left). For EM-Seq, since the majority of samples have a beta value (call threshold) >0.1%, the PPA can be expected to exceed 80%. Positive clinical samples with a mixture of positive and negative promoter methylation calls across all genes in the epigenome detection and EMSeq panel. Negative (cancer-free) clinical samples with mainly negative promoter methylation calls across genes in the epigenome detection and EMSeq panel. [Modes for carrying out the invention]
[0037] Detailed explanation BRCA1 promoter methylation (PM) is an early-onset event in cancer, present in 3–65.2% of all breast tumors and 30–65% of triple-negative tumors, depending on the subtype. BRCA1 promoter methylation is associated with homologous recombination repair deficiency (HRR), early onset of breast and ovarian cancer, and improved clinical response to adjuvant chemotherapy. To date, no diagnostic assay exists that comprehensively evaluates both BRCA1 promoter methylation and genomic alterations in cell-free circulating tumor DNA (ctDNA). Here, we have established a detection method to investigate both promoter methylation status and genomic alterations, which had not been achieved before, and to quantify methylation regardless of the presence or absence of epigenetic allele status. This multi-mode detection of BRCA1 PM and genomic alterations in a cohort of breast cancer patients using an epigenome detection platform including methyl-binding domain distribution enables liquid biopsy assays and genome-wide methylation detection that examine more than 800 genes. BRCA1 PM was evaluated in ctDNA from 1016 patients with advanced breast cancer, along with genome sequencing of over 800 genes. PM profiling of 398 cancer-related genes was performed using an epigenome methylation detection assay. The predefined promoter regions of each covered gene were analyzed, including a novel approach to promoter definition. Methylation scores were calculated for each gene in each sample and used as a basis for PM calling. The limit of detection (LoD) was determined by in silico and experimental titration of ctDNA from clinical samples and cell lines with known gene PMs into plasma from cancer-free donors.
[0038] Furthermore, establishing the detection approach described above makes it possible to determine epigenetic allele status at a systematic level. Allele-specific methylation patterns play a crucial role in regulating gene expression and maintaining normal cellular function, and disruption of these patterns may contribute to pathogenesis, including carcinogenesis. Imprinting is a form of allele-specific methylation pattern in which one allele of a gene is methylated and silenced, depending on whether it is inherited from the mother or father. Differential methylation of imprinted genes and the resulting single-allele expression are important for normal development and physiological function, and abnormal changes in these imprinting patterns (loss or acquisition of methylation) may lead to increased susceptibility to developmental disorders and diseases, including cancer. For example, loss of imprinting (LOI) results in the expression of both alleles of a gene that is normally imprinted, potentially doubling the expression of genes that promote cell growth, a feature common to various cancers. See, for example, Figure 10 Panel A. In additional cases, typically unmethylated and active tumor suppressor genes may become methylated on one allele. This methylation can silence gene expression from that allele, potentially contributing to cancer progression if the other allele is lost or mutated. A well-known example is the p16 gene (CDKN2A), which can undergo hypermethylation in various cancers, including melanoma, bladder cancer, and others. See, for example, Figure 10, Panel B. Partial allele-specific methylation patterns (see Figure 10, Panel C) may have a more subtle impact on gene function compared to complete allele methylation. This selective methylation can occur in specific regions of a gene, such as promoters, enhancers, or other regulatory elements that affect the gene's transcriptional activity in a cell-type-specific manner. In cancer, partial methylation of the promoter region of a tumor suppressor gene can downregulate gene expression without completely silencing the gene.This partial methylation can occur only in specific CpG islands within the promoter region. Furthermore, methylation of enhancer regions can modulate enhancer activity and therefore indirectly affect the expression of genes associated with these enhancers. Partial methylation of enhancer regions may result in altered gene expression profiles that contribute to carcinogenesis.
[0039] Current approaches either omit testing both the genomic and epigenomic characteristics of patient samples, or conduct multiple tests separately. Omitting genomic or epigenomic information can result in prescribing cancer treatments that are known to be ineffective, or withholding cancer treatments that are known to be effective if both genomic and epigenomic information were available. For example, a patient with the KRASG12C biomarker might be prescribed a KRAS inhibitor, but if epigenomic information indicates that the KRAS promoter is methylated and therefore the gene is silenced, the KRAS inhibitor would be ineffective. On the other hand, a patient without a detected BRCA1 mutation might not be prescribed a PARP inhibitor, but if epigenomic information indicates that the BRCA1 promoter is methylated and therefore the gene is silenced, the patient would be a good candidate for a PARP inhibitor. Often, multiple tests are not conducted for a variety of reasons, including a lack of sufficient patient samples. Other drawbacks include non-reimbursement, inconvenience, and lack of available commercial offerings.
[0040] Cancer can be indicated by epigenetic variations, such as methylation. Examples of methylation alterations in cancer include localized increases in DNA methylation at CpG islands in transcription start sites (TSSs) of genes involved in normal growth regulation, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation may be associated with abnormal loss of transcriptional ability of the genes involved and occurs at least as frequently as point mutations and deletions, contributing to altered gene expression. DNA methylation profiling can be used to detect regions of the genome with varying degrees of methylation ("differentially methylated regions" or "DMRs") that are altered during development or perturbed by disease, such as cancer or any cancer-related disease. The genomes of cancer cells have imbalances in the DNA methylation patterns described above, and therefore in the functional packaging of DNA. Thus, abnormalities in chromatin composition can be linked to methylation alterations and, when analyzed together, may contribute to the enhancement of cancer profiling. MBD distribution can be used in combination with fragment mix data, such as fragments mapped to start and stop positions (correlated with nucleosome location), fragment length, and associated nucleosome occupancy, to improve biomarker detection rates in chromatin structure analysis in hypermethylation studies.
[0041] Methylation profiling can involve determining methylation patterns across different regions of the genome. For example, molecules can be allocated based on their degree of methylation (e.g., the relative number of methylation sites per molecule), sequenced, and then the sequences of molecules in different compartments can be mapped to a reference genome. This can reveal regions of the genome that are more or less highly methylated compared to other regions. Thus, genomic regions can have different degrees of methylation, in contrast to individual molecules.
[0042] Nucleic acid molecules can be characterized by modifications, which may include various chemical modifications or protein modifications (i.e., epigenetic modifications). Non-limiting examples of chemical modifications include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation involves the addition of a methyl group to cytosine at a CpG site (cytosine followed by guanine in the nucleic acid sequence). In some embodiments, DNA methylation involves the addition of a methyl group to adenine, such as at N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation involves the addition of a methyl group to the 5C position of cytosine to produce 5-methylcytosine (m5c). In some embodiments, methylation includes derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxylcytosine (caryboxylcytosine) (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine to produce 3-methylcytosine (3mC). Other examples include N6-methyladenine or glycosylation. DNA methylation involves the addition of a methyl group to DNA (e.g., CpG) and can alter the expression of methylated DNA regions. Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA within a promoter region is methylated, gene transcription may be repressed. DNA methylation is crucial for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruption of epigenetic regulation, such as suppression, can lead to disease, such as cancer. Promoter methylation in DNA can indicate cancer.
[0043] A CpG dimolacule is the dinucleotide CpG (cytosine-phosphate-guanine, i.e., cytosine followed by guanine in the 5'→3' direction of the nucleic acid sequence) on the sense strand of a double-stranded DNA molecule and its complementary CpG on the antisense strand. A CpG dimolacule can be either fully methylated or partially methylated (methylated on only one strand). CpG dinucleotides are underpresented in the normal human genome, and the majority of CpG dinucleotide sequences are transcriptionally inactive (e.g., in the periconomere regions of chromosomes and in DNA heterochromatin regions of repeat elements) and methylated. However, many CpG islands, particularly around transcription start sites (TSSs), are protected from such methylation.
[0044] Protein modification includes binding to chromatin components, particularly histones (including their modified forms), and binding to other proteins, such as proteins involved in replication or transcription. This disclosure provides methods for processing and analyzing nucleic acids having varying degrees of modification, such that the nature of their original modifications is correlated with nucleic acid tags and can be decoded by sequencing the tags during nucleic acid analysis. The genetic variation of nucleic acid modifications in a sample, including single-stranded (e.g., ssDNA or RNA) or double-stranded molecules (e.g., dsDNA), can then be correlated with the degree of modification (epigenetic variation) of that nucleic acid in the original sample.
[0045] DNA loss can reduce the presence of one or more types of DNA, making it difficult to detect the presence of one or more types of DNA, such as cfDNA. In one or more additional scenarios, existing methods for measuring DNA methylation, such as enrichment or depletion methods, may have relatively high levels of resolution, e.g., about 100 base pairs (bp) to about 200 bp, which can make it difficult to accurately determine the amount of DNA methylation. The accuracy of determining DNA methylation can affect the accuracy of the tumor fraction estimate for a sample. Since the tumor fraction can be used to determine whether a sample originates from an object in which a tumor is present, the accuracy of determining the tumor fraction estimate can affect diagnostic and / or treatment decisions for an individual.
[0046] More specifically, the techniques described herein enable the quantification of promoter region methylation. Jedatricib is an intravenously administered PI3K and mTOR inhibitor that has been shown to be safe in patients with metastatic breast cancer, either alone or in combination with oral therapy. Previous studies have shown that PI3K inhibitors reduce the nucleotide pool necessary for DNA synthesis and phase S progression. Furthermore, inhibition of PI3K / mTOR may interfere with the interaction between PI3K and homologous recombination complexes, thereby increasing the PARP enzyme's dependence on DNA repair. Based on this data, a combination of a PI3K inhibitor and a PARP inhibitor may lead to a novel, non-chemotherapy treatment option for TNBC with wild-type BRCA, potentially improving the moderate PFS seen with PARP inhibitors as a monotherapy in BRCA1 / 2 mutant advanced breast cancer. The hypothesis for this trial is that jedatricib sensitizes advanced TNBC or BRCA1 / 2 mutant breast cancer to talazoparib-based PARP inhibition. Of particular interest is determining the recommended phase 2 dose for the combination of jedatricib and talazoparib, and evaluating the efficacy of this combination in advanced HER2-negative breast cancer that is triple-negative or BRCA1 / 2 positive (mutated / deleted). sample
[0047] The sample may be any biological sample isolated from the subject. The sample may be a body sample. Examples of samples include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsy, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, fluids of the intercellular space, gingival crevicular exudate, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Preferably, the sample is a body fluid, in particular blood and its fractions, as well as urine. The sample may be in the form originally isolated from the subject, or it may have been subjected to further processing to remove or add components, such as cells, or to enrich one component with another. Therefore, preferred body fluids for analysis are plasma or serum containing cell-free nucleic acids. Samples can be isolated or obtained from subjects and transported to the site of sample analysis. Samples can be stored or shipped at a desired temperature, e.g., room temperature, 4°C, -20°C, and / or -80°C. Samples can be isolated or obtained from subjects at the site of sample analysis. Subjects may be humans, mammals, animals, companion animals, service animals, or pets. Subjects may have cancer. Subjects may not have cancer or detectable symptoms of cancer. Subjects may have been treated with one or more cancer treatments, e.g., chemotherapy, antibodies, vaccines, or biology. Subjects may be in remission. Subjects may or may not have been diagnosed as susceptible to cancer or any cancer-related gene mutation / disorder.
[0048] The volume of plasma may depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, the volume could be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled may be 5–20 mL.
[0049] The sample may contain varying amounts of nucleic acids, including genome equivalents. For example, a sample of about 30 ng of DNA may contain about 10,000 (10 4 ) Contains haploid human genome equivalents, and in the case of cfDNA, approximately 200 billion (2 × 10⁻¹⁶). 11 It may contain ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, it may contain about 600 billion individual molecules.
[0050] The sample may contain nucleic acids from different origins, e.g., from cells and cell-free cells of the same subject, or from cells and cell-free cells of different subjects. The sample may contain nucleic acids with mutations. For example, the sample may contain DNA with germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of the subject. Somatic mutations refer to mutations originating from somatic cells of the subject, e.g., cancer cells. The sample may contain DNA with cancer-associated mutations (e.g., cancer-associated somatic mutations). The sample may contain epigenetic variants (i.e., chemical or protein modifications), where the epigenetic variant is associated with the presence of genetic variants, e.g., cancer-associated mutations. In some embodiments, the sample contains epigenetic variants associated with the presence of genetic variants, where the sample does not contain genetic variants.
[0051] Exemplary amounts of cell-free nucleic acids in the sample before amplification range from approximately 1 fg to approximately 1 μg, for example, 1 pg to 200 ng, 1 ng to 100 ng, and 10 ng to 1000 ng. For example, the amount may be up to approximately 600 ng, up to approximately 500 ng, up to approximately 400 ng, up to approximately 300 ng, up to approximately 200 ng, up to approximately 100 ng, up to approximately 50 ng, or up to approximately 20 ng of cell-free nucleic acid molecules. The amount may be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity may be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.
[0052] Cell-free nucleic acids are nucleic acids that are not contained in cells or are not otherwise bound to cells, or in other words, nucleic acids that remain in a sample after intact cells have been removed. Examples of cell-free nucleic acids include DNA, RNA, and their hybrids, which include genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, nucleolar small RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death processes, such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free embryonic DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.
[0053] Cell-free nucleic acids have an exemplary size distribution of approximately 100–500 nucleotides, with molecules of 110–approximately 230 nucleotides accounting for about 90% of the molecules, the mode being approximately 168 nucleotides, and a second minor peak in the range of 240–440 nucleotides. Cell-free nucleic acids can be isolated from body fluids by fractionation or partitioning steps that separate the cell-free nucleic acids found in the solution from intact cells and other insoluble components of the body fluid. Partitioning steps may include techniques such as centrifugation or filtration. Alternatively, cells in the body fluid may be lysed, and the cell-free nucleic acids and cellular nucleic acids may be processed together. Generally, after buffer addition and washing steps, nucleic acids can be precipitated with alcohol. Further washing steps, such as using silica-based columns, may be used to remove impurities or salts. Nonspecific bulk carrier nucleic acids, such as Cot1-DNA, DNA, or proteins, for bisulfite sequencing, hybridization, and / or ligation may be added throughout the reaction in certain aspects of the procedure, for example, to optimize yield.
[0054] Following such processing, the sample may contain various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA may be converted to double-stranded form so that they can be included in subsequent processing and analysis steps.
[0055] analyte The analytes may include nucleic acid analytes and non-nucleic acid analytes. This disclosure provides for detecting genetic variations in biological samples from a subject. The biological samples may include polynucleotides derived from cancer cells. The polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNA), or any combination thereof. The biological samples may include tumor tissue derived from a biopsy, for example. In some cases, the biological samples may include blood or saliva. In certain cases, the biological samples may include cell-free DNA ("cfDNA") or circulating tumor DNA ("ctDNA"). Cell-free DNA may be present in blood, for example.
[0056] Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphorylated proteins, specific phosphorylated or acetylated variants of proteins, amidated variants of proteins, hydroxylated variants of proteins, methylated variants of proteins, ubiquitinated variants of proteins, sulfated variants of proteins, viral proteins (e.g., viral capsids, viral envelopes, viral coats, viral accessories, viral glycoproteins, viral spikes, etc.), extracellular and intracellular proteins, antibodies, and antigen-binding fragments. Non-nucleic acid analytes include receptors, antigens, surface proteins, transmembrane proteins, surface antigen classification proteins, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, intercellular interaction protein complexes, antigen presentation complexes, major histocompatibility complexes, engineered T cell receptors, T cell receptors, B cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modifications of cell surface proteins (e.g., phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, or lipid addition), gap junctions, and adhesion junctions.
[0057] In general, systems, apparatus, methods, and compositions can be used to analyze any number of analytes, including both nucleic acid analytes and non-nucleic acid analytes. For example, the number of analytes analyzed could be at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, at least about 1,000, at least about 10,000, at least about 100,000 or more different analytes present within a certain area of the sample or within individual features of the substrate. Methods for performing multiplexed assays to analyze two or more different analytes will be discussed in later sections of this disclosure.
[0058] One or more nucleic acid analytes and / or non-nucleic acid analytes constitute a set of intermolecular interactions in the biological system under test (e.g., cells), and these intermolecular interactions can be thought of as an "interactome"—intermolecular interactions occurring between molecules belonging to different biochemical families (e.g., proteins, nucleic acids, lipids, carbohydrates), and also between molecules within a given family. In various embodiments, the interactome is a protein-DNA interactome (a network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes). In other embodiments, the interactome refers to a protein-protein interaction network (PPI), or protein-protein interaction network (PIN). The methods described herein enable the testing and analysis of interactomes. Techniques such as proteogenomics (e.g., whole-genome sequencing, whole-exome sequencing, and RNA-seq and mass spectrometry) can support the testing of interactomes.
[0059] analysis This method can be used to diagnose a condition in a subject, particularly the presence of cancer; to characterize the condition (e.g., to stage cancer or determine cancer heterogeneity); to monitor the response to treatment of the condition; and to determine the risk of developing the condition or the prognosis of the subsequent course of the condition. This disclosure may also be useful in determining the effectiveness of a particular treatment option. A successful treatment option may increase the amount of copy number variation or rare mutations detected in the subject's blood, because if the treatment is successful, more cancer cells may be killed and DNA may be shed. In other cases, this may not occur. In another embodiment, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a treatment. Furthermore, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.
[0060] The types and number of cancers that can be detected include blood cancers, brain cancers, lung cancers, skin cancers, nasal cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, skin cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, solid tumors, heterogeneous tumors, and homogeneous tumors. Cancer type and / or stage can be detected from genetic variations, including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, alterations in chromosomal structure, gene fusions, chromosome fusions, gene shortenings, gene amplifications, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0061] Genetic and other analyte data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in terms of both composition and staging. Genetic profiling data may enable the characterization of specific subtypes of cancer. This characterization may be important for the diagnosis or treatment of that particular subtype. This information may also provide subjects or workers with clues about the prognosis of a particular type of cancer, enabling either the subjects or workers to adapt treatment options as the disease progresses. Some cancers may become more invasive and genetically unstable as they progress. Other cancers may remain benign, inactive, or dormant. The systems and methods of this disclosure may be useful in determining disease progression.
[0062] This analysis is also useful in determining the effectiveness of specific treatment options. A successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood, because if the treatment is successful, more cancer cells may be killed and DNA may be shed. In other cases, this may not occur. In another example, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a treatment. Furthermore, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.
[0063] This method can also be used to detect genetic variations in conditions other than cancer. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection, and a particular immune state can be monitored. In this example, copy number variation analysis can be performed over time to create a profile of how a particular disease may progress. Using the detection of copy number variations, or even rare mutations, it is possible to determine how a population of pathogens changes during the course of infection. This can be particularly important during chronic infections, such as HIV / AIDS or hepatitis infections, where viruses can change their life cycle and / or mutate into more virulent forms during the course of infection. Because immune cells attempt to destroy transplanted tissue, this method can be used to determine or profile the rejection activity of the host body in order to monitor the state of transplanted tissue, as well as to modify the course of treatment or prevent rejection.
[0064] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of an abnormal condition in a subject. Such methods may include, for example, the step of generating a gene profile of extracellular polynucleotides derived from the subject, where the gene profile includes multiple data obtained by analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may result in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of cancer. In other examples, the heterogeneity may constitute multiple disease lesions. Furthermore, in the example of cancer, there may be multiple tumor lesions, possibly one or more lesions resulting from metastasis spreading from the primary site.
[0065] This method can be used to generate or profile fingerprints or sets of data that aggregate genetic information from different cells in heterogeneous diseases. These data sets may include, alone or in combination, analysis of copy number variations and mutations.
[0066] This method can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods described herein do not involve diagnosing, prognosing, or monitoring a fetus, and therefore do not pertain to non-invasive prenatal testing. In other embodiments, these methodologies may be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in prenatal subjects where DNA and other polynucleotides can co-circulate with maternal molecules.
[0067] Determination of the 5-methylcytosine pattern of nucleic acids Bisulfite-based sequencing and its variations provide means for determining the methylation patterns of nucleic acids. In some embodiments, determining the methylation pattern includes distinguishing 5-methylcytosine (5mC) from unmethylated cytosine. In some embodiments, determining the methylation pattern includes distinguishing N6-methyladenine from unmethylated adenine. In some embodiments, determining the methylation pattern includes distinguishing 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC) from unmethylated cytosine. Examples of bisulfite sequencing, but not limited to these, include oxidative bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reductive bisulfite sequencing (redBS-seq).
[0068] Oxidative bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC by first converting 5hmC to 5fC and then proceeding with bisulfite sequencing as previously described. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish between 5mc and 5hmC. In TAB-seq, 5hmC is protected by glucosylation. Then, 5mC is converted to 5caC using the Tet enzyme, and then bisulfite sequencing is proceeded as previously described. Reductive bisulfite sequencing is used to distinguish 5fC from modified cytosine.
[0069] Generally, bisulfite sequencing involves dividing a nucleic acid sample into two aliquots and treating one aliquot with bisulfite. The bisulfite converts native cytosines and certain modified cytosine nucleotides (e.g., 5-formylcytosine or 5-carboxylcytosine) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not. Comparison of the nucleic acid sequences of molecules from the two aliquots reveals which cytosines were converted to uracil and which were not. Consequently, modified and unmodified cytosines can be determined. Initially dividing the sample into two aliquots is inconvenient for samples containing only small amounts of nucleic acids and / or composed of heterogeneous cell / tissue origins, such as body fluids containing cell-free DNA.
[0070] This disclosure provides methods for enabling bisulfite sequencing and its deformation. These methods function by linking nucleic acids in a population to a capture moiety, i.e., a label that can be captured or immobilized. Capture moieties include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing specific nucleotide sequences, haptens recognized by antibodies, and magnetically adsorbable particles. Extraction moieties may be members of binding pairs, e.g., biotin / streptavidin or hapten / antibody. In some embodiments, a capture moiety attached to a sample is captured by its binding pair attached to an isolateable moiety, e.g., a magnetically adsorbable particle or a larger particle that can be settled by centrifugation. The capture moiety may be any type of molecule that enables affinity separation of nucleic acids having a capture moiety from nucleic acids lacking a capture moiety. Exemplary capture moieties are biotin enabling affinity separation by binding to streptavidin that is bound to or can be bound to a solid phase, or oligonucleotides enabling affinity separation through binding to complementary oligonucleotides that are bound to or can be bound to a solid phase. After the capture portion is attached to the sample nucleic acid, the sample nucleic acid is used as a template for amplification. After amplification, the original template remains attached to the capture portion, but the amplicon is not attached to the capture portion.
[0071] The capture portion can be ligated to the sample nucleic acid as a component of an adapter that can also provide amplification and / or sequencing primer binding sites. In some methods, adapters are ligated to both ends of the sample nucleic acid, so that both adapters have a capture portion. Preferably, any cytosine residue in the adapter is modified, for example, by 5-methylcytosine, to protect it from the action of bisulfites. In some cases, the capture portion is ligated to the original template by a uracil residue that can be cleaved by a cleavable linkage (e.g., photocleavable desthiobiotin-TEG or USER® enzyme, Chem. Commun. (Camb). 2015 Feb 21; 51(15): 3266-3269), in which case the capture portion can be removed if desired.
[0072] The amplicon is denatured and brought into contact with an affinity reagent for the capture tag. The original template binds to the affinity reagent, while the nucleic acid molecule produced by amplification does not. Therefore, the original template can be separated from the nucleic acid molecule produced by amplification.
[0073] After separation or distribution, each population of nucleic acids (i.e., the original template and the amplified product) can be subjected to bisulfite treatment, with the original template population being treated and the amplified product not. Alternatively, the amplified product can be subjected to bisulfite treatment, while the original template population is not. After such treatment, each population can be amplified (in the case of the original template population, uracil is converted to thymine). The populations can also be subjected to biotin probe hybridization for enrichment. Then, each population is analyzed and its sequences are compared to determine which cytosines were 5-methylated (or 5-hydroxymethylated) in the original population. Unmodified C is indicated by the detection of T nucleotides (corresponding to unmethylated cytosines converted to uracil) in the template population and C nucleotides at the corresponding positions in the amplified population. Modified C is indicated in the original sample by the presence of C at the corresponding positions in the original template and the amplified population.
[0074] In some embodiments, the method utilizes sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation of molecularly tagged DNA libraries. This process is carried out by adapter labeling (e.g., biotin), DNA-seq amplification of the entire library, parent molecule recovery (e.g., pull-down with streptavidin beads), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosine at single-nucleotide resolution by sequential NGS-pre-amplification of parent library molecules with and without bisulfite treatment. This can be achieved by modifying one of the two adapter strands of a 5-methylated NGS-adapter (directional adapter; Y-shaped / fork-shaped with 5-methylcytosine replacement) used in BIS-seq with labeling (e.g., biotin). The adapter is ligated to the sample DNA molecule and amplified (e.g., by PCR). Since only the parent molecule will have the labeled adapter end, the parent molecule can be selectively recovered from their amplified progeny by a label-specific capture method (e.g., streptavidin-magnetic beads). Because the parent molecule retains the 5-methylation mark, bisulfite conversion of the captured library yields a single-nucleotide resolution 5-methylation state during BIS-seq, preserving molecular information for the corresponding DNA-seq. In some embodiments, the bisulfite-treated library can be combined with an untreated library by adding a sample-tagged DNA sequence in a standard multiplexed NGS workflow before enrichment / NGS. Bioinformatics analysis can then be performed for genomic alignment and 5-methylated base identification, similar to the BIS-seq workflow. In summary, this method provides the ability to selectively recover the parent, ligated molecule with the 5-methylcytosine mark after library amplification, thereby enabling parallel processing of bisulfite-converted DNA. This overcomes the destructive nature of bisulfite treatment on the quality / sensitivity of DNA-seq information extracted from the workflow.Using this method, the recovered ligated parental DNA molecule (via a labeled adapter) allows for the parallel application of amplification of a complete DNA library and processing to elicit epigenetic DNA modifications. While this disclosure discusses, but is not limited to, the use of the BIS-seq method for identifying cytosine 5-methylated (5-methylcytosine), variations of BIS-seq have been developed for identifying hydroxymethylated cytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxylcytosine. These methodologies can be implemented in conjunction with the sequential / parallel library preparation described herein.
[0075] Alternative methods for modified nucleic acid analysis This disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylation, histone linkage, and other modifications described above). Some such methods involve contacting a population of nucleic acids having varying degrees of modification (e.g., 0, 1, 2, 3, 4, 5, or more methyl groups per nucleic acid molecule) with an adapter, and then fractionating the population according to the degree of modification. The adapter binds to either one or both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags such that the number of tag combinations results in a low probability, e.g., 95, 99, or 99.9%, that two nucleic acids having the same start and end points accept the same tag combination. After binding to the adapter, the nucleic acid is amplified from a primer that binds to a primer binding site in the adapter. The adapter may contain the same or different primer binding sites, whether they have the same or different tags, but preferably the adapter contains the same primer binding site. After amplification, the nucleic acid is contacted with an active agent that binds to nucleic acids, preferably those having modifications (e.g., such active agents described previously). The nucleic acids are separated from binding to the activator into at least two compartments with different degrees of modification. For example, if the activator has affinity for nucleic acids with modifications, nucleic acids with overexpression of the modification (compared to the median expression in the population) will preferentially bind to the activator, while nucleic acids with underexpression of the modification will not bind to the activator or will be more easily eluted from the activator. After separation, the different compartments can then be subjected to further processing steps, which typically include further amplification and sequence analysis, in parallel but separately. The sequence data from the different compartments can then be compared.
[0076] Nucleic acids can be ligated to a Y-shaped adapter containing primer binding sites and a tag at both ends. The molecule is amplified. The amplified molecule is then fractionated by contacting it with an antibody that preferentially binds to 5-methylcytosine, resulting in two compartments. One compartment contains the original molecule lacking methylation and the amplified copy with lost methylation. The other compartment contains the original DNA molecule with methylation. The two compartments are then processed and sequenced separately, with further amplification of the methylated compartment. The sequence data of the two compartments can then be compared. In this example, the tag is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish between different molecules within these compartments, and thus it is possible to determine whether reads with the same start and end points are based on the same molecule or different molecules.
[0077] This disclosure provides further methods for analyzing a population of nucleic acids in which at least a portion of the nucleic acids contain one or more modified cytosine residues, e.g., 5-methylcytosine and any of the other modifications described above. These methods involve contacting the population of nucleic acids with an adapter containing one or more cytosine residues modified at the 5C position, e.g., 5-methylcytosine. Preferably, all cytosine residues of such an adapter are also modified, or all such cytosines within the primer-binding region of the adapter are modified. The adapter binds to both ends of the nucleic acid molecules in the population. Preferably, the adapter contains a sufficient number of different tags such that the number of tag combinations results in a low probability, e.g., 95, 99, or 99.9%, that two nucleic acids having the same start and end points accept the same tag combination. The primer-binding sites of such an adapter may be the same or different, but are preferably the same. After binding to the adapter, the nucleic acids are amplified from a primer that binds to the primer-binding site of the adapter. The amplified nucleic acids are split into a first aliquot and a second aliquot. The first aliquot is assayed for sequence data with or without further processing. Therefore, the sequence data for the molecules in the first aliquot is determined independently of the initial methylation state of the nucleic acid molecules. The nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosine to uracil. The bisulfite-treated nucleic acids are then subjected to amplification, primed with a primer to the original primer-binding site of the adapter attached to the nucleic acid. At this point, only the nucleic acid molecules originally attached to the adapter (separate from the amplified product) are amplified because these nucleic acids retain cytosine at the primer-binding site of the adapter, while the amplified product loses methylation of the cytosine residue that was converted to uracil by bisulfite treatment. Therefore, only the original molecules in the population, which are at least partially methylated, are amplified. After amplification, these nucleic acids are subjected to sequence analysis.By comparing the sequence determined from the first aliquot with the sequence determined from the second aliquot, it may be possible to indicate, in particular, which cytosines within the nucleic acid population were subjected to methylation.
[0078] Distribution of a sample into multiple subsamples; sample morphology; analysis of epigenetic features. In certain embodiments described herein, a population of different forms of nucleic acids (e.g., a captured set of highly methylated and hypomethylated DNA in a sample, e.g., cfDNA as described herein) can be physically distributed based on one or more characteristics of the nucleic acids before further analysis, e.g., differential modification or isolation of nucleic acid bases, tagging, and / or sequencing. This approach can be used, for example, to determine whether a particular sequence is highly methylated or hypomethylated. In some embodiments, highly methylated variable epigenetic target regions are analyzed to determine whether they exhibit the highly methylated characteristics of tumor cells, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit the hypomethylated characteristics of tumor cells. Furthermore, by distributing a heterogeneous population of nucleic acids, rare signals can be increased, for example, by enriching rare nucleic acid molecules that are more dominant in one fraction (or compartment) of the population. For example, genetic variations present in highly methylated DNA but less so (or absent) in less methylated DNA can be more easily detected by distributing the sample into highly methylated and less methylated nucleic acid molecules. By analyzing multiple fractions of the sample, multidimensional analysis of a single locus of the genome or a nucleic acid species can be performed, thus achieving higher sensitivity.
[0079] In some cases, heterogeneous nucleic acid samples are distributed into two or more compartments (e.g., at least three, four, five, six, or seven compartments). In some embodiments, each distribution is differentially tagged. The tagged distributions can then be pooled together for collective sample preparation and / or sequencing. The distribution-tagging-pooling steps may be performed more than once, with each distribution being based on different characteristics (examples are provided herein) and tagged using differential tags that identify it with other distributions and distribution means.
[0080] Examples of features that can be used for distribution include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting distribution may include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, distribution based on cytosine modifications (e.g., cytosine methylation) or methylation is commonly performed and, if necessary, combined with at least one additional distribution step which may be based on any of the aforementioned DNA features or forms. In some embodiments, a heterogeneous population of nucleic acids is distributed into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine versus other types of methylation, e.g., adenine methylation and / or cytosine hydroxymethylation), as well as association with one or more proteins, e.g., histones, and the level of association. Alternatively, the heterogeneous nucleic acid population may be distributed between nucleic acid molecules associated with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively, the heterogeneous nucleic acid population may be distributed between single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively, the heterogeneous nucleic acid population may be distributed based on the length of the nucleic acid (e.g., molecules up to 160 bp and molecules longer than 160 bp).
[0081] In some cases, each compartment (representing a different nucleic acid morphology) is differentially labeled before sequencing, and the compartments are pooled together. In other cases, the different morphologies are sequenced separately. In some embodiments, a population of different nucleic acids is distributed into two or more different compartments. Each compartment represents a different nucleic acid morphology, and the first compartment (also referred to as a subsample) contains DNA with a higher proportion of cytosine modifications than the second subsample. Each compartment is tagged separately. The first subsample is subjected to a procedure that affects a first nucleic acid base in the DNA of the first subsample differently from a second nucleic acid base in the DNA, where the first nucleic acid base is modified or unmodified, and the second nucleic acid base is a modified or unmodified nucleic acid base different from the first, and the first and second nucleic acid bases have the same base-pairing specificity. The tagged nucleic acids are pooled together before sequencing. Sequence reads are obtained, and analysis is performed in silico, including distinguishing the first nucleic acid base from the second nucleic acid base in the DNA of the first partial sample. Tags can be used to sort reads from different distributions. Analysis for detecting genetic variants can be performed at the distribution level as well as at the whole nucleic acid population level. For example, the analysis may include in silico analysis to determine genetic variants in nucleic acids in each distribution, such as CNVs, SNVs, indels, and fusions. In some cases, in silico analysis may include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the location of nucleosomes in chromatin. Higher coverage may correlate with higher nucleosome occupancy in genomic regions, while lower coverage may correlate with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0082] The sample may contain nucleic acids that have undergone modifications, including post-replication modifications to nucleotides, and that have been altered in their typically non-covalent binding to one or more proteins.
[0083] In one embodiment, the nucleic acid population is obtained from serum, plasma, or blood samples from subjects suspected of having a neoplasm, tumor, or cancer, or who have been previously diagnosed with a neoplasm, tumor, or cancer. The nucleic acid population includes nucleic acids with varying levels of methylation. Methylation can result from any one or more post-replication or post-transcriptional modifications. Post-replication modifications include modifications of nucleotide cytosines, particularly those at the 5-position of the nucleic acid base, e.g., 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine. The affinity agent may be an antibody with desired specificity, its natural binding partner or variant (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or an artificial peptide selected, for example, by phage display to have specificity for a given target.
[0084] Examples of the capture portions intended herein include the methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) described herein, including proteins such as antibodies that preferentially bind to MeCP2 and 5-methylcytosine. Similarly, the distribution of different forms of nucleic acids can be carried out using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. With respect to some affinity agents and modifiers, binding to the active ingredient may occur in an essentially present or absent manner, depending on whether the nucleic acid has the modifier, while separation may be to a degree. In such cases, nucleic acids with over-presented modifiers will bind to the active ingredient to a greater extent than nucleic acids with under-presented modifiers. Alternatively, nucleic acids with modifiers may bind in an present or absent manner. Nevertheless, various levels of modifiers can be sequentially eluted from the binder.
[0085] For example, in some embodiments, the partitioning may be binary or based on the degree / level of modification. For instance, all methylated fragments can be partitioned from unmethylated fragments using a methyl-binding domain protein (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Subsequent partitioning may involve eluting fragments with different levels of methylation by adjusting the salt concentration in the solution containing the methyl-binding domain and the binding fragments. As the salt concentration increases, fragments with higher levels of methylation are eluted. In some cases, the final partitions represent nucleic acids with different degrees of modification (over-expression or under-expression of the modification). Over-expression and under-expression may be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in nucleic acids in a sample is 2, then nucleic acids containing more than 2 5-methylcytosine residues are over-expressing this modification, and nucleic acids with 1 or 0 5-methylcytosine residues are under-expressing. The effect of affinity separation is to enrich nucleic acids that are overexpressing the modification in the conjugated phase and nucleic acids that are underexpressing the modification in the unconjugated phase (i.e., in solution). Nucleic acids in the conjugated phase can be eluted before subsequent processing.
[0086] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be partitioned using sequential elution. For example, a low-methylation section (e.g., no methylation) can be separated from the methylated section by contacting the nucleic acid population with MBD from the kit, which is attached to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. Then, one or more elution steps are performed sequentially to elute nucleic acids with different levels of methylation. For example, a first set of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, e.g., at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After eluting such methylated nucleic acids, magnetic separation is used again to separate nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps themselves can be repeated to create various compartments, such as a low-methylation compartment (representing no methylation), a methylated compartment (representing low levels of methylation), and a high-methylation compartment (representing high levels of methylation).
[0087] In some methods, nucleic acids bound to the active ingredient used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can enrich nucleic acids with a degree of modification close to the mean or median (i.e., an intermediate value between nucleic acids that remained bound to the solid phase when the sample was first brought into contact with the active ingredient and nucleic acids that were not bound to the solid phase). Affinity separation results in at least two, sometimes three or more, partitions of nucleic acids with different degrees of modification. The partitions, still distinct, ligate the nucleic acids of at least one partition, and usually two or three (or more) partitions, to nucleic acid tags, usually provided as components of an adapter, and nucleic acids in different partitions receive different tags that identify one partition member as another. Tags ligated to nucleic acid molecules of the same partition may be the same or different from each other. However, if they are different from each other, the tags may have a common coding portion to identify the molecule to which they are attached as belonging to a particular partition. For further details regarding the portioning of nucleic acid samples based on characteristics such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some embodiments, nucleic acid molecules may be fractionated into different portions based on which nucleic acid molecules are bound to a particular protein or fragment and which are not bound to that particular protein or fragment.
[0088] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activity. Examples of proteins that can bind to DNA and function as criteria for fractionation include, but are not limited to, protein A and protein G. Nucleic acid molecules can be fractionated based on protein-bound regions using any preferred method. Examples of methods used to fractionate nucleic acid molecules based on protein-bound regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric flow-field separation (AF4).
[0089] In some embodiments, nucleic acid distribution is carried out by contacting the nucleic acid with the methylation-binding domain ("MBD") of a methylation-binding protein ("MBP"). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads such as Dynabeads® M-280 streptavidin via a biotin linker. Distribution to fractions with different degrees of methylation can be carried out by eluting the fractions by increasing the NaCl concentration.
[0090] An exemplary method for identifying molecular tags in a library distributed by MBD-beads using NGS is as follows:
[0091] Physical distribution of extracted DNA samples (e.g., plasma DNA extracted from human samples) using a methyl-binding domain protein-bead purification kit. All eluates from the process are stored for downstream processing.
[0092] Parallel application of differential molecular tags and adapter sequences enabling NGS to each compartment. For example, ligating hypermethylated, residual methylated ("washed"), and hypomethylated compartments with NGS-adapters having molecular tags.
[0093] All molecularly tagged compartments are reassembled and then amplified using adapter-specific DNA primer sequences.
[0094] Enrichment / hybridization of the total library, which has been combined and amplified again. Targeting the desired genomic region (e.g., cancer-specific genetic variants and differentially methylated regions).
[0095] Re-amplification of enriched total DNA libraries, and addition of sample tags. Different samples are pooled and subjected to multiple assays using NGS instruments.
[0096] Bioinformatics analysis of NGS data. This involves identifying unique molecules using molecular tags, and similarly deconvolving samples into differentially MBD-distributed molecules. This analysis allows for obtaining relative information about 5-methylcytosine for genomic regions simultaneously with standard gene sequencing / variant detection.
[0097] Examples of MBPs intended in this specification include, but are not limited to, the following: (a) MeCP2 is a protein that preferentially binds to 5-methylcytosine rather than unmodified cytosine. (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethylcytosine rather than unmodified cytosine. (c)FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 preferably bind to 5-formyl-cytosine more than unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)). (d) An antibody specific to one or more methylated nucleotide bases.
[0098] Generally, elution is a function of the number of methylation sites per molecule, with molecules having more methylation elute at higher salt concentrations. To elute DNA into separate populations based on the degree of methylation, a series of elution buffers with progressively increasing NaCl concentrations can be used. The salt concentration can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process yields three partitions. Molecules may be brought into contact with a solution containing molecules with methyl-binding domains at a first salt concentration, causing the molecules to bind to a capture site, such as streptavidin. At the first salt concentration, one population of molecules binds to the MBD, while another population remains unbound. The unbound population can be separated as a "low-methylated" population. For example, the first compartment representing the low-methylated form of DNA remains unbound at low salt concentrations, e.g., 100 mM or 160 mM. The second compartment, representing intermediate methylated DNA, is eluted using an intermediate salt concentration, e.g., between 100 mM and 2000 mM. This compartment is also separated from the sample. The third compartment, representing the highly methylated form of DNA, is eluted using a high salt concentration, e.g., at least about 2000 mM.
[0099] This disclosure provides further methods for analyzing a population of nucleic acids in which at least a portion of the nucleic acids contains one or more modified cytosine residues, e.g., 5-methylcytosine and any of the other modifications described above. In these methods, after distribution, a portion of the nucleic acid sample is brought into contact with an adapter containing one or more cytosine residues modified at the 5C position, e.g., 5-methylcytosine. Preferably, all cytosine residues of such an adapter are also modified, or all such cytosines within the primer-binding region of the adapter are modified. The adapter binds to both ends of the nucleic acid molecule in the population. Preferably, the adapter contains a sufficient number of different tags such that the number of tag combinations results in a low probability, e.g., 95, 99, or 99.9%, that two nucleic acids having the same start and end points accept the same tag combination. The primer-binding sites of such an adapter may be the same or different, but preferably they are the same. After binding to the adapter, the nucleic acid is amplified from a primer that binds to the primer-binding site of the adapter. The amplified nucleic acid is split into a first aliquot and a second aliquot. The first aliquot is assayed for sequence data with or without further processing. Thus, the sequence data for the molecule in the first aliquot is determined independently of the initial methylation state of the nucleic acid molecule. The nucleic acid molecule in the second aliquot is subjected to a procedure that affects the first nucleic acid base in the DNA differently than the second nucleic acid base in the DNA, where the first nucleic acid base contains cytosine modified at position 5 and the second nucleic acid base contains unmodified cytosine. This procedure may be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. The nucleic acid subjected to the procedure is then amplified using a primer to the original primer binding site of the adapter ligated to the nucleic acid. At this point, only the nucleic acid molecule originally ligated to the adapter (separate from its amplified product) is amplified because these nucleic acids retain cytosine at the primer binding site of the adapter, while the amplified product loses methylation of the cytosine residue that was converted to uracil by bisulfite treatment.Therefore, only the original molecules in the population that are at least partially methylated undergo amplification. After amplification, these nucleic acids are subjected to sequence analysis. By comparing the sequence determined from the first aliquot with the sequence determined from the second aliquot, it may be possible to indicate, in particular, which cytosines in the nucleic acid population were subjected to methylation.
[0100] Such analysis can be performed using the following exemplary procedure. After distribution, the methylated DNA is ligated at both ends to a Y-shaped adapter containing a primer binding site and a tag. The cytosine in the adapter is modified at position 5 (e.g., 5-methylation). The modification of the adapter serves to protect the primer binding site in subsequent conversion steps (e.g., bisulfite treatment, TAP conversion, or any other conversion that does not affect the modified cytosine but does affect the unmodified cytosine). After ligation of the adapter, the DNA molecule is amplified. The amplified product is split into two aliquots for sequencing with and without conversion. The aliquot not subjected to conversion can be subjected to sequencing analysis with or without further processing. The other aliquot is subjected to a procedure that affects the first nucleic acid base in the DNA differently from the second nucleic acid base in the DNA, where the first nucleic acid base contains the cytosine modified at position 5 and the second nucleic acid base contains the unmodified cytosine. This procedure may be bisulfite treatment or another procedure to convert unmodified cytosine to uracil. When contacted with a primer specific to the original primer binding site, only the primer binding site protected by the cytosine modification can support amplification. Therefore, only the original molecule is subjected to further amplification, and the copy from the first amplification is not subjected to further amplification. The further amplified molecule is then subjected to sequence analysis. The sequences from the two aliquots can then be compared. Similar to the separation scheme described above, the nucleic acid tag on the adapter is used to distinguish nucleic acid molecules within the same compartment, rather than to distinguish between methylated and unmethylated DNA.
[0101] A step of subjecting a first partial sample to a procedure that affects the first nucleic acid base in the DNA of the first partial sample differently from the second nucleic acid base in the DNA. A method disclosed herein is a step of subjecting a first partial sample to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity. In some embodiments, if the first nucleic acid base is modified or unmodified adenine, then the second nucleic acid base is modified or unmodified adenine; if the first nucleic acid base is modified or unmodified cytosine, then the second nucleic acid base is modified or unmodified cytosine; if the first nucleic acid base is modified or unmodified guanine, then the second nucleic acid base is modified or unmodified guanine; and if the first nucleic acid base is modified or unmodified thymine, then the second nucleic acid base is modified or unmodified thymine (for the purposes of this step, modified and unmodified uracil are encompassed under modified thymine).
[0102] In some embodiments, the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine. For example, the first nucleic acid base may contain unmodified cytosine (C), and the second nucleic acid base may contain one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleic acid base may contain C, and the first nucleic acid base may contain one or more of mC and hmC. Other combinations are also possible, for example, when one of the first and second nucleic acid bases contains mC and the other contains hmC, as shown in the above summary of the invention and the following discussion.
[0103] In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA involves bisulfite conversion. Treatment with bisulfite converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxylcytosine (caC)) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Therefore, when using bisulfite conversion, the first nucleic acid base may include one or more of unmodified cytosine, 5-formylcytosine, 5-carboxylcytosine, or other cytosine forms affected by bisulfite, and the second nucleic acid base may include one or more of mC and hmC, e.g., mC and optionally hmC. Sequencing of the bisulfite-treated DNA identifies the position read as cytosine as an mC or hmC position. On the other hand, positions read as T are identified as T, or bisulfite-sensitive forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxylcytosine. Therefore, by performing bisulfite conversion on the first partial sample described herein, it becomes easier to identify mC or hmC-containing positions using sequence reads obtained from the first partial sample. For an illustrative description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.
[0104] In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes oxidative bisulfite (Ox-BS) conversion. In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes Tet-assisted bisulfite (TAB) conversion. In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes Tet-assisted conversion using a substituted borane reducing agent, which may optionally be 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In some embodiments, a procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes chemically assisted conversion using a substituted borane reducing agent, which may optionally be 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the procedure that affects a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes APOBEC coupling epigenetic (ACE) conversion.
[0105] In some embodiments, the procedure for affecting a first nucleic acid base in the DNA of a first partial sample differently from a second nucleic acid base in the DNA includes, for example, enzymatic conversion of the first nucleic acid base, similar to EM-Seq. See, for example, Vaisvila R, et al. (2019) EM-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by deaminase (e.g., APOBEC3A), and then the unmodified cytosines can be deaminated using deaminase (e.g., APOBEC3A) to convert them to uracil.
[0106] In some embodiments, the procedure for affecting a first nucleic acid base in the DNA of a first partial sample differently from that affecting a second nucleic acid base in the DNA includes separating the DNA that originally contains the first nucleic acid base from the DNA that originally does not contain the first nucleic acid base.
[0107] In some embodiments, the first nucleic acid base is modified or unmodified adenine, and the second nucleic acid base is modified or unmodified adenine. In some embodiments, the modified adenine is N 6 -methyl adenine (mA). In some embodiments, the modified adenine is N 6 -Methyladenine (mA), N 6 -Hydroxymethyladenine (hmA), or N 6 - One or more of the formyladenine (fA) compounds.
[0108] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to isolate DNA containing modified bases such as mA from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37:1155-62. Antibodies against various modified nucleic acid bases, including halogenated forms such as 5-bromouracil, and thymine / uracil forms are commercially available. Various modified bases can also be detected based on changes in their base-pairing specificity. For example, hypoxanthine is a modified form of adenine that can result from deamination and is read as G in sequencing. For example, see U.S. Patent 8,486,630; Brown, Genomes, 2nd Ed., John Wiley & Sons, Inc., New York, NY, 2002, chapter 14, "Mutation, Repair, and Recombination."
[0109] Enrichment / capture steps, amplification, adapter, barcode In some embodiments, the methods disclosed herein include the step of capturing one or more sets of target regions of DNA, such as cfDNA. Capture can be carried out using any preferred approach known in the Art. In some embodiments, the capture step includes contacting the DNA to be captured with a set of target-specific probes. The set of target-specific probes may have any of the features described herein with respect to a set of target-specific probes, including, but not limited to, those in the embodiments described above and in the sections relating to probes below. The capture step can be carried out on one or more partial samples prepared during the methods disclosed herein. In some embodiments, DNA is captured from at least a first or second partial sample, for example, at least a first and a second partial sample. When a first partial sample is subjected to a separation step (for example, separating DNA that originally contains a first nucleic acid base (e.g., hmC) from DNA that originally does not contain the first nucleic acid base, e.g., an hmC-seal), the capture step can be performed on any, any two, or all of the DNA that originally contains the first nucleic acid base (e.g., hmC), the DNA that originally does not contain the first nucleic acid base, and the second partial sample. In some embodiments, the partial samples are differentially tagged (e.g., as described herein), then pooled, and then subjected to capture.
[0110] The capture step can be carried out using conditions suitable for specific nucleic acid hybridization, which generally depend to some extent on the characteristics of the probe, such as length and base composition. Those skilled in the art are familiar with suitable conditions, taking into account the general knowledge in the art regarding nucleic acid hybridization. In some embodiments, a complex is formed between the target-specific probe and DNA.
[0111] In some embodiments, the method described herein includes a step of capturing cfDNA obtained from a test subject for multiple sets of target regions. The target regions include epigenetic target regions, which may exhibit differences in methylation levels and / or fragmentation patterns depending on whether they originate from a tumor or healthy cells. The target regions also include sequence-variable target regions, which may exhibit differences in sequence depending on whether they originate from a tumor or healthy cells. The capture step produces a captured set of cfDNA molecules, in which cfDNA molecules corresponding to the sequence-variable target region set are captured in a higher capture yield than cfDNA molecules corresponding to the epigenetic target region set. For further consideration of the capture step, capture yields, and related embodiments, see WO2020 / 160414, which is incorporated herein by reference for any purpose.
[0112] In some embodiments, the method described herein includes the step of contacting cfDNA obtained from a test subject with a set of target-specific probes, wherein the set of target-specific probes is configured to capture cfDNA corresponding to a sequence-variable target region set with a higher capture yield than cfDNA corresponding to an epigenetic target region set.
[0113] To analyze sequence-variable target regions with sufficient confidence or accuracy, sequencing to a greater depth may be required than may be necessary to analyze epigenetic target regions. Therefore, capturing cfDNA corresponding to a set of sequence-variable target regions with a higher capture yield than cfDNA corresponding to a set of epigenetic target regions may be beneficial. The amount of data required to determine fragmentation patterns (e.g., to test for perturbations at transcription start sites or CTCF binding sites) or the abundance of fragments (e.g., in highly methylated and hypomethylated compartments) is generally less than the amount of data required to determine the presence or absence of cancer-associated sequence mutations. By capturing target region sets with different yields, it may be easier to sequence target regions to different depths of sequencing in the same sequencing trial (e.g., using a pooled mixture and / or in the same sequencing cell).
[0114] In various embodiments, the method further includes a step of sequencing the captured cfDNA to varying degrees of sequencing depth, for example, with respect to epigenetic and sequence-variable target region sets, consistent with the discussion herein. In some embodiments, the target-specific probe-DNA complex is separated from DNA not bound to the target-specific probe. For example, if the target-specific probe is covalently or noncovalently bound to a solid support, washing or aspiration steps can be used to separate the unbound material. Alternatively, chromatography can be used if the complex has different chromatographic properties from the unbound material (for example, if the probe contains a ligand that binds to a chromatography resin).
[0115] As discussed in detail elsewhere in this specification, a set of target-specific probes may include multiple sets, for example, probes for a set of sequence-variable target regions and probes for a set of epigenetic target regions. In some such embodiments, the capture step is performed simultaneously in the same container using probes for the sequence-variable target regions and probes for the epigenetic target regions, for example, the probes for the sequence-variable target regions and probes for the epigenetic target regions are in the same composition. This approach results in a relatively streamlined workflow. In some embodiments, the concentration of probes for the sequence-variable target regions is higher than the concentration of probes for the epigenetic target regions.
[0116] Alternatively, the capture step may be performed using a sequence-variable target region probe set in a first container and an epigenetic target region probe set in a second container, or the contact step may be performed using the sequence-variable target region probe set for a first time and in the first container, and the epigenetic target region probe set for a second time before or after the first time. This approach makes it possible to prepare a first composition containing captured DNA corresponding to a separate sequence-variable target region set and a second composition containing captured DNA corresponding to an epigenetic target region set. The compositions may be processed separately as desired (e.g., to fractionate based on methylation as described elsewhere in this specification) and then recombined in proportions suitable for further processing and analysis, such as for sequencing.
[0117] In some embodiments, DNA is amplified. In some embodiments, amplification is performed before the capture step. In some embodiments, amplification is performed after the capture step.
[0118] In some embodiments, the adapter is included in the DNA. This can be done in conjunction with the amplification procedure, for example, by providing the adapter to the 5' portion of the primer, as described above. Alternatively, the adapter can be added by other approaches such as ligation.
[0119] In some embodiments, the DNA may be accompanied by a tag that is either a barcode or contains a barcode. The tag can facilitate the identification of the nucleic acid's origin. For example, after pooling multiple samples for parallel sequencing, the barcode may be used to identify the origin (e.g., target) from which the DNA originates. This can be done concurrently with the amplification procedure, for example, by providing the barcode on the 5' portion of the primer, as described above. In some embodiments, the adapter and tag / barcode are provided by the same primer or primer set. For example, the barcode may be positioned at the 3' of the adapter and the 5' of the portion that hybridizes to the target of the primer. Alternatively, the barcode may be added together with the adapter, if necessary, on the same ligation substrate, by other approaches, such as ligation.
[0120] Further details regarding amplification, tagging, and barcodes are discussed below in the section “General Features of the Method,” which can be put into practice to an operational degree with any of the embodiments described above as well as in the Introduction and Summary sections.
[0121] Captured set In some embodiments, a captured set of DNA (e.g., cfDNA) is provided. With respect to the disclosed method, the captured set of DNA can be obtained, for example, by performing a capture step after a distribution step, as described herein. The captured set may include DNA corresponding to a sequence variable target region set, DNA corresponding to an epigenetic target region set, or a combination thereof. In some embodiments, the quantity of captured sequence variable target region DNA is greater than the quantity of captured epigenetic target region DNA, after normalization for differences in targeting region size (footprint size).
[0122] Alternatively, a first captured set and a second captured set can be obtained, each containing DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set, respectively. The first captured set and the second captured set can be combined to obtain a combined captured set.
[0123] In some embodiments, the captured set, which includes DNA corresponding to the sequence variable target region set and DNA corresponding to the epigenetic target region set, contains captured sets of the combinations discussed above, in which case the DNA corresponding to the sequence variable target region set is at a higher concentration than the DNA corresponding to the epigenetic target region set, for example, 1.1 to 1.2 times higher, 1.2 to 1.4 times higher, 1.4 to 1.6 times higher, 1 0.6 to 1.8 times higher concentration, 1.8 to 2.0 times higher concentration, 2.0 to 2.2 times higher concentration, 2.2 to 2.4 times higher concentration, 2.4 to 2.6 times higher concentration, 2.6 to 2.8 times higher concentration, 2.8 to 3.0 times higher concentration, 3.0 to 3.5 times higher concentration, 3.5 to 4.0 times, 4.0 to 4.5 times higher concentration, 4.5 to 5.0 times higher concentration, 5.0 to 5.5 times higher concentration, 5.5 to 6.0 times higher concentration, 6.0 to 6.5 times higher concentration, 6. 5 to 7.0 times higher concentration, 7.0 to 7.5 times higher concentration, 7.5 to 8.0 times higher concentration, 8.0 to 8.5 times higher concentration, 8.5 to 9.0 times higher concentration, 9.0 to 9.5 times higher concentration, 9.5 to 10.0 times higher concentration, 10 to 11 times higher concentration, 11 to 12 times higher concentration, 12 to 13 times higher concentration, 13 to 14 times higher concentration, 14 to 15 times higher concentration, 15 to 16 times higher concentration, 16 to 17 times higher concentration, 17 to 18 times higher concentration It can exist at concentrations 18 to 19 times higher, 19 to 20 times higher, 20 to 30 times higher, 30 to 40 times higher, 40 to 50 times higher, 50 to 60 times higher, 60 to 70 times higher, 70 to 80 times higher, 80 to 90 times higher, 90 to 100 times higher, 10 to 20 times higher, 10 to 40 times higher, 10 to 50 times higher, 10 to 70 times higher, or 10 to 100 times higher. Depending on the degree of the concentration difference, the normalization with respect to the footprint size of the target region is explained, as discussed in the definition section.
[0124] Epigenetic target region set An epigenetic target region set may include one or more types of target regions that can distinguish DNA from that derived from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, e.g., non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. An epigenetic target region set may also include, for example, one or more control regions described herein. In some embodiments, an epigenetic target region set has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target region set has footprints ranging from 100 to 1000 kb, for example, 100 to 200 kb, 200 to 300 kb, 300 to 400 kb, 400 to 500 kb, 500 to 600 kb, 600 to 700 kb, 700 to 800 kb, 800 to 900 kb, and 900 to 1,000 kb.
[0125] Highly methylated variable target region In some embodiments, the epigenetic target region set includes one or more hypermethylated variable target regions. Generally, a hypermethylated variable target region refers to a region where, for example in a cfDNA sample, the observed elevated level of methylation indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by neoplastic cells such as tumor or cancer cells. See, for example, Kang et al, Genome Biol. 18: 53 (2017) and the references cited therein. In one example, a hypermethylated variable target region may include a region in cancerous tissue that does not necessarily have different methylation compared to DNA from the same type of healthy tissue, but has different methylation (e.g., more methylation) compared to cfDNA that is typical in healthy subjects. For example, if the presence of cancer leads to increased cell death, e.g., increased apoptosis of cells in the tissue type corresponding to cancer, such cancer can be detected, at least partially, using such hypermethylated variable target regions. In some embodiments, the hypermethylated variable target regions include one or more genomic regions in which cfDNA molecules in those regions have no different methylation status in cancer subjects compared to cfDNA derived from healthy subjects, but the presence / increased quantity of hypermethylated cfDNA in those regions indicates a specific tissue type (e.g., cancer origin) and is shown as cfDNA with increased apoptosis (e.g., tumor shedding) in circulation.
[0126] Highly methylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a stochastic method called CancerLocator using highly methylated target regions derived from breast, colon, kidney, liver, and lung. In some embodiments, highly methylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the highly methylated target regions include one, two, three, four, or five subsets of highly methylated target regions that collectively exhibit high methylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0127] In some embodiments, a probe for a set of epigenetic target regions includes a probe specific to one or more highly methylated variable target regions. The highly methylated variable target regions may be any of the above. For example, in some embodiments, a probe specific to a highly methylated variable target region includes a plurality of loci listed in Table 1, e.g., a probe specific to at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1. In some embodiments, a probe specific to a highly methylated variable target region includes a plurality of loci listed in Table 2, e.g., a probe specific to at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2. In some embodiments, probes specific to highly methylated variable target regions include probes specific to a plurality of loci listed in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2. In some embodiments, for each locus included as a target region, there may be one or more probes having a hybridization site that binds between the transcription start site and the stop codon (or the last stop codon in the case of a gene that is alternatively spliced). In some embodiments, one or more probes bind within 300 bp, e.g., 200 or 100 bp, of the listed locations. In some embodiments, the probes have hybridization sites that overlap with the locations listed above. In some embodiments, probes specific to hypermethylated target regions include probes specific to one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0128] Low-methylation variable target region Overall hypomethylation is a phenomenon commonly observed in various cancers. See, for example, Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (a review article mentioning the observation of hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions such as repeating elements, e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA, as well as intergenetic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Therefore, in some embodiments, a set of epigenetic target regions may include hypomethylated variable target regions, where the reduced level of observed methylation indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by neoplastic cells such as tumor or cancer cells. For example, a hypomethylated variable target region may include regions in cancerous tissue that are not necessarily methylated differently from DNA derived from the same type of healthy tissue, but are methylated differently (e.g., less methylated) than cfDNA that is typical in healthy subjects. For example, if the presence of cancer leads to increased cell death, e.g., apoptosis of cells of the tissue type corresponding to cancer, such cancer can be detected, at least partially, using such a hypomethylated variable target region. In some embodiments, a hypomethylated variable target region includes one or more genomic regions in which cfDNA molecules in those regions are not methylated differently in cancer subjects compared to cfDNA derived from healthy subjects, but the presence / increased quantity of hypomethylated cfDNA in those regions indicates a specific tissue type (e.g., cancer origin) and is indicated as cfDNA with increased apoptosis into circulation (e.g., tumor shedding).
[0129] In some embodiments, the low-methylation variable target region includes repeat elements and / or intergenetic regions. In some embodiments, the repeat elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and / or satellite DNA.
[0130] Exemplary specific genomic regions exhibiting cancer-related hypomethylation include nucleotides 8403565–8953708 and 151104701–151106035 on human chromosome 1. In some embodiments, the hypomethylation variable target region overlaps with or includes one or both of these regions.
[0131] In some embodiments, a probe for a set of epigenetic target regions includes a probe specific to one or more hypomethylated variable target regions. The hypomethylated variable target regions may be any of the above. For example, a probe specific to one or more hypomethylated variable target regions may include probes for repeating elements, such as LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA, where intergenetic regions that are normally methylated in healthy cells may show reduced methylation in tumor cells.
[0132] In some embodiments, probes specific to low-methylation variable target regions include probes specific to repeat elements and / or intergenetic regions. In some embodiments, probes specific to repeat elements include probes specific to one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and / or satellite DNA.
[0133] Exemplary probes specific to genomic regions exhibiting cancer-related hypomethylation include probes specific to human chromosome 1 nucleotides 8403565–8953708 and / or 151104701–151106035. In some embodiments, probes specific to hypomethylation variable target regions include probes specific to regions overlapping with or containing human chromosome 1 nucleotides 8403565–8953708 and / or 151104701–151106035.
[0134] Probes for detecting a panel of regions include those for detecting target genomic regions (hotspot regions), as well as nucleosome recognition probes (e.g., KRAS codons 12 and 13). Furthermore, probes for detecting a panel of regions can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. The regions used herein may also include non-hotspot regions optimized based on nucleosome location and GC model.
[0135] In some embodiments, DNA (e.g., cfDNA) is obtained from subjects having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects suspected of having cancer. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects suspected of having a tumor. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects having a neoplasm. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects suspected of having a neoplasm. In some embodiments, DNA (e.g., cfDNA) is obtained from subjects who have achieved remission from a tumor, cancer, or neoplasm (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the embodiments described above, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, may be of the lung. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the colon or rectum. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the breast. In some embodiments, the cancer, tumor, or neoplasm, or suspected cancer, tumor, or neoplasm, is of the prostate. In any of the embodiments described above, the subject may be a human subject.
[0136] In some embodiments, the sequence-variable target region probe set has a footprint of at least 0.5kb, for example, at least 1kb, at least 2kb, at least 5kb, at least 10kb, at least 20kb, at least 30kb, or at least 40kb. In some embodiments, the epigenetic target region probe set has a footprint ranging from 0.5 to 100kb, for example, in the range of 0.5 to 2kb, 2 to 10kb, 10 to 20kb, 20 to 30kb, 30 to 40kb, 40 to 50kb, 50 to 60kb, 60 to 70kb, 70 to 80kb, 80 to 90kb, and 90 to 100kb.
[0137] In some embodiments, probes specific to a sequence-variable target region set include probes specific to target regions from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0138] composition containing captured DNA A combination comprising a first population and a second population of captured DNA is provided herein. The first population may comprise or be obtained from DNA having a higher proportion of cytosine modifications than the second population. The first population may comprise a first nucleic acid base form originally present in the DNA with altered base-pairing specificity and a second nucleic acid base with unaltered base-pairing specificity, where the first nucleic acid base form originally present in the DNA before alteration of base-pairing specificity is a modified or unaltered nucleic acid base, and the second nucleic acid base is a modified or unaltered nucleic acid base different from the first nucleic acid base, and the first nucleic acid base form originally present in the DNA before alteration of base-pairing specificity and the second nucleic acid base have the same base-pairing specificity. The second population does not comprise a first nucleic acid base form originally present in the DNA with altered base-pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleic acid base is modified or unaltered cytosine, and the second nucleic acid base is modified or unaltered cytosine. The first nucleic acid base and the second nucleic acid base may be any of those discussed herein, in the summary of the invention, or in relation to the step of subjecting a first partial sample to a procedure that affects the first nucleic acid base in the DNA of the first partial sample differently from the second nucleic acid base in the DNA.
[0139] In some embodiments, the first group includes array tags selected from a first set of one or more array tags, and the second group includes array tags selected from a second set of one or more array tags, wherein the second set of array tags is different from the first set of array tags. The array tags may include barcodes.
[0140] In some embodiments, the first population includes protected hmC, e.g., glucosylated hmC. In some embodiments, the first population has been subjected to one of the conversion procedures discussed herein, e.g., bisulfite conversion, Ox-BS conversion, TAB conversion, ACE conversion, TAP conversion, TAPSβ conversion, or CAP conversion. In some embodiments, the first population has been subjected to protection of hmC, followed by deamination of mC and / or C. In some embodiments of the combination, the first population includes or is derived from DNA having a higher proportion of cytosine modifications than the second population, and the first population includes a first subpopulation and a second subpopulation, where the first nucleic acid bases are modified or unmodified nucleic acid bases, the second nucleic acid bases are modified or unmodified nucleic acid bases different from the first nucleic acid bases, and the first and second nucleic acid bases have the same base-pairing specificity. In some embodiments, the second population does not include the first nucleic acid bases. In some embodiments, the first nucleic acid base is modified or unmodified cytosine, and the second nucleic acid base is modified or unmodified cytosine, and optionally the modified cytosine is mC or hmC. In some embodiments, the first nucleic acid base is modified or unmodified adenine, and the second nucleic acid base is modified or unmodified adenine, and optionally the modified adenine is mA.
[0141] In some embodiments, the first nucleic acid base (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleic acid base (e.g., modified cytosine) is the product of hysgen cycloaddition to β-6-azido-glucosyl-5-hydroxymethylcytosine, which contains an affinity label (e.g., biotin).
[0142] In any of the combinations described herein, the captured DNA may include cfDNA. The captured DNA may have any of the characteristics described herein with respect to the captured set, including, for example, that the concentration of DNA corresponding to the sequence variable target region set (normalized with respect to footprint size as described above) is higher than that of DNA corresponding to the epigenetic target region set. In some embodiments, the DNA of the captured set includes a sequence tag, which can be attached to the DNA as described herein. Generally, including a sequence tag results in a DNA molecule that differs from the naturally occurring, untagged form.
[0143] The combination may further include the probe sets or sequencing primers described herein, each of which may differ from naturally occurring nucleic acid molecules. For example, the probe sets described herein may include a capture portion, and the sequencing primers may include labels that do not exist in nature.
[0144] Computer system The methods of the present disclosure can be implemented using or with the assistance of a computer system. For example, such a method may include the steps of: distributing a sample into a plurality of partial samples, including a first partial sample and a second partial sample, wherein the first partial sample contains a higher proportion of DNA having cytosine modifications than the second partial sample; subjecting the first partial sample to a procedure that affects a first nucleic acid base in the DNA of the first partial sample differently from a second nucleic acid base in the DNA, wherein the first nucleic acid base is a modified or unmodified nucleic acid base, the second nucleic acid base is a modified or unmodified nucleic acid base different from the first nucleic acid base, and the first and second nucleic acid bases have the same base-pairing specificity; and sequencing the DNA in the first partial sample and the DNA in the second partial sample in such a manner that the first and second nucleic acid bases in the DNA of the first partial sample are distinguishable.
[0145] In one embodiment, the Disclosure provides a non-temporary computer-readable medium containing computer-executable instructions for a method that, when executed by at least one electronic processor, includes: collecting cfDNA from a subject of test; capturing a plurality of sets of target regions from the cfDNA, wherein the plurality of target region sets include a sequence-variable target region set and an epigenetic target region set, thereby producing a captured set of cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a greater depth than the captured cfDNA molecules of the epigenetic target region set; obtaining a plurality of sequence reads generated from the sequencing of the captured cfDNA molecules by a nucleic acid sequencer; mapping the plurality of sequence reads to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and the mapped sequence reads corresponding to the epigenetic target region set to determine the likelihood that the subject has cancer.
[0146] The code may be pre-compiled and configured for use in a machine having a processor adapted to run the code, or it may be compiled during runtime. The code may be supplied in a programming language that can be selected to allow the code to run in pre-compiled fashion or as-compiled fashion.
[0147] Further details regarding computer systems and networks, databases, and computer program products are also provided, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), each of which is thus incorporated herein by reference in its entirety.
[0148] Cancer and other diseases This method can be used to diagnose a condition in a subject, particularly the presence of cancer; to characterize the condition (e.g., to stage cancer or determine cancer heterogeneity); to monitor the response to treatment of the condition; and to determine the risk of developing the condition or the prognosis of the subsequent course of the condition. This disclosure may also be useful in determining the effectiveness of a particular treatment option. A successful treatment option may increase the amount of copy number variation or rare mutations detected in the subject's blood, because if the treatment is successful, more cancer cells may be killed and DNA may be shed. In other cases, this may not occur. In another case, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a treatment.
[0149] Furthermore, if the cancer is observed to be in remission after treatment, this method can be used to monitor residual disease or disease recurrence.
[0150] In some embodiments, the methods and systems disclosed herein can be used to identify customized or targeted therapies for treating a given disease or condition in a patient, based on a classification of whether the nucleic acid variant is of somatic or germline origin. Typically, the disease under consideration is a certain type of cancer. Non-specific examples of such cancers include cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, dysplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia These include clinocyte leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, progenitor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal cancer (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, and acinar cell carcinoma. These include prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.The type and / or stage of cancer can be detected from genetic variations, including mutations, rare mutations, indels, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, alterations in chromosomal structure, gene fusions, chromosome fusions, gene shortening, gene amplification, gene duplication, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in nucleic acid 5-methylcytosine.
[0151] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profiling data may enable the characterization of specific cancer subtypes, which may be important for the diagnosis or treatment of that particular subtype. This information may also provide subjects or practitioners with clues about the prognosis of a particular cancer type, enabling either the subjects or practitioners to adapt treatment options as the disease progresses. Some cancers may become more invasive and genetically unstable as they progress. Other cancers may remain benign, inactive, or dormant. The systems and methods of this disclosure may be useful in determining disease progression.
[0152] Furthermore, the methods of this disclosure can be used to characterize heterogeneity of an abnormal condition in a subject. Such methods may include, for example, the step of generating a gene profile of extracellular polynucleotides derived from the subject, where the gene profile includes multiple data obtained by analysis of copy number variations and rare mutations. In some embodiments, the abnormal condition is cancer. In some embodiments, the abnormal condition may result in a heterogeneous genomic population. In the example of cancer, it is known that some tumors contain tumor cells at different stages of cancer. In other examples, the heterogeneity may constitute multiple disease lesions. Furthermore, in the example of cancer, there may be multiple tumor lesions, possibly one or more lesions resulting from metastasis spreading from the primary site.
[0153] This method can be used to generate or profile fingerprints or sets of data that aggregate genetic information from different cells in heterogeneous diseases. These data sets may include, individually or in combination, analyses of copy number variations, epigenetic variations, and mutations.
[0154] This method can be used to diagnose, prognose, monitor, or observe cancer or other diseases. In some embodiments, the methods described herein do not involve diagnosing, prognosing, or monitoring a fetus, and therefore do not apply to non-invasive prenatal testing. In other embodiments, these methodologies may be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in prenatal subjects where DNA and other polynucleotides can co-circulate with maternal molecules.
[0155] Non-exclusive examples of other gene-based diseases, disorders, or conditions that may be evaluated using the methods and systems disclosed herein as appropriate include: achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cat-crying syndrome, Crohn's disease, cystic fibrosis, Darkham's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombosis, familial hypercholesterolemia, familial Mediterranean fever, and fragile X syndrome. Examples include syndromes such as Gaucher disease, hemochromatosis, hemophilia, holoprosencephalopathy, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatocardiafacial syndrome, WAGR syndrome, and Wilson's disease.
[0156] In some embodiments, the method herein includes the step of detecting the presence or absence of DNA originating from or derived from tumor cells at a pre-selected point in time after previous cancer treatment of a subject previously diagnosed with cancer, using a set of sequence information obtained as described herein. The method may further include the step of determining a cancer recurrence score for the subject of test, indicating the presence or absence of DNA originating from or derived from tumor cells. Once a cancer recurrence score has been determined, the score can be further used to determine a cancer recurrence status. A cancer recurrence status may be, for example, a risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. A cancer recurrence status may be, for example, a low or less risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. In certain embodiments, a cancer recurrence status may result in either a risk of cancer recurrence or a low or less risk of cancer recurrence if the cancer recurrence score is equal to a predetermined threshold.
[0157] In some embodiments, the cancer recurrence score is compared to a predetermined cancer recurrence threshold, and the subject is classified as a candidate for further cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or as not a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, if the cancer recurrence score is equal to the cancer recurrence threshold, the classification may be either candidate for further cancer treatment or not a candidate for treatment.
[0158] The methods described above may further include any applicable features (one or more) set out elsewhere in this Specification, including the section on methods for determining the risk of cancer recurrence in a study subject and / or classifying a study subject as a candidate for subsequent cancer treatment.
[0159] A method for determining the risk of cancer recurrence in a study subject and / or classifying the study subject as a candidate for subsequent cancer treatment. In some embodiments, the methods provided herein are for determining the risk of cancer recurrence in a test subject. In some embodiments, the methods provided herein are for classifying a test subject as a candidate for subsequent cancer treatment.
[0160] Any such method may include the step of collecting DNA (e.g., originating from or obtained from tumor cells) from a test subject diagnosed with cancer at one or more pre-selected time points after one or more previous cancer treatments to the test subject. The subject may be any of the subjects described herein. The DNA may be cfDNA. The DNA may be obtained from a tissue sample.
[0161] One such method may include a step of capturing multiple sets of target regions from DNA derived from a subject, wherein the multiple sets of target regions include a sequence-variable target region set and an epigenetic target region set, thereby producing a captured set of DNA molecules. The capture step can be carried out according to any of the embodiments described elsewhere in this specification. In one such method, the prior cancer treatment may include surgery, administration of a therapeutic composition, and / or chemotherapy.
[0162] One of these methods may include a step of sequencing the captured DNA molecule, thereby producing a set of sequence information. Captured DNA molecules of a sequence-variable target region set can be sequenced to a greater depth than captured DNA molecules of an epigenetic target region set.
[0163] Any such method may include the step of detecting the presence or absence of DNA originating from or derived from tumor cells at a pre-selected point in time using a set of sequence information. The detection of the presence or absence of DNA originating from or derived from tumor cells can be carried out according to any of the embodiments described elsewhere in this specification.
[0164] A method for determining the risk of cancer recurrence in a test subject may include the step of determining a cancer recurrence score for the test subject, indicating the presence or absence, or amount, of DNA originating from or derived from tumor cells. The cancer recurrence score can be further used to determine a cancer recurrence status. A cancer recurrence status may be, for example, a risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. A cancer recurrence status may be, for example, a low or less high risk of cancer recurrence if the cancer recurrence score is above a predetermined threshold. In certain embodiments, a cancer recurrence status may result in either a risk of cancer recurrence or a low or less high risk of cancer recurrence if the cancer recurrence score is equal to a predetermined threshold.
[0165] A method for classifying a test subject as a candidate for subsequent cancer treatment may include comparing the test subject's cancer recurrence score to a predetermined cancer recurrence threshold, thereby classifying the test subject as a candidate for subsequent cancer treatment if the cancer recurrence score exceeds the cancer recurrence threshold, or as not a candidate for treatment if the cancer recurrence score falls below the cancer recurrence threshold. In certain embodiments, if the cancer recurrence score is equal to the cancer recurrence threshold, this may result in either being classified as a candidate for subsequent cancer treatment or not a candidate for treatment. In some embodiments, the subsequent cancer treatment may include chemotherapy or administration of a therapeutic composition.
[0166] One such method may include a step of determining the disease-free survival (DFS) period for the study subjects based on a cancer recurrence score. For example, the DFS period could be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.
[0167] In some embodiments, the set of sequence information includes a sequence variable target region sequence, and the step of determining the cancer recurrence score may include determining at least a first subscore indicating the amount of SNVs, insertions / deletions, CNVs and / or fusions present within the sequence variable target region sequence.
[0168] In some embodiments, the number of mutations in the sequence variable target region, selected from one, two, three, four, or five, is sufficient to result in a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from one, two, or three.
[0169] In some embodiments, the set of sequence information includes epigenetic target region sequences, and the step of determining the cancer recurrence score includes determining a second subscore indicating the amount of molecules (derived from the epigenetic target region sequences) that represent an epigenetic state different from the DNA found in a corresponding sample derived from a healthy subject (e.g., cfDNA found in a blood sample derived from a healthy subject, or DNA found in a tissue sample derived from a healthy subject, where the tissue sample is of the same tissue type as that obtained from the test subject). These abnormal molecules (i.e., molecules having an epigenetic state different from the DNA found in a corresponding sample derived from a healthy subject) may correspond to cancer-related epigenetic changes, e.g., methylation of a hypermethylated variable target region and / or perturbed fragmentation of a fragmented variable target region, where “perturbed” means different from the DNA found in a corresponding sample derived from a healthy subject.
[0170] In some embodiments, a second subscore being classified as positive for cancer recurrence is sufficient if the proportion of molecules corresponding to the hypermethylated variable target region set and / or fragmented variable target region set exhibiting hypermethylation within the hypermethylated variable target region set and / or abnormal fragmentation within the fragmented variable target region set is greater than or equal to a value in the range of 0.001% to 10%. The range may be 0.001% to 1%, 0.005% to 1%, 0.01% to 5%, 0.01% to 2%, or 0.01% to 1%.
[0171] In some embodiments, any such method may include determining the fraction of tumor DNA from a molecular fraction of a set of sequence information that exhibits one or more characteristics indicative of originating from tumor cells. This can be done, for example, for molecules corresponding to some or all of an epigenetic target region, including one or both of a hypermethylation variable target region and a fragmentation variable target region (hypermethylation of the hypermethylation variable target region and / or abnormal fragmentation of the fragmentation variable target region can be considered indicative of originating from tumor cells). This can be done for molecules corresponding to a sequence variable target region, such as molecules that contain changes consistent with cancer, such as SNVs, indels, CNVs, and / or fusions. The fraction of tumor DNA can be determined based on a combination of molecules corresponding to an epigenetic target region and molecules corresponding to a sequence variable target region.
[0172] The determination of the cancer recurrence score can be at least partially based on the fraction of tumor DNA, where the fraction of tumor DNA being greater than a threshold in the range of 10 -11 ~1 or 10 -10 ~1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, the fraction of tumor DNA being greater than or equal to a threshold in the range of 10 -10 ~10 -9 、10 -9 ~10 -8 、10 -8 ~10 -7 、10 -7 ~10 -6 、10 -6 ~10 -5 、10 -5 ~10 -4 、10 -4 ~10 -3 、10 -3 ~10 -2 、or 10 -2 ~10 -1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. In some embodiments, the fraction of tumor DNA is at least 10 -7A value greater than a threshold is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence. The determination that the tumor DNA fraction is greater than a threshold, for example, the threshold corresponding to any of the embodiments described above, can be made based on cumulative probability. For example, a sample was considered positive if the cumulative probability of the tumor fraction being greater than the threshold in any of the aforementioned ranges exceeded a probability threshold of at least 0.5, 0.75, 0.9, 0.95, 0.98, 0.99, 0.995, or 0.999. In some embodiments, the probability threshold is at least 0.95, for example, 0.99.
[0173] In some embodiments, the set of sequence information includes a sequence variable target region sequence and an epigenetic target region sequence, and the step of determining the cancer recurrence score includes determining a first subscore indicating the amount of SNVs, insertions / deletions, CNVs and / or fusions present in the sequence variable target region sequence and a second subscore indicating the amount of abnormal molecules in the epigenetic target region sequence, and combining the first and second subscores to obtain the cancer recurrence score. When combining the first and second subscores, they can be combined by independently applying thresholds to each subscore (e.g., more than a predetermined number of mutations in the sequence variable target region (e.g., >1) and higher than a predetermined fraction of abnormal molecules in the epigenetic target region (i.e., molecules having a different epigenetic state than the DNA found in the corresponding sample derived from a healthy subject; e.g., tumors)), or by training a machine learning classifier to determine the state based on multiple positive and negative training samples.
[0174] In some embodiments, a combined score value in the range of -4 to 2 or -3 to 1 is sufficient for the cancer recurrence score to be classified as positive for cancer recurrence.
[0175] In any embodiment in which the cancer recurrence score is classified as positive for cancer recurrence, the subject's cancer recurrence status may be one of risk for cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment.
[0176] In some embodiments, the cancer is one of the types of cancer described elsewhere in this specification, for example, colorectal cancer. Treatment and related administrations
[0177] In certain embodiments, the methods disclosed herein relate to identifying and administering a customized treatment to a patient, taking into account the state of a nucleic acid variant, whether of somatic or germline origin. In some embodiments, essentially any cancer treatment (e.g., surgery, radiotherapy, chemotherapy, and / or similar) can be included as part of these methods. Typically, a customized treatment comprises at least one immunotherapy (or immunotherapy agent). Immunotherapy generally refers to a method of enhancing the immune response to a given type of cancer. In certain embodiments, immunotherapy refers to a method of enhancing the T-cell response to a tumor or cancer.
[0178] In certain embodiments, the status of nucleic acid variants derived from samples of somatic or germline origin from a subject can be compared to a database of comparison subjects obtained from a reference population to identify customized or targeted therapies for that subject. Typically, the reference population includes patients with the same cancer or disease type as the subject of study, and / or patients who are receiving or have received the same treatment as the subject of study. Customized or targeted therapies (or treatments) can be identified if the nucleic acid variants and comparison subjects meet certain classification criteria (e.g., substantially or substantially matched).
[0179] In certain embodiments, the customized treatments described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized treatments (e.g., immunotherapeutic agents) may also be administered by means such as buccal, sublingual, rectal, vaginal, urethral, topical, intraocular, intranasal, and / or intraauricular, and the administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.
[0180] While preferred embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are presented merely as examples. The present invention is not intended to be limited by any specific examples provided herein. Although the present invention is described in relation to the above specification, the descriptions and illustrations of embodiments herein are not intended to be construed as limiting. Those skilled in the art will readily come up with numerous variations, changes, and substitutions without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to any specific description, configuration, or relative proportion described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments disclosed herein can be used in the practice of the present invention. Accordingly, this disclosure is intended to encompass all such alternatives, modifications, variations, or equivalents. The scope of the present invention is defined by the following claims, and it is intended that the methods and structures that fall within these claims and their equivalents are also encompassed thereto.
[0181] While the foregoing disclosure is described in some detail as explanations and examples for clarity and understanding, it will be apparent to those skilled in the art that various variations in form and detail can be made without departing from the true scope of this disclosure and can be implemented within the scope of the appended claims. For example, all methods, systems, computer-readable media, and / or component features, steps, elements, or other embodiments can be used in various combinations.
[0182] Cancer treatment and therapy In some cases, as cancer treatment, without limitation, imatinib, gefitinib, afatinib, dacomitinib, sunitinib, sorafenib, vandetanib, brivanib, cabozantib, neratinib, tivantinib, bevacizumab, cictumumab, darotuzumab, figtumumab, rilotumumab, onartuzumab, ganitumumab, ramucirumab, ridafololimus, tesirolimus, everolimus, BMS-690514, BMS-754807, EMD Examples include 525797, GDC-0973, GDC-0941, MK-2206, AZD6244, GSK1120212, PX-866, XL821, IMC-A12, MM-121, PF-02341066, RG7160, and Sym004. Suitable antibodies for use as anti-EGFR therapy include cetuximab (brand name: Erbitux) and panitumumab (brand name: Vectibex). In some cases, EGFR tyrosine kinase inhibitors, such as gefitinib (brand name: Iressa), erlotinib (brand name: Tarceva), lapatinib, canertinib, and cetuximab, are used for cancer treatment.
[0183] In some cases, treatments can be used in combination, such as anti-EGFR therapy and anti-EGFR therapy. Anti-EGFR therapy can be used in combination with any combination of chemotherapy agents or chemotherapy regimens, such as FOLFOX (fluorouracil [5-FU] / leucovorin / oxaliplatin) and FOLFIRI (5-FU / leucovorin / irinotecan).
[0184] In some cases, cancer treatments are administered to the target. In some cases, cancer treatments are administered in combination with other treatments, such as non-anti-EGFR therapy and anti-EGFR therapy.
[0185] Genetic analysis Genetic analysis includes the detection of nucleotide sequence variants and copy number variations. Genetic variants can be determined by sequencing. Sequencing methods perform large-scale parallel sequencing, meaning that at least 100,000, 1 million, 10 million, 100 million, or 1 billion polynucleotide molecules can be sequenced simultaneously (or sequentially). Sequencing methods are not limited to these, but include high-throughput sequencing, pyrosequencing, synthetic sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, synthetic single-molecule sequencing (SMSS) (Helicos), large-scale parallel sequencing, clonal single-molecule array (Solexa), shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or Nanopore platforms, and any other sequencing methods known in the art.
[0186] Sequencing can be performed more efficiently by performing sequence capture, that is, by enriching the sample with sequences containing the target sequence of interest, such as the KRAS and / or EGFR genes or portions thereof containing sequence variant biomarkers. Sequence capture can be performed using immobilized probes that hybridize to the target of interest.
[0187] Cell-free DNA may contain small amounts of tumor DNA mixed with germline DNA. Sequencing methods that increase the sensitivity and specificity of detecting tumor DNA, particularly gene sequence variants and copy number variations, may be useful in the methods of the present invention. Such methods are described, for example, in WO2014 / 039556. According to these methods, molecules can be detected with sensitivity down to 0.1% or higher, and these signals can also be distinguished from typical noise in current sequencing methods. Increased sensitivity and specificity from blood-based cfDNA samples can be achieved using a variety of methods. One method involves highly efficient tagging of DNA molecules in the sample, e.g., tagging of at least 50%, 75%, or 90% of the polynucleotides in the sample. This increases the likelihood that low-abundance target molecules in the sample will be tagged and subsequently sequenced, significantly increasing the sensitivity of detecting the target molecules.
[0188] Another method involves molecular tracking, which identifies sequence reads generated redundantly from the original parent molecule and assigns the most likely base identity at each locus or position of the parent molecule. This significantly increases the specificity of detection and reduces the frequency of false positives by reducing noise caused by amplification and sequencing errors.
[0189] Using the methods of this disclosure, genetic variations in non-uniquely tagged initial starting gene material (e.g., rare DNA) at concentrations of less than 5%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, or less than 0.01% can be detected with specificity of at least 99%, 99.9%, 99.99%, 99.999%, 99.9999%, or 99.99999%. Subsequently, sequence reads of the tagged polynucleotides can be tracked to generate consensus sequences for the polynucleotides with error rates of less than or equal to 2%, less than or equal to 1%, less than or equal to 0.1%, or less than or equal to 0.01%.
[0190] Copy number variation determination may involve determining a quantitative measure of polynucleotides in a sample that are mapped to a gene locus, such as the EGFR gene or the KRAS gene. The quantitative measure can be a number. Once the total number of polynucleotides mapped to a locus is determined, this number can be used in a standard method for determining copy number variation at that locus. The quantitative measure can be normalized against a standard. One method allows the quantitative measure at a test locus to be standardized against a quantitative measure of polynucleotides mapped to a control locus in the genome, e.g., a gene with a known copy number. Another method allows the quantitative measure to be compared against the amount of nucleic acid in the original sample. For example, the quantitative measure can be compared against an expected measure for diploidy. Another method allows the quantitative measure to be normalized against a measure derived from a control sample, and the normalized measures at different loci can be compared. Another method involves quantifying the parental or original molecule mapped to the locus in the sample, rather than the number of sequence reads. Copy number variation can be gene amplification, deletion, or shortening. Amplification can involve 3, 4, 5, 6, 7, 8, 9, 10, or 10 or more copies of a gene. Deletion or shortening can involve 0 or 1 copy of a gene.
[0191] An example of a method for detecting copy number variations may include an array. The array may include multiple capture probes. The capture probes may be oligonucleotides bound to the surface of the array. The capture probes may bind to at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, or twelve genes listed in Table 1. DNA derived from the target can be labeled before hybridization for detection (e.g., using a fluorophore).
[0192] In other examples, the gene of interest can be amplified using a primer that recognizes the gene of interest. The primer may hybridize to a gene upstream and / or downstream of a specific region of interest (e.g., upstream of a mutation site). A detection probe may hybridize to the amplification product. The detection probe may hybridize specifically to a wild-type sequence or a mutant / variant sequence. The detection probe may be labeled with a detectable label (e.g., a fluorophore). Detection of the wild-type or mutant sequence can be performed by detecting the detectable label (e.g., fluorescence imaging). In the case of copy number variations, the gene of interest can be compared to a reference gene. Differences in copy numbers between the gene of interest and the reference gene may indicate gene amplification or deletion / shortening. Examples of suitable platforms for carrying out the methods described herein include digital PCR platforms, such as Fluidigm digital arrays. [Examples]
[0193] The following examples are provided for illustrative purposes to illustrate various embodiments of the present invention and are not intended to limit the invention in any way. These examples, together with the methods described herein, represent currently preferred embodiments and are illustrative, and are not intended to limit the scope of the invention. Modifications and other uses that fall within the scope of the spirit of the invention as defined by the claims will be conceivable to those skilled in the art.
[0194] (Example 1) Promoter methylation-related treatment efficacy: In one example, it is interesting to understand whether jedatricib sensitizes advanced TNBC or BRCA1 / 2 mutant breast cancer to talazoparib-assisted PARP inhibition. In a Phase I trial, three patients were classified as having a partial response (PR) and five patients as having stable disease (SD). Histological genomics could not account for confounding scenarios between patients with disease progression and those without. Importantly, an alternative mechanism to gene inactivation exists, and this mechanism can be explained by promoter methylation.
[0195] PI3K inhibitors are thought to reduce the nuclear pool (potentially leading to increased replication errors / fork stoppages / DNA repair). This is thought to increase reliance on repair mechanisms and PARP. PI3K inhibitors are also thought to interfere with the interaction of PI3K with the HR complex. This is thought to increase reliance on PARP for DNA repair.
[0196] Here, the BRCA promoter methylation status of these patients is assessed by analyzing TNBC samples and clinical outcome data, which includes calculating the HRD score by adding promoter methylation data.
[0197] (Example 2) MLH1 promoter methylation test: In another example, this test can identify patients at risk of hereditary / familial forms of colorectal cancer or Lynch syndrome-associated tumor types. MLH1 promoter hypermethylation (and often BRAF V600E positivity) is associated with sporadic forms of colorectal cancer (CRC).
[0198] (Example 3) BRCA1 promoter methylation test: In another embodiment, this can be included as a variant type of homologous recombination repair deficiency-associated tumor type (brca, ovca, panc, prca). Many HRD-associated tumors exhibit single copy loss or rearrangement of BRCA1 without a second hit. Some of these cases have promoter hypermethylation of the remaining allele, which may result in biallelic loss of BRCA1.
[0199] (Example 4) MGMT promoter methylation test: In another embodiment, promoter hypermethylation associated with benefits can be incorporated when treated with a particular type of chemotherapy.
[0200] (Example 5) Allele expression and promoter methylation: The above-described techniques for quantifying promoter methylation levels make it possible to determine methylation levels in an allele-specific manner.
[0201] Using the methods and compositions described herein, a sample is characterized by treating the methylation levels of one or more classification regions, including the determination of a quantitative measure. The determination of a quantitative measure may include the steps of combining multiple nucleic acids derived from at least one of the blood or tissue of interest with a solution containing a certain amount of methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution, and performing multiple washes of the nucleic acid-MBD protein solution with a salt solution to produce several nucleic acid fractions. In some cases, the individual nucleic acid fractions have a threshold number of methylated cytosines within regions of multiple nucleic acids having at least a threshold cytosine-guanine content. Subsequently, one of the multiple washes is performed with a solution having a single concentration of sodium chloride (NaCl) to produce one nucleic acid fraction among several nucleic acid fractions having a range of binding strengths to the MBD protein.
[0202] It can be determined that a first nucleic acid fraction is associated with a first compartment of multiple compartments of nucleic acid, and a first molecular barcode can be attached to the nucleic acid of the first nucleic acid fraction; then, it can be determined that a second nucleic acid fraction is associated with a second compartment of multiple compartments of nucleic acid, and a second molecular barcode can be attached to the nucleic acid of the second nucleic acid fraction, where the first compartment corresponds to a first range of binding strength to MBD protein, and the first molecular barcode is contained within a first set of molecular barcodes associated with the first compartment; and the second compartment corresponds to a second range of binding energy to MBD protein, which is different from the first range of binding strength to MBD protein, and the second molecular barcode is contained within a second set of molecular barcodes associated with the second compartment.
[0203] Using the methods and compositions described above, the ratio of the number of molecules overlapping with the classification region, normalized by the total positive control molecule, can be determined, where the molecule exhibits a threshold amount of methylated cytosine. In some cases, this quantitative measure is compared to a predetermined threshold to call the methylation state of one or more classification regions. In other cases, the step of determining the ratio also includes filtering molecules based on at least a threshold amount of methylated cytosine and / or determining the methylation level of one or more classification regions based on the number of methylated CpGs.
[0204] As a general rule, the importance of specific DNA methylation patterns for developmentally appropriate gene expression is most clearly demonstrated in relation to imprinting loci. While genes are normally expressed from both maternal and paternal alleles, at imprinting loci, only the maternal or only the paternal allele is expressed. In some cases, this restriction may be limited to specific tissues or time periods during development.
[0205] The methylation status of DNA surrounding the imprinting locus also exhibits a pattern specific to each allele. The location of differentially methylated domains or regions (DMDs or DMRs) is variable, and the expressed allele may exhibit both low-methylated and / or high-methylated domains. Parental allele-specific methylation patterns can induce allele-specific expression. In cancer settings, one example is the H19 / Igf2 and Rasgrf1 loci, whose DMRs possess enhancer-blocking activity and bind to CTCF in a methylation-sensitive manner. CTCF bound to unmethylated DMRs blocks the interaction of enhancers necessary for Igf2 and Rasgrf1 expression with their promoters. When the DMRs are methylated and CTCF binding is inhibited, this blockade is relieved, and expression becomes possible.
[0206] (Example 6) Promoter methylation silencing, imprinting: The conventional understanding is that imprinted genes are "silenced," which is a form of monoallelic expression originating from either the maternal or paternal allele. In cancer, some copies of silenced imprinted genes may be reactivated, thereby resulting in expression from both alleles. The loss of monoallelic gene regulation is referred to as loss of imprinting (LOI), and furthermore, amplification of activated copies of imprinted genes, without any effect on the methylation of the silenced copies, has also been observed in cancer cell lines
[19] . In such cases, the imprinted gene may be expressed at two or more transcription sites rather than just one. Therefore, an increased number of transcription site detections of imprinted genes in the cell nucleus can be used as a potential cancer biomarker. Existing reports have shown that these transcription sites can be visualized and labeled using nascent RNA or premRNA in situ hybridization (ISH) methods targeting introns, and that this method is applicable to studies of transcriptional regulation in both imprinted genes.
[0207] Using the methods and compositions described herein, a sample is characterized by processing the methylation level of one or more classification regions, including the determination of a quantitative scale, which includes determining the ratio of the number of molecules overlapping with the classification region, normalized by a total positive control molecule, where the molecules exhibit a threshold amount of methylated cytosine. In some cases, this quantitative scale is compared to a predetermined threshold to call the methylation state of one or more classification regions. Also in some cases, the step of determining the ratio includes filtering molecules based on at least a threshold amount of methylated cytosine and / or determining the methylation level of one or more classification regions based on the number of methylated CpGs.
[0208] (Example 7) Other forms of expression resulting from epigenetic allele states, imprinting: As described, the conventional understanding is that imprinted genes leading to LOIs make cells more susceptible to cell transformation and tumorigenesis due to abnormal biallelelic expression (for example, the imprinted IGF2 locus is thought to promote tumorigenesis by inhibiting apoptosis in colorectal cancer, lead to hyperproliferative defects in lung, colon, and ovarian cancers, and result in LOIs of other imprinted genes such as H19, PEG3, MEST, and PLAGL1 in various cancers).
[0209] Nevertheless, while LOIs are typically associated with the silencing of active alleles, their primary role in downregulation of reported imprinted genes in cancer remains poorly understood. For example, in esophageal cancer, an IGF2 LOI was specifically associated with downregulation of expression and improved survival. Similarly, in prostate cancer, no increase in IGF2 expression was observed despite an LOI. Despite the recognized strong association of LOIs in cancer, this fragmented evidence suggests that the current paradigm regarding the role of LOIs in cancer (i.e., growth-promoting and tumor-promoting expression) needs further evaluation.
[0210] The methods described above enable the determination of imprinted gene networks in which these genes are simultaneously regulated. In parallel, copy number variation (CNV) may be a significant cause of imprinting deregulation in cancer, and the methods and techniques described above lead to multi-mode detection of genomic and epigenomic features.
[0211] The methods and techniques described above support the systematic analysis of LOI or other forms of allele expression, which remains inadequate. While monoallelic expression is better understood, only a few regions in humans have been adequately characterized, and existing methods used to detect abnormal monoallelic expression at a single imprinting locus in cancer fail to understand the evaluated tissue-specific imprinting patterns. Furthermore, the practical applicability of existing high-throughput methods is significantly hindered by the need for genotyping. The techniques described above enable the systematic profiling of (i) monoallelic expression / allelelic expression including imprinting loci, and (ii) their dysregulation and dedysregulation (e.g., LOI) in cancer.
[0212] (Example 8) Epigenetic regulatory imbalances can also increase the plasticity of tumor cells; Epigenetic allele expression for determining tumor heterogeneity: Given the obvious importance of epigenetic regulation, including allele expression, and its role in cancer pathogenesis, it is interesting to apply the methods and compositions described above in situations where tumor heterogeneity is confirmed. Various cancers (e.g., breast cancer) are highly complex heterogeneous diseases, forming tumor subpopulations with distinct phenotypic features at the molecular level. Differences in DNA methylation patterns between different cell subpopulations can drive phenotypic changes, which is useful for providing novel insights into intratumoral epigenetic heterogeneity in breast cancer. In some cases, manual observation of epiaelelic expression has been used to identify differential epiaelelic patterns between the central and peripheral parts of the tumor and to characterize tumor subpopulations with distinct methylation patterns. While methods for epiaelelic imbalances can be calculated based on the Jensen-Shannon divergence, it will be readily apparent to those skilled in the art that various other methods for calculating variations can be utilized. This technique allows for the identification of consecutive CpGs (e.g., four consecutive CpGs) encompassed within the same read as a single epiarelyl. Given that the methylation state of a CpG is either methylated or unmethylated, an epiarelyl contains 16 possible methylation patterns. Divergence (e.g., entropy divergence) can be used to quantify the differences between the methylation patterns of one or more samples.
[0213] Here, the methylation pattern of the tumor (e.g., the center of the tumor) may be more disordered than the tumor periphery, as a result of higher epigenetic heterogeneity, and consequently, genes with higher epigenetic heterogeneity also exhibit higher transcriptional heterogeneity. Using the methods and techniques described above, this can be systematically analyzed to assess the entire set of epigenetic states within the tumor. Four consecutive CpGs covered by the same read were defined as one epiarel.
[0214] (Example 9) Measurement of epigenetic promoter methylation divergence, epirelic diversity, and epigenetic loading with respect to shifts: In other cases, loci with variations in epigenetic alleles can be calculated using the compositional entropy equation (e.g., Methclone(Methyclone)). Here, the epigenetic state of each locus, including cytosine methylation of four consecutive CpG dinucleotides, supports 16 possible CpG methylation patterns at these loci as a single epirelic. An epigenetic shift at a locus may be considered significant if, when comparing one or more samples, the proportion of epirelics at these sites undergoes a statistically significant entropy shift with respect to their composition (calculated by delta-Boltzmann entropy ΔS < Δ90).
[0215] To determine the overall magnitude of genome-wide epirelic shifts, similar to tumor mutational loads, in the form of epirelic loading calculations, the determination of epigenetic status per million loci can be applied to normalize the varying depths of coverage and the number of measured loci per sample. Epigenetic shifts may include both the acquisition and / or loss of epirelics between two samples. Epirelic and systematic epigenetic locus measurements can be determined using methylome data from the methods and compositions described above.
[0216] In various embodiments, orthogonal methylome sequencing methods (e.g., em-SEQ, ERRBS) can be used to validate subsets of specimens. This systematic analysis makes it possible to determine the genetic and epigenetic heterogeneity of tumors and characterize independent, biologically distinct phenomena, each likely possessing its own unique functional significance. The degree of epigenetic allele burden may or may not include other factors, such as age and other clinical parameters, somatic mutations affecting epigenetic modifier genes (e.g., DNMT3A, TET2, and IDH1 / 2), the behavior of dominant epigenetic alleles during clonal evolution in a similar or distinct manner to that of genetic alleles, and further, when monitored longitudinally, at continuous time points, among the ongoing dynamics and patterns of genetic and epigenetic alleles.
[0217] In various cases, these measurements can be used to identify one or more of the epierelecal pattern dynamics and somatic mutational burdens during disease progression, by classifying them using machine learning algorithms (e.g., vector support machines) and / or various databases. For example, diagnostic criteria can be divided into diseases with dominant epierelecal diversity and low somatic mutations (e.g., epigenetically driven) and other diseases with lower epierelecal diversity and higher mutational burdens (e.g., genetically driven). The latter, as the disease progresses, result in increased epigenetic diversity. Here, in both cases, the genetic clonal composition is primarily stable, although examples of genetic clonal stability may also be identified. If there is no association between epigenetic instability and genetic instability or specific somatic mutations, alternative modes of dominant heterogeneity in diagnosed patients can be assessed as including one genetically dominant form and one epigenetically dominant form.
[0218] (Example 10) Epigenetic shift, allele expression measurement: Here, the ability to determine system-wide biological gene activation / inactivation and associated clinical cancer risk, which depends on the epigenetic allele state, including the methylation status of epialleles and the quantification of methylation events across several CpGs, can be considered a form of haplotype definition (e.g., considering both the methylation status of individual CpGs within the sequence read as well as the mean methylation level of the sequence read itself).
[0219] Using exemplary default values (minimum number of CpG sites: 2, minimum mean methylation beta value for CpG: 0.5), and employing data based on next-generation sequencing (NGS), optional thresholding measurements define a subpopulation of the target epiarelels based on the minimum number of cytosines and mean methylation levels in various sequence contexts (e.g., CpG, CHG, or CHH). Thresholding parameters may be sufficiently tunable to target the desired population of epiarelels; sites, maximum mean methylation beta value for non-CpG sites: 0.1). Optional thresholding for sequence reads without thresholding involves calculating the methylation beta value for all genomic locations as the ratio of the number of methylated cytosines to the total number of methylated and unmethylated cytosines: b = C / (C+T). In contrast, when read thresholding is performed (the default mode of action), the level of methylation at each genomic location, i.e., the variant epirele frequency (VEF), can be calculated as the ratio of the number of methylated cytosines (Ca) in read pairs that passed the threshold to the total number of methylated and unmethylated cytosines in all read pairs: VEF = Ca / (C + T).
[0220] By adjusting the range to include the level of extended genomic regions rather than individual bases, the VEF can be made equal to the ratio of the number of read pairs that pass the threshold (Na) to the total number of read pairs that overlap with the region of interest (N): VEF = Na / N. This makes it possible to define groups of epiarelels (i.e., individual methylation patterns) with similar methylation characteristics by setting a threshold, where VEF effectively represents the frequency of this group of epiarelels that pass the threshold at the level of individual cytosines or at the level of extended genomic regions.
[0221] In either case, methylation beta and VEF values from the default reporting format with read threshold settings can be generated from any number of BAM files, without prior hypothesis, provided the experimental configuration allows for calling methylation at a base-by-base level. Both of these values effectively represent methylation levels at each genomic location and can therefore be used directly as input to other bioinformatics tools, including, but not limited to, differential methylation analysis tools.
Claims
1. A step of detecting methylation in one or more promoter regions of at least one of multiple genes, The steps include generating multiple methylation calls to quantify the methylation of one or more promoter regions, and Methods that include...
2. The method according to claim 1, comprising the step of obtaining a sample.
3. The method according to claim 1, comprising the step of obtaining a sample.
4. The method according to claim 1, comprising the step of processing the amount of methylation in one or more promoter regions to characterize the sample.
5. The method according to claim 1, wherein the step of characterizing the sample includes HRD, cancer-derived promoter methylation, familial morphology of colorectal cancer, or Lynch syndrome tumor type.
6. The method according to claim 1, wherein the promoter includes a region 5 kb upstream of the transcription start site (TSS), and the 5 kb region is further refined using one or more of the exclusion of a custom panel region, methylation peaks found in clinical samples, and peaks found in normal samples.
7. The method according to claim 6, wherein the TSS is defined at the transcript level.
8. The method according to claim 6, wherein the TSS is defined at the gene level.
9. The method according to claim 1, comprising the step of determining the ratio of the number of molecules overlapping with the target region, normalized by the total number of positive control molecules.
10. The method according to claim 9, wherein the step of determining the ratio includes filtering the molecules based on at least the number of overlapping CpGs.
11. The method according to claim 1, wherein the step of quantifying the methylation of one or more promoter regions is based on the number of methylated CpGs.
12. The method according to claim 1, comprising the step of refining one or more promoter regions based on at least literature annotations, common methylation peak locations, and / or public datasets.
13. The method according to claim 1, wherein the gene comprises a tumor suppressor gene, an HRR gene, and an IO gene.
14. The method according to claim 13, wherein the HRR gene comprises at least BRCA1 and BRCA2.
15. The method according to claim 1, comprising the step of comparing with a minimum methylation threshold derived from a population of training samples.
16. The method according to claim 15, wherein the training sample includes a sample free of cancer.
17. The minimum methylation threshold for calling is a minimum molecular count of 1 to 100, and the minimum methylation score per gene is the 95th quantile in normality + 8 × 10⁻⁶. 5 The method according to claim 15, comprising at least one of the following: or being the maximum value of median + 5 * median absolute deviation.
18. The method according to claim 1, wherein the step of quantifying the methylation of one or more promoter regions predicts the therapeutic response.
19. The method according to claim 18, wherein the step of quantifying the methylation of one or more promoter regions is combined with the MSI-H state.
20. The method according to claim 18, wherein the treatment comprises one or more of the following: an immune checkpoint inhibitor, a poly(ADP-ribose) polymerase (PARP) inhibitor, a kinase inhibitor, or an aromatase inhibitor, or a PI3K and mTOR inhibitor.
21. The method according to claim 20, wherein the immune checkpoint inhibitor is pembrolizumab.
22. The method according to claim 20, wherein the poly(ADP-ribose) polymerase (PARP) inhibitor is olaparib or talazoparib.
23. The method according to claim 20, wherein the treatment is a combination of a PI3K and mTOR inhibitor and a poly(ADP-ribose) polymerase (PARP) inhibitor.
24. The method according to claim 23, wherein the PI3K and mTOR inhibitor is jedatricib, and the poly(ADP-ribose) polymerase (PARP) inhibitor is talazoparib.
25. Each step involves determining at least one promoter region from among multiple genes obtained from multiple samples, The steps include determining the methylation score for the promoter region and generating multiple methylation calls and / or quantifications of promoter methylation, The steps include processing the aforementioned multiple methylation calls to generate a prediction that the test sample represents the genomic state, and Methods that include...
26. A computing system having one or more hardware processors and memory obtains sequencing reads derived from the target sample. The steps include determining one or more classification regions corresponding to multiple genes contained in the sample, The steps include determining the methylation level of one or more classification regions by generating a quantitative scale derived from the sequencing reads in the target sample, and Methods that include...
27. The method according to claim 26, comprising the step of obtaining a sample.
28. The method according to claim 26, comprising the step of obtaining a sample.
29. The method according to claim 26, comprising the step of processing the methylation levels of one or more classification regions to characterize the sample.
30. The method according to claim 29, wherein the step of characterizing the sample includes determining the HRD status and promoter methylation associated with cancer.
31. The method according to any one of the claims, wherein the quantitative measure comprises determining the ratio of the number of molecules overlapping with the classification region, normalized by a total positive control molecule, and the molecule exhibits a threshold amount of methylated cytosine.
32. The method according to any one of the claims, comprising comparing the quantitative scale with a predetermined threshold and calling the methylation state of one or more classification regions.
33. The method according to any one of the claims, wherein the step of determining the ratio includes filtering the molecules based on at least a threshold amount of methylated cytosine.
34. The method according to any one of the claims, wherein determining the methylation level of one or more classification regions is based on the number of methylated CpGs.
35. The method according to any one of the claims, wherein the classification region includes a promoter region.
36. The method according to any one of the claims, wherein the one or more classification regions individually correspond to genomic regions where the methylation rate of cytosine in the genomic region of nucleic acids derived from cells obtained from a subject in which cancer is present differs from the methylation rate of cytosine in the genomic region of nucleic acids derived from cells obtained from a subject in which cancer is absent.
37. The method according to any one of the claims, wherein the plurality of samples and additional samples include cell-free nucleic acids.
38. The computing system performs a training process using the training data to generate a model, wherein the training process is performed by: The computing system determines one or more additional weights for individual samples included in the training data based on cancer indicators for each individual sample that are within a threshold confidence level. Steps including The method according to any of the claims, including
39. The cancer index for each individual sample is outside the range of the threshold confidence level, and the method is The computing system includes the step of applying a penalty to the weight of each individual sample during the training process, The method according to any of the above claims.
40. The computing system performs one or more first iterations of the training process on the model using a portion of the training data, using one or more machine learning algorithms. A step of generating first output data for the model using the computing system based on the one or more first iterations of the training process, wherein the first output data corresponds to one or more first additional indicators indicating the presence of cancer in a first individual subject among the plurality of subjects, and the first individual subject corresponds to the portion of the training data. The method according to any of the claims, including
41. The computing system performs the steps of combining the first output data and the training data to produce additional training data, The computing system performs one or more second iterations of the training process on the model using a portion of the additional training data. A step of the computing system generating second output data for the model based on the one or more second iterations of the training process, wherein the second output data indicates one or more second additional indicators of the presence of cancer in a second individual subject among the plurality of subjects, and the second individual subject corresponds to the portion of the additional training data. The method according to any of the claims, including
42. The method according to any one of the claims, wherein the weights for each of the plurality of classification regions are determined based on the first output data and the second output data.
43. The computing system determines, during one or more iterations of the training process, that several indicators of the presence of cancer are at least threshold values for one or more samples included in the training data. The computing system determines whether or not modifications to one or more weights of the model are made or only minimally. The method according to any of the claims, including
44. The computing system determines, during one or more iterations of the training process, that the index of the additional number of cancers present is below a threshold for one or more additional samples included in the training data. The computing system determines that the modification to one or more additional weights of the model is greater than the minimum amount. The method according to any of the claims, including
45. The process involves combining multiple nucleic acids derived from at least one of the target blood or tissue with a solution containing a certain amount of methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution, and A step of producing several nucleic acid fractions by performing multiple washes of the nucleic acid-MBD protein solution with a salt solution, wherein each nucleic acid fraction has a threshold number of methylated cytosines within the region of the multiple nucleic acids having at least a threshold cytosine-guanine content. The method according to any of the claims, including
46. The method according to claim 20, wherein one of the plurality of washes is performed using a solution having a single concentration of sodium chloride (NaCl) to produce one nucleic acid fraction among several nucleic acid fractions having a series of binding strengths to the MBD protein.
47. A step of determining that a first nucleic acid fraction is associated with a first compartment of a plurality of compartments of nucleic acid, wherein the first compartment corresponds to a first range of binding strength to the MBD protein, A step of binding a first molecular barcode to the nucleic acid of the first nucleic acid fraction, wherein the first molecular barcode is included in a first set of molecular barcodes associated with the first compartment. A step of determining that a second nucleic acid fraction is associated with a second compartment of the plurality of compartments of nucleic acid, wherein the second compartment corresponds to a second range of binding energy to the MBD protein, which is different from the first range of binding strength to the MBD protein. A step of binding a second molecular barcode to the nucleic acid of the second nucleic acid fraction, wherein the second molecular barcode is included in a second set of molecular barcodes associated with the second compartment. The method according to any of the claims, including
48. A step of producing at least a portion of the plurality of samples used to produce sequencing reads, by combining at least a portion of the aforementioned nucleic acid fractions with a restriction enzyme that cleaves one amount of one or more molecules having unmethylated cytosine, The threshold amount of methylated cytosine corresponds to the lowest frequency of methylated cytosine within the region having at least the threshold cytosine-guanine content, step The method according to any of the claims, including
49. A computing system having one or more hardware processors and memory obtains sequencing reads derived from the target sample. The steps include determining one or more classification regions corresponding to multiple genes contained in the sample, A step of determining the methylation level of one or more classification regions by generating a quantitative scale that includes the ratio of the number of molecules overlapping with the classification regions, normalized by a total positive control molecule, wherein the molecule exhibits a threshold amount of methylated cytosine. The steps include comparing the quantitative scale with a predetermined threshold and calling the methylation state of one or more classification regions. Methods that include...