Tissue-specific methylation markers
By measuring the methylation status of free DNA molecules in biological samples of organisms, the problems of high cost and analytical challenges in measuring circulating DNA in the prior art are solved, and the accurate identification of DNA in different tissues and early detection of cancer are achieved.
Patent Information
- Application Number
- CN202510148128.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-20
- Filing Date
- 2019-03-15
- Publication Date
- 2025-05-09
AI Technical Summary
Existing genome-wide bisulfite sequencing methods are expensive and pose analytical challenges, making it difficult to provide a more cost-effective method to measure circulating DNA from different tissues.
By obtaining free DNA molecules from biological samples of an organism suffering from the first tissue cancer, the methylation status of the target sequence is determined to determine whether the DNA is from the second tissue and to determine whether the organism has cancer located in the second tissue according to the absolute amount.
This method provides a more cost-effective way to detect circulating free DNA, accurately identify the tissue of origin of DNA, and has the advantages of early detection and progression-free survival in cancer management.
Smart Images

Figure CN119955906A_ABST
Abstract
Description
[0001] This application is a divisional application of an application with an application date of March 15, 2019, application number 201980032333.3, and invention name “Tissue-specific methylation markers”. Cross-references
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 643,649, filed on March 15, 2018, and U.S. Provisional Application No. 62 / 769,928, filed on November 20, 2018, each of which is incorporated herein by reference in its entirety. Background Art
[0003] Quantitative measurement of DNA from different tissues to circulating DNA can potentially provide important information about the presence of many different pathological conditions. However, existing methods involving whole genome bisulfite sequencing are relatively expensive and can be challenging to analyze. A more cost-effective method for measuring DNA from different tissues would be useful.
[0004] The detection of circulating free DNA from cancer cells, often referred to as liquid biopsies, is increasingly being used in the management of cancer patients. For example, detection of epidermal growth factor receptor (EGFR) mutations in plasma can correlate well with mutational status in tumor tissue and can predict responsiveness to EGFR tyrosine kinase inhibitors. In addition to point mutations, other cancer-associated genetic and genomic alterations, including copy number alterations and cleavage pattern alterations, can also be detected in the free plasma of cancer patients. Patients identified by screening plasma DNA may have a significantly earlier stage distribution and superior progression-free survival compared with patients who are not screened. Summary of the invention
[0005] One aspect of the present disclosure provides a method for determining whether an organism suffering from cancer in a first tissue suffers from cancer located in a second tissue, the method comprising: (a) obtaining free DNA molecules from a first biological sample of the organism suffering from cancer in the first tissue; (b) assaying the free DNA molecules to determine a first methylation state of a target sequence in the free DNA molecules, wherein the first methylation state of the target sequence indicates that the free DNA molecules including the target sequence are from a second tissue of the organism, wherein the first tissue and the second tissue are different; (c) determining an absolute amount of free DNA molecules from the first biological sample including the target sequence having the first methylation state; and (d) determining whether the organism suffers from cancer located in the second tissue based on the absolute amount.
[0006] In some cases, the methylation state includes a methylation level. In some cases, the assay includes separating the free DNA molecules including the target sequence from the first biological sample. In some cases, the assay includes separating the free DNA molecules including the target sequence in an oil emulsion. In some cases, the assay includes hybridizing the free DNA molecules including the target sequence with a probe. In some cases, the probe hybridizes with the target sequence. In some cases, the hybridization affinity of the probe to the target sequence depends on the first methylation state of the target sequence in the first biological sample. In some cases, when the methylation site of the target sequence is methylated in the first biological sample, the probe hybridizes with the target sequence. In some cases, when the methylation site of the target sequence is not methylated in the first biological sample, the probe hybridizes with the target sequence. In some cases, the assay includes detecting the hybridization of the probe with the target sequence.
[0007] In some cases, the determination includes amplifying the free DNA molecules. In some cases, the amplification includes using a pair of primers. In some cases, the hybridization affinity of at least one primer in the pair of primers to the target sequence depends on the first methylation state of the target sequence. In some cases, when the methylation site of the target sequence is methylated in the first biological sample, the at least one primer in the pair of primers hybridizes with the target sequence. In some cases, when the methylation site of the target sequence is not methylated in the first biological sample, the at least one primer in the pair of primers hybridizes with the target sequence. In some cases, the determination includes converting the non-methylated cytosine residues in the free DNA molecules into uracil via bisulfite. In some cases, the determination includes methylation-aware sequencing of the free DNA molecules from the first biological sample.
[0008] In some cases, the target sequence includes at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10 methylation sites. In some cases, the target sequence includes at least 5 methylation sites. In some cases, the first methylation state includes the methylation density of a single site in the target sequence, the distribution of methylation / non-methylation sites on a continuous region in the target sequence, the methylation pattern or level of each single methylation site in the target sequence, or non-CpG methylation. In some cases, the target sequence includes a higher methylation density in the first tissue than in the second tissue. In some cases, the first methylation state includes the methylation density of a single site in the target sequence, and the methylation density of the target sequence in the first tissue is at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or at least 95%. In some cases, the methylation density included in the first tissue of the target sequence is greater than 50%. In some cases, the target sequence includes a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5% or 0% in the second tissue. In some cases, the target sequence includes a methylation density of less than 20% in the second tissue. In some cases, the target sequence includes a lower methylation density in the first tissue than in the second tissue. In some cases, the target sequence includes a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5% or 0% in the first tissue. In some cases, the target sequence includes a methylation density of less than 50% in the first tissue. In some cases, the target sequence includes a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% or 100% in the second tissue. In some cases, the target sequence includes a methylation density of greater than 80% in the second tissue.
[0009] In some cases, the first tissue comprises liver tissue, and the target sequence comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to SEQ ID NO: 1. In some cases, the first tissue comprises liver tissue, and wherein the determining comprises amplifying using primers comprising SEQ ID NO: 2, primers comprising SEQ ID NO: 3, or both, or using a detectably labeled probe comprising SEQ ID NO: 4 for detecting the target sequence. In some cases, the amplifying further comprises using primers comprising SEQ ID NO: 5, primers comprising SEQ ID NO: 6, or both, or using a detectably labeled probe comprising SEQ ID NO: 7 for detecting the target sequence.
[0010] In some cases, the first tissue comprises colon tissue, and the target sequence comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to SEQ ID NO: 8. In some cases, the first tissue comprises colon tissue, and wherein the amplifying comprises using a primer comprising SEQ ID NO: 9, a primer comprising SEQ ID NO: 10, or both, or using a detectably labeled probe comprising SEQ ID NO: 11 for detecting the target sequence. In some cases, the methylation-specific amplification further comprises using a primer comprising SEQ ID NO: 12, a primer comprising SEQ ID NO: 13, or both, or using a detectably labeled probe comprising SEQ ID NO: 14 for detecting the target sequence.
[0011] In some cases, the cancer is selected from the group consisting of bladder cancer, bone cancer, brain tumors, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastrointestinal cancer, hematopoietic malignancies, head and neck squamous cell carcinoma, leukemia, liver cancer, lung cancer, lymphoma, myeloma, nasal cancer, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, ovarian cancer, prostate cancer, sarcoma, gastric cancer, melanoma, and thyroid cancer. In some cases, the cancer comprises hepatocellular carcinoma or colorectal cancer.
[0012] In some cases, the method further comprises determining the classification of the cancer in the second tissue. In some cases, the determination of the classification of the cancer in the second tissue comprises evaluating free nucleic acid molecules from a second biological sample of the organism. In some cases, the evaluation comprises determining a methylation profile, copy number variation, single polymorphism (SNP) profile, or cleavage pattern of free nucleic acid molecules from the second biological sample. In some cases, the evaluation comprises determining the amount of free nucleic acid molecules from the second biological sample of a pathogen. In some cases, the second biological sample is the same as the first biological sample. In some cases, wherein the second biological sample is different from the first biological sample.
[0013] One aspect of the present disclosure provides a system configured to perform the methods provided herein.
[0014] One aspect of the present disclosure provides a non-transitory computer-readable medium comprising a series of instructions for controlling a computer system to perform the method disclosed herein.
[0015] One aspect of the present disclosure provides a method for analyzing a biological sample of an organism. The method may include: based on the methylation state of a first tissue-specific marker, amplifying a first tissue-specific marker in a free DNA molecule from the biological sample, wherein the first tissue-specific marker includes a predetermined sequence having one or more differential methylation sites, wherein the one or more differential methylation sites have a first methylation state in a first tissue of the organism and a second methylation state in other tissues of the organism, and wherein the first methylation state and the second methylation state are different; identifying the tissue of origin of the free DNA molecule by detecting the amplification of the first tissue-specific marker; and determining the absolute amount of the free DNA molecule from the first tissue of the organism in the biological sample.
[0016] In some cases, the method further comprises converting non-methylated cytosine residues to uracil via bisulfite prior to the amplification. In some cases, the method further comprises separating the free DNA molecule from other DNA molecules in the biological sample prior to the amplification. In some cases, the amplification comprises using a methylation-specific primer complementary to the first tissue-specific marker and annealing to at least a portion of the one or more differentially methylated sites. In some cases, the identification of the tissue of origin comprises determining that the free DNA molecule is from the first tissue if the first tissue-specific marker in the free DNA molecule is amplified by a primer (the primer is configured to amplify a first tissue-specific marker methylated in the first methylation state).
[0017] In some cases, the one or more differential methylation sites have a higher methylation density in the first tissue than in the second tissue. In some cases, the methylation density of the one or more differential methylation sites in the first tissue is at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or at least 95%. In some cases, the one or more differential methylation sites have a methylation density greater than 50% in the first tissue. In some cases, the one or more differential methylation sites have a methylation density in the second tissue of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5% or 0%. In some cases, the methylation density of the one or more differential methylation sites in the second tissue is less than 20%.
[0018] In some cases, the one or more differential methylation sites have a lower methylation density in the first tissue than in the second tissue. In some cases, the methylation density of the one or more differential methylation sites in the first tissue is at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5% or 0%. In some cases, the one or more differential methylation sites have a methylation density of less than 50% in the first tissue. In some cases, the methylation density of the one or more differential methylation sites in the second tissue is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% or 100%. In some cases, the one or more differential methylation sites have a methylation density greater than 80% in the second tissue.
[0019] In some cases, the first tissue comprises liver tissue, and the first tissue-specific marker comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to SEQ ID NO: 1. In some cases, the first tissue comprises liver tissue, and wherein the amplification comprises using a primer comprising SEQ ID NO: 2, a primer comprising SEQ ID NO: 3, or both, or using a detectably labeled probe comprising SEQ ID NO: 4 for detecting the first tissue-specific marker. In some cases, the amplification further comprises using a primer comprising SEQ ID NO: 5, a primer comprising SEQ ID NO: 6, or both, or using a detectably labeled probe comprising SEQ ID NO: 7 for detecting the first tissue-specific marker. In some cases, the first tissue comprises colon tissue, and the first tissue-specific marker comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98%, or 99% identity to SEQ ID NO: 8. In some cases, the first tissue comprises colon tissue, and wherein the amplifying comprises using a primer comprising SEQ ID NO: 9, a primer comprising SEQ ID NO: 10, or both, or using a detectably labeled probe comprising SEQ ID NO: 11 for detecting the first tissue-specific marker. In some cases, the methylation-specific amplification further comprises using a primer comprising SEQ ID NO: 12, a primer comprising SEQ ID NO: 13, or both, or using a detectably labeled probe comprising SEQ ID NO: 14 for detecting the first tissue-specific marker.
[0020] In some cases, the method further comprises determining the amount of free DNA molecules derived from a second tissue in the biological sample based on the methylation pattern of a second tissue-specific marker, wherein the first tissue and the second tissue are different. In some cases, the second tissue belongs to the organism.
[0021] In some cases, the method further comprises diagnosing, monitoring or predicting cancer in the first tissue based on the absolute amount of the free DNA molecules from the first tissue of the organism in the biological sample. In some cases, the cancer is selected from: bladder cancer, bone cancer, brain tumor, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastrointestinal cancer, hematopoietic malignancies, head and neck squamous cell carcinoma, leukemia, liver cancer, lung cancer, lymphoma, myeloma, nasal cancer, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, ovarian cancer, prostate cancer, sarcoma, gastric cancer, melanoma and thyroid cancer. In some cases, the cancer comprises hepatocellular carcinoma or colorectal cancer. In some cases, the diagnosis or monitoring comprises determining the size of the tumor in the first tissue based on the absolute amount of the free DNA molecules from the first tissue of the organism in the biological sample. In some cases, the diagnosis or monitoring comprises determining whether the cancer is transferred to the first tissue based on the absolute amount of the free DNA molecules from the first tissue of the organism in the biological sample.
[0022] In some cases, the first tissue comprises a transplanted organ. In some cases, the methods provided herein further comprise assessing organ transplantation based on the absolute amount of the free DNA molecules from the first tissue in the biological sample.
[0023] Another aspect of the present disclosure provides a composition for determining the amount of free DNA molecules from the liver of an organism in a biological sample. The composition may include a pair of primers for amplifying a liver-specific marker based on the methylation state / level of the liver-specific marker, wherein the liver-specific marker includes a polynucleotide sequence with at least 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 1. In some cases, the pair of primers includes a primer including SEQ ID NO: 2 and a primer including SEQ ID NO: 3. In some cases, the composition also includes a detectably labeled probe including SEQ ID NO: 4 for detecting the liver-specific marker. In some cases, the composition also includes a primer including SEQ ID NO: 5 and a primer including SEQ ID NO: 6. In some cases, the composition also includes a detectably labeled probe including SEQ ID NO: 7 for detecting the liver-specific marker.
[0024] Another aspect of the present disclosure provides a composition for determining the amount of free DNA molecules from the colon of an organism in a biological sample, comprising a pair of primers for amplifying the colon-specific marker based on the methylation state / level of the colon-specific marker, wherein the colon-specific marker comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 8. In some cases, the pair of primers comprises a primer comprising SEQ ID NO: 9 and a primer comprising SEQ ID NO: 10. In some cases, the composition further comprises a detectably labeled probe comprising SEQ ID NO: 11 for detecting the colon-specific marker. In some cases, the composition further comprises a primer comprising SEQ ID NO: 12 and a primer comprising SEQ ID NO: 13. In some cases, the composition further comprises a detectably labeled probe comprising SEQ ID NO: 14 for detecting the colon-specific marker. Incorporation by reference
[0025] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The novel features of the present disclosure are particularly set forth in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by referring to the following detailed description and accompanying drawings which set forth illustrative embodiments in which the principles of the present disclosure are utilized, in which: Figure 1 Schematic diagram of the origin of plasma free DNA in patients with colorectal cancer liver metastases.
[0027] Figure 2 Correlation between fractionated concentrations of liver-derived DNA in the plasma of liver transplant recipients based on liver-specific methylation marker analysis (droplet digital PCR) and donor-specific allele analysis by sequencing is shown.
[0028] Figure 3A and Figure 3B The absolute concentrations of liver-derived DNA obtained using droplet digital PCR in the plasma of healthy subjects, chronic HBV carriers, cirrhotic patients, and HCC patients are shown ( Figure 3A ) and graded concentrations ( Figure 3B ).
[0029] Figure 4A and Figure 4B Figure 2 shows the maximum size of tumors and the absolute concentration of liver-derived DNA obtained using droplet digital PCR in plasma from HCC patients ( Figure 4A ) and graded concentrations ( Figure 4B ) between them.
[0030] Figure 5A-5D Plasma concentrations of colon-derived DNA and liver-derived DNA obtained by droplet digital PCR in healthy subjects and CRC patients with or without liver metastases are shown. Figure 5A ) Absolute concentration of colon-derived DNA, ( Figure 5B ) Graded concentration of colon-derived DNA, ( Figure 5C ) absolute concentration of liver-derived DNA and ( Figure 5D ) Graded concentrations of liver-derived DNA.
[0031] Figure 6 ROC curves for distinguishing CRC patients with and without liver metastasis using absolute and fractional concentrations of liver-derived DNA and colon-derived DNA are shown. "AUC" means area under the curve.
[0032] Figure 7 Shown are the methylation densities of CpG sites within the protein tyrosine kinase 2β (PTK2B) gene region.
[0033] Figure 8 Sestrin 3 ( SESN3 ) Methylation density of CpG sites within gene regions.
[0034] Fig. 9 Shown are ROC curves for distinguishing HCC patients from non-HCC patients using the absolute concentration and fractional concentration of liver-derived DNA.
[0035] Fig.10 A computer control system is shown that may be programmed or otherwise configured to implement the methods provided herein.
[0036] Fig.11 A diagram illustrating the methods and systems disclosed herein is shown. DETAILED DESCRIPTION
[0037] Provided herein are methods, compositions and systems for quantifying free nucleic acid molecules (e.g., free DNA molecules, such as plasma DNA) from specific tissues using tissue-specific markers. Also provided herein are clinical applications of these markers, such as, but not limited to, diagnosis, monitoring and prediction of cancer, detection of metastatic cancer, and organ transplantation assessment in some cases.
[0038] In addition to cancer-specific changes, there may be a general increase in DNA released from organs affected by cancer into the circulation. For example, the level of liver-derived DNA may be elevated in patients with liver cancer. Without wishing to be bound by a particular theory, the elevated levels of DNA released from organs affected by cancer may be due to direct release of DNA from tumor cells or increased turnover of non-tumor cells invaded by cancer. This increase in tissue-specific DNA release can also be observed in organ transplant recipients who experience acute rejection due to increased cell turnover. In some cases, the proportional contribution of cells in the transplanted organ to plasma DNA determined by methylation deconvolution is well correlated with the proportional contribution determined based on donor-specific allele analysis.
[0039] In some cases, methylation deconvolution using whole-genome bisulfite sequencing can be challenging due to relatively high costs and long run times. In addition, in some cases, methylation deconvolution may determine the relative or proportional contributions of different organs rather than the absolute concentration of DNA originating from each organ. In cases where DNA from more than one organ is released into the circulation, measuring the absolute concentration of DNA originating from a specific organ can provide more information. For example, in patients with colorectal cancer (CRC) metastases to the liver, increased amounts of DNA will be released from the liver into the circulation. However, due to an even greater increase in DNA released by tumor cells originating from the colon, a paradoxical decrease in the proportion of liver DNA in the plasma may occur. Therefore, it may be useful to develop methods that can accurately determine the absolute amount of DNA with tissue-specific methylation patterns.
[0040] I. Tissue-specific markers Provided herein are tissue-specific markers that can identify the tissue of origin of free DNA molecules. In some cases, tissue-specific markers are polynucleotide sequences of organism genomes. In some cases, tissue-specific markers include differential methylation regions (DMRs), which are identified based on the methylation state of one or more differential methylation sites contained in the marker polynucleotide sequence. In some cases, one or more differential methylation sites include one or more CpG sites. In some cases, one or more differential methylation sites include one or more non-CpG sites. In some cases, the tissue-specific markers discussed herein can be referred to as target sequences.
[0041] In some cases, a differentially methylated site of a tissue-specific marker has a first methylation state in a first tissue of an organism and a second methylation state in a different second tissue of the organism. The first methylation state and the second methylation state may be different, so that the first tissue and the second tissue may be distinguished based on the methylation state of the tissue-specific marker.
[0042] In some cases, a differentially methylated site of a tissue-specific marker has a first methylation state in a first tissue of an organism and a second methylation state in all other tissues of the organism. The first methylation state and the second methylation state can be different, so that the first tissue can be distinguished from all other tissues of the organism based on the methylation state of the tissue-specific marker.
[0043] In some cases, tissue-specific markers include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 50 differential methylation sites. In some cases, tissue-specific markers include at least 5 differential methylation sites. Methylated nucleotides or methylated nucleotide bases can refer to the presence of methyl moieties on nucleotide bases, wherein there is no methyl moiety in generally recognized typical nucleotide bases. For example, cytosine does not contain methyl moieties on its pyrimidine ring, but 5-methylcytosine contains methyl moieties at 5 positions of its pyrimidine ring. In this case, cytosine is not a methylated nucleotide, and 5-methylcytosine is a methylated nucleotide. In another example, thymine contains methyl moieties at 5 positions of its pyrimidine ring, however, for purposes herein, when thymine is present in DNA, it is not considered to be a methylated nucleotide, because thymine is a typical nucleotide base of DNA. The typical nucleoside bases of DNA are thymine, adenine, cytosine and guanine. The typical bases of RNA are uracil, adenine, cytosine and guanine. Accordingly, "methylation site" can be a position where methylation occurs or may occur in the target gene nucleic acid region. For example, a position containing CpG is a methylation site, in which cytosine may be methylated or may not be methylated." Site" can correspond to a single site, which can be a single base position or a group of related base positions, for example, a CpG site. Methylation site can refer to a CpG site or a non-CpG site of a DNA molecule that may be methylated. A CpG site can be a region of a DNA molecule as follows: in this region, a guanine nucleotide is followed by a guanine nucleotide in a base linear sequence along its 5' to 3' direction, and this region is easily methylated by a naturally occurring event in vivo or by a chemical action performed in vitro to make the nucleotide methylated event. A non-CpG site may be a region that does not have a CpG dinucleotide sequence but is also susceptible to methylation by events occurring naturally in vivo or by chemical methylation of nucleotides performed in vitro.A locus or region may correspond to a region that includes multiple sites.
[0044] The methylation state of tissue-specific markers may include: the methylation density of each site in the marker region, the distribution of methylated / non-methylated sites on continuous regions within the marker, the methylation pattern or level of each single methylated site within a marker containing more than one site, and non-CpG methylation. In some cases, the methylation state of tissue-specific markers includes the methylation level (or methylation density) of each differentially methylated site. For a given methylation site, the methylation density may refer to the fraction of nucleic acid molecules methylated at a given methylation site in the total number of nucleic acid molecules of interest containing the methylation site. For example, the methylation density of the first methylation site in liver tissue may refer to the fraction of liver DNA molecules methylated at the first site in the total liver DNA molecules. In some cases, the methylation state includes the coherence of the methylation / non-methylation state between each differentially methylated site.
[0045] In some cases, a tissue-specific marker includes a methylation site that is hypermethylated in a first tissue and hypomethylated in a second tissue. For example, a tissue-specific marker may include one or more methylation sites that are hypermethylated in liver tissue, thereby indicating that the one or more methylation sites have a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% in liver tissue; in contrast, the one or more methylation sites may have a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5%, or 0% in other tissues (e.g., but not limited to, blood cells, lungs, esophagus, stomach, small intestine, colon, pancreas, bladder, heart, and brain). Tissue-specific markers can include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 30, at least 50 methylation sites that are hypermethylated in a first tissue and hypomethylated in a second tissue. In some cases, tissue-specific markers include at least 5 methylation sites that are hypermethylated in a first tissue and hypomethylated in a second tissue.
[0046] Tissue-specific markers may include at most 300 base pairs (bp), at most 250 bp, at most 225 bp, at most 200 bp, at most 190 bp, at most 185 bp, at most 180 bp, at most 175 bp, at most 170 bp, at most 169 bp, at most 168 bp, at most 167 bp, at most 166 bp, at most 165 bp, at most 164 bp, at most 163 bp, at most 162 bp, at most 161 bp, at most 160 bp, at most 150 bp, at most 140 bp, at most 120 bp, or at most 100 bp. In some cases, tissue-specific markers include at most 166 bp.
[0047] In some cases, the liver-specific markers provided herein include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 50 methylation sites that are hypermethylated in liver tissue and hypomethylated in other tissues. In some cases, the liver-specific markers provided herein include at least 5 methylation sites that are hypermethylated in liver tissue and hypomethylated in other tissues. Each methylation site that is hypermethylated in liver tissue can have a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or at least 95% in liver tissue, and a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5% or 0% in other tissues (such as, but not limited to, blood, brain, thymus, pancreas, kidney). In some cases, each methylation site that is hypermethylated in liver tissue may have a methylation density in liver tissue greater than 50%, while a methylation density in other tissues (such as, but not limited to, blood cells, lung, esophagus, stomach, small intestine, colon, pancreas, bladder, heart, and brain) less than 20%.
[0048] In some cases, tissue-specific markers include methylation sites that are hypomethylated in a first tissue and hypermethylated in a second tissue. In some cases, tissue-specific markers include methylation sites that are hypomethylated in a first tissue and hypermethylated in other tissues. For example, a tissue-specific marker may include one or more methylation sites that are hypomethylated in liver tissue, which may mean that one or more methylation sites have a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5% or 0% in liver tissue; on the contrary, one or more methylation sites may have a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or at least 95% in other tissues (e.g., but not limited to, blood cells, lungs, esophagus, stomach, small intestine, colon, pancreas, bladder, heart, and brain). The tissue-specific marker may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 30, at least 50 methylation sites that are hypomethylated in a first tissue and hypermethylated in a second tissue. In some cases, the tissue-specific marker includes at least 5 methylation sites that are hypomethylated in a first tissue and hypermethylated in a second tissue. In some cases, the liver-specific markers provided herein include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 50 methylation sites that are hypomethylated in liver tissue and hypermethylated in other tissues. In some cases, the liver-specific markers provided herein include at least 5 methylation sites that are hypomethylated in liver tissue and hypermethylated in other tissues. Each methylation site that is hypomethylated in liver tissue can have a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10% or at most 5% in liver tissue, and a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% or 100% in other tissues (such as, but not limited to, blood, brain, thymus, pancreas, kidney). In some cases, each methylation site that is hypomethylated in liver tissue may have a methylation density of less than 50% in liver tissue and greater than 80% in other tissues (e.g., but not limited to, blood cells, lung, esophagus, stomach, small intestine, colon, pancreas, bladder, heart, and brain).
[0049] In some cases, the colon-specific markers provided herein include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 50 methylation sites that are hypermethylated in colon tissue and hypomethylated in other tissues. In some cases, the colon-specific markers provided herein include at least 5 methylation sites that are hypermethylated in colon tissue and hypomethylated in other tissues. Each methylation site that is hypermethylated in colon tissue can have a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% in colon tissue, and a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10%, at most 5%, or 0% in other tissues (e.g., but not limited to, blood, brain, thymus, pancreas, kidney). In some cases, each methylation site that is hypermethylated in colon tissue can have a methylation density greater than 50% in colon tissue and less than 20% in other tissues (e.g., but not limited to, blood cells, lung, esophagus, stomach, small intestine, liver, pancreas, bladder, heart, and brain).
[0050] In some cases, the colon-specific markers provided herein include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20, at least 25, at least 30, at least 50 methylation sites that are hypomethylated in colon tissue and hypermethylated in other tissues. In some cases, the colon-specific markers provided herein include at least 5 methylation sites that are hypomethylated in colon tissue and hypermethylated in other tissues. Each methylation site that is hypomethylated in colon tissue may have a methylation density of at most 50%, at most 40%, at most 30%, at most 20%, at most 10% or at most 5% in colon tissue, and a methylation density of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95% or 100% in other tissues (e.g., but not limited to, blood, brain, thymus, pancreas, kidney). In some cases, each methylation site that is hypomethylated in colon tissue may have a methylation density of less than 50% in colon tissue and greater than 80% in other tissues (e.g., but not limited to, blood cells, lung, esophagus, stomach, small intestine, liver, pancreas, bladder, heart, and brain).
[0051] Also provided herein is a liver-specific marker for identifying liver-derived DNA molecules. The liver-specific marker can be located in the exon region of the protein tyrosine kinase 2β (PTK2B) gene on chromosome 8. The eight CpG sites in the liver-specific DMR can be hypermethylated in the liver and hypomethylated in other tissues and blood cells. The liver-specific marker provided herein can include a polynucleotide sequence with at least about 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 1. The liver-specific marker provided herein can include SEQ ID NO: 1.
[0052] Also provided herein are colon-specific markers for identifying colon-derived DNA molecules. The colon-specific marker can be located in the exon region of the Sestrin 3 (SESN3) gene on chromosome 11. All six CpG sites located within the colon-specific DMR can be hypermethylated in the colon and hypomethylated in other tissues and blood cells. The colon-specific marker can include a polynucleotide sequence having at least about 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 8.
[0053] As used herein, the term "identity" or "percent identity" between two or more nucleotide or amino acid sequences can be determined by aligning the sequences to achieve optimal comparison purposes (e.g., a gap can be introduced into the sequence of the first sequence). The nucleotides at the corresponding positions can then be compared, and the percent identity between the two sequences can be a function of the number of identical positions shared by the sequences (i.e., % identity = number of identical positions / total number of positions × 100). For example, a position in the first sequence can be occupied by the same nucleotide as the corresponding position in the second sequence, and the molecules are identical at that position. The percent identity between the two sequences can be a function of the number of identical positions shared by the sequences, which needs to be introduced to achieve optimal alignment of the two sequences, taking into account the number of gaps and the length of each gap. In some cases, the length of the sequence aligned for comparison purposes is at least about: 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 95% of the length of the reference sequence. A BLAST® search can determine the homology between two sequences. Homology can be between the entire length of the two sequences or between parts of the entire length of the two sequences. The two sequences can be genes, nucleotide sequences, protein sequences, peptide sequences, amino acid sequences or fragments thereof. The actual comparison of the two sequences can be accomplished by well-known methods, for example, using a mathematical algorithm. Non-limiting examples of such mathematical algorithms can be those described in Karlin, S. and Altschul, S., Proc. Natl. Acad. Sci. USA, 90-5873-5877 (1993). Such algorithms can be incorporated into NBLAST and XBLAST programs (version 2.0), as described in Altschul, S. et al., Nucleic Acids Res., 25:3389-3402 (1997). When using BLAST and Gapped BLAST programs, any relevant parameters of the corresponding programs (e.g., NBLAST) can be used. For example, the parameters for sequence comparison can be set to score = 100, word length = 12, or can be changed (e.g., W = 5 or W = 20). Other examples include the algorithm of Myers and Miller, CABIOS (1989), ADVANCE, ADAM, BLAT, and FASTA.
[0054] The present invention also provides a method for identifying tissue-specific markers. The method may include comparing the methylation status on the genome in different tissue samples. Publicly available databases, such as the RoadMapEpigenomics Project (Roadmap Epigenomics Consortium et al. Nature 2015;518:317-30) and the BLUEPRINT project (Martens et al. Haematologica 2013;98:1487-9), can be used for bioinformatics analysis to screen potential tissue-specific markers. In some cases, experimental verification is required. For example, methylation-aware sequencing, such as bisulfite sequencing, can be performed to verify the methylation status among different tissues. In some cases, methylation-specific amplification can also be used for relatively more targeted verification.
[0055] II. Methods for analyzing free DNA molecules in tissues Provided herein is a method for analyzing a biological sample of an organism. The method provided herein can determine the free nucleic acid molecules from an organism tissue, such as the absolute amount of free DNA molecules. The method can include identifying the free DNA molecules from a biological sample as free DNA molecules from a first tissue of an organism when the free DNA molecules include a first tissue-specific marker with a first methylation state. In some cases, the first tissue-specific marker includes a predetermined sequence with one or more differential methylation sites. In some cases, the first tissue-specific marker can have a first methylation state in a first tissue of an organism, and have a second methylation state in other tissues of the organism, and the first methylation state and the second methylation state can be different.
[0056] As provided herein, method can include the methylation state of assessing free DNA molecules. In some cases, method can also include before amplification, non-methylated cytosine residues are converted into uracil through bisulfite. In some cases, method can include by any other method conversion methylated or non-methylated cytosine residues, make it possible to distinguish the residues converted by subsequent detection methods, such as primer-based amplification. In some cases, free DNA molecules can be digested with methylation sensitive enzymes, when methylation sites are methylated or non-methylated, methylation sensitive enzymes digest DNA molecules at one or more specific methylation sites. Therefore, methylation sensitive enzymes can be used to distinguish methylated and non-methylated DNA molecules. Non-limiting examples of methylation-sensitive enzymes that can be used in the methods provided herein can include Aat II, Acc II, Aor13H I, Aor51H I, BspT104 I, BssH II, Cfr10 I, Cla I, Cpo I, Eco52 I, HaeII, Hap II, Hha I, Mlu I, Nae I, Not I, Nru I, Nsb I, PmaC I, Psp1406 I, Pvu I, Sac II, Sal I, Sma I, SnaB I, and any combination thereof.
[0057] The method provided herein may include amplifying a first tissue-specific marker in a free DNA molecule from a biological sample based on the methylation state of the first tissue-specific marker. The method may also include identifying the tissue of origin of the free DNA molecule by detecting the amplification of the first tissue-specific marker. The method may also include determining the absolute amount of the free DNA molecules in the biological sample of the first tissue from an organism. Amplification reaction may refer to the process of replicating nucleic acid once or multiple times. In some cases, amplification methods include but are not limited to polymerase chain reaction (PCR), self-maintaining sequence reaction, ligase chain reaction, rapid amplification of cDNA ends, polymerase chain reaction and ligase chain reaction, Q-β phage amplification, chain displacement amplification or splicing overlap extension polymerase chain reaction. In some cases, a single nucleic acid molecule, for example, is amplified by digital PCR.
[0058] In some cases, amplification includes using a methylation-specific primer complementary to the first tissue-specific marker and annealing to at least a portion of one or more methylation sites. A methylation-specific primer may refer to a primer that can distinguish between methylated and non-methylated target sequences. For a given target sequence containing a methylation site, a methylation-specific primer may be designed to cover at least a portion of the methylation site. In some cases, when bisulfite conversion is performed before amplification, the nucleotide residues on the methylation-specific primer for a given methylation site may be designed to be complementary to the unconverted cytosine residues at the site to detect the methylated target sequence, and for detecting the non-methylated target sequence, the nucleotide residues on the methylation-specific primer may be designed to be complementary to the converted residues (e.g., uracil residues) at the site.
[0059] In some cases, the method may further include separating the free DNA molecules from other DNA molecules in the biological sample prior to amplification. By physically separating the free DNA molecules, a "digital" assay of the free DNA molecules in the biological sample may be performed, such as, but not limited to, digital PCR, e.g., droplet digital PCR. The separation method may be any method known to the skilled person, such as, but not limited to, separation by arrays of microplates, capillaries, oil emulsions, and miniaturized chambers. The digital PCR reaction may be performed using any technique known in the art, such as microfluidics-based, or emulsion-based, e.g., BEAMing (Dressman et al. Proc Natl Acad Sci USA 2003; 100 : 8817-8822).
[0060] In some cases, the first tissue includes liver tissue, and the method includes using the above-mentioned liver-specific markers. In some cases, liver-specific markers, such as SEQ ID NO: 1, include sites that are highly methylated in the liver but hypomethylated in other tissues. In these cases, the method may include primers for detecting methylated liver tissue-specific markers ("primers for methylation determination"). For example, primers for methylation determination may include primers including SEQ ID NO: 2, primers including SEQ ID NO: 3, or both. In some cases, primers provided herein are used for amplification reactions after bisulfite conversion. Alternatively or cumulatively, the method may include using a detectably labeled probe including SEQ ID NO: 4 to detect methylated liver tissue-specific markers. Optionally, the method may also include using primers for detecting non-methylated liver tissue-specific markers ("primers for non-methylation determination"). Primers for non-methylation determination may include primers including SEQ ID NO: 5, primers including SEQ ID NO: 6, or both. Alternatively or cumulatively, the method may comprise using a detectably labeled probe comprising SEQ ID NO: 7 to detect a non-methylated liver tissue-specific marker.
[0061] Primer, probe or oligonucleotide can be used interchangeably in this article, and can refer to more than one, for example, 2, 4, 6, 8, 10, 14, 18, 20 or 40 nucleotides or chemically modified nucleotides linked together by phosphodiester bonds. Primer can include about 20 to about 30 nucleotides. Primer can include at least 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38 or 40 nucleotides. Chemical modifications of primer composition nucleotides can be introduced to modify some properties of primers, for example, increase its stability, increase hybridization specificity and use detectable signal markers. The chemical modifications that can be used in the compositions provided herein can include covalent modification, attachment chemistry and nucleotide base modification. Chemical modifications may include phosphorylation, addition of biotin, cholesteryl-TEG, amino modifiers (e.g., C6, C12, or dT), azides, alkynes, thiol modifiers, fluorophores, dark quenchers, and spacers. Primers may include one or more phosphorothioate bonds, or one or more modified bases, such as 2-aminopurine, 2,6-diaminopurine (2-amino-dA), 5-bromo-dU, deoxyuridine, reverse dT, reverse dideoxy-T, dideoxy-C, 5-methyl dC, deoxyguanosine, superT®, superG®, locked nucleic acid, 5-nitroindole, 2'-O-methyl RNA bases, hydroxymethyl dC, iso-dC, iso-dG, fluoro bases (e.g., fluoro-C, U, A, G, T), and 2'-O-methoxy-ethyl bases (e.g., 2-methoxyethoxy A, G, U, C, T). Detectably labeled probes may include any detectable label known to those skilled in the art, e.g., any suitable fluorophore. PCR reactions (e.g., digital PCR or real-time PCR) may be monitored by any suitable optical, magnetic, electronic, or any other technique available to those skilled in the art.
[0062] In some cases, the first tissue includes colon tissue, and the method includes using the colon-specific markers described above. In some cases, colon-specific markers, such as SEQ ID NO: 8, include sites that are highly methylated in the colon but hypomethylated in other tissues. In these cases, primers for methylation determination may include primers including SEQ ID NO: 9, primers including SEQ ID NO: 10, or both. In some cases, primers provided herein are used for amplification reactions after bisulfite conversion. Alternatively or cumulatively, the method may include using a detectably labeled probe including SEQ ID NO: 11 to detect methylated colon-specific markers. Optionally, the method may also include using primers for detecting non-methylated colon-specific markers ("primers for non-methylation determination"). Primers for non-methylation determination may include primers including SEQ ID NO: 12, primers including SEQ ID NO: 13, or both. Alternatively or cumulatively, the method may comprise using a detectably labeled probe comprising SEQ ID NO: 14 to detect an unmethylated colon-specific marker.
[0063] The methods provided herein, for example, using digital PCR technology, can make it possible to directly determine the actual number of target DNA molecules without the use of a calibrator. Other techniques, such as certain sequencing-based methods, such as, but not limited to, bisulfite sequencing and non-bisulfite-based methylation-aware sequencing using the PacBio sequencing platform, can determine the relative or fractional concentration of the DNA of the target tissue relative to other tissues. Absolute quantity can refer to the absolute count of DNA molecules, or in some cases, it can also refer to the concentration of DNA molecules, such as the number, mole or weight per volume, such as, copy number / mL, mole / L or mg / L. Absolute quantity analysis as provided herein can be useful in the case of releasing an increased amount of DNA from more than one type of tissue. On the other hand, methylation deconvolution analysis based on sequencing of free nucleic acid molecules, such as disclosed in U.S. Patent Application No. 14 / 803,692, can provide a reading of the source tissue of the free nucleic acid in the form of a fractional contribution, for example, a first tissue contributes A% of the free nucleic acid from a biological sample, and a second tissue contributes B% of the free nucleic acid from the same biological sample.
[0064] In some cases, the methods, compositions and systems provided herein can also utilize technologies such as real-time PCR, sequencing and microarrays to perform methylation analysis on free nucleic acids. In some cases, the absolute number of free nucleic acids with tissue-specific markers, such as counting of positive reactions in digital PCR assays, may not be directly derived from methylation analysis by some technologies. However, this absolute number can be calculated indirectly based on the concentration (relative or graded) of free nucleic acids carrying tissue-specific markers, for example, by considering the total number or concentration of free nucleic acids in a given volume of biological sample. In some cases, sequencing that can be used in the methods provided herein can include chain termination sequencing, hybridization sequencing, Illumina sequencing (e.g., using reversible terminator dyes), ion torrent semiconductor sequencing, mass spectrometry sequencing, massively parallel sequencing (MPSS), Maxam-Gilbert sequencing, nanopore sequencing, polony sequencing, pyrophosphate sequencing, shotgun sequencing, single molecule real-time (SMRT) sequencing, SOLiD sequencing (using four fluorescently labeled double base probe hybridization), universal sequencing or any combination thereof. Microarrays with probes targeted to methylation sites can also be used in the methods provided herein to analyze the methylation status of free DNA molecules.
[0065] In some cases, the methods provided herein also include determining the amount of free DNA molecules derived from a second tissue in the biological sample based on the methylation pattern of a second tissue-specific marker, wherein the first tissue and the second tissue are different. The second tissue can belong to the same organism. The second tissue can also be from a different organism, for example, a fetus in a pregnant woman.
[0066] III. Methods for Diagnosis, Monitoring and Prognosis of Cancer Provided herein are methods for diagnosing, monitoring, and predicting cancer. The methods may include: determining the absolute amount of free DNA molecules from a first tissue as described above; and diagnosing, monitoring, and predicting cancer in the first tissue based on the absolute amount of free DNA molecules from the first tissue of the organism in the biological sample.
[0067] The absolute amount of free DNA molecules from the first tissue can be related to the condition of the first tissue. For example, the amount of liver-derived plasma DNA molecules can increase due to increased release of DNA molecules from liver tissue due to tumor growth. In other cases, for example, increased cell turnover due to organ transplantation may also lead to increased plasma DNA released from tissue accompanying the transplant.
[0068] The methods provided herein can include determining the size of a tumor in a first tissue based on the absolute amount of free DNA molecules from a first tissue of an organism in a biological sample. A predetermined comparison diagram associating the amount of a target free DNA molecule with tumor size can be used for determining tumor size. Detection of tumor size can help diagnose, monitor and predict cancer.
[0069] Methods provided herein can include determining whether cancer has metastasized to a first tissue based on the absolute amount of free DNA molecules from a first tissue of an organism in a biological sample. Compared to the fractionated amount of target DNA molecules, the absolute amount of target DNA molecules determined by the methods provided herein can provide a desired distinction between cancer patients with and without metastasis.
[0070] The cancer types to which the methods, compositions and systems provided herein are applicable can include bladder cancer, bone cancer, brain tumors, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastrointestinal cancer, hematopoietic malignancies, squamous cell carcinoma of the head and neck, leukemia, liver cancer, lung cancer, lymphoma, myeloma, nasal cancer, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, ovarian cancer, prostate cancer, sarcoma, gastric cancer or thyroid cancer. The metastatic tissues assessed by the methods provided herein can include bladder, bone, brain, breast, cervix, colon, esophagus, gastrointestinal tract, blood, head, neck, liver, lung, lymph node, nose, nasopharynx, mouth, oropharynx, ovary, prostate, skin, stomach or thyroid.
[0071] As will be readily appreciated by those skilled in the art, cancer cells can spread locally by moving into nearby normal tissues, can spread locally to nearby lymph nodes, tissues or organs, and can spread to distant parts of the body. The spread of cancer from an initial first tissue to a second tissue can be referred to as metastasis, and thus such cancers can be referred to as metastatic cancers. Exemplary cancer metastasis types to which the methods, compositions and systems provided herein can be applied can include metastasis occurring at the sites listed in Table 1.
[0072] Table 1. Exemplary cancer metastasis sites
[0073] In some cases, when combined with other techniques available to those skilled in the art, the methods, compositions and systems provided herein can be used for diagnosing, monitoring and predicting cancer. In some cases, the detection of other molecular markers (e.g., in nucleic acids, e.g., DNA, RNA), such as copy number aberrations (CNA), single nucleotide polymorphisms (SNPs), genetic mutations, germline mutations, somatic mutations, nucleic acids from pathogens (e.g., viruses, e.g., Epstein-Barr virus), the size of free nucleic acids and the cracking mode of free nucleic acids, can also be used in combination with the methods, compositions and systems provided herein. The combination of technology can help promote the detection of cancer levels, including, but not limited to, whether cancer exists, the stage of cancer, the size of tumors, the loss or amplification (e.g., duplication or triplicate) of how many chromosome regions are involved and / or the measurement of the severity of other cancers. The cancer level can be many other features. The level can be zero. The cancer level can also include pre-malignant or precancerous conditions associated with loss or amplification.
[0074] In some cases, detection of copy number aberrations (CNAs), such as the methods disclosed in U.S. Patent No. 8,741,811, can be used in combination with the methods provided herein. As described above, in some cases, based on the absolute amount of free nucleic acids from a first tissue, the methods provided herein can determine whether the cancer has metastasized to the first tissue. On the other hand, detection of CNA can help identify the origin of metastatic cancer cells in the first tissue. In some cases, analysis of the cleavage pattern of free nucleic acids, such as the methods disclosed in U.S. Patent Application No. 15 / 218,497, can be used in combination with the methods, compositions, and systems provided herein. In some cases, the subject methods, compositions, and systems can be used in combination with any available method for detecting, monitoring, or predicting cancer in a subject. In addition to the above detection methods, any appropriate test may be performed, such as tumor biomarker testing (e.g., alpha-fetoprotein (AFP) for liver cancer, ALK gene for non-small cell lung cancer, prostate-specific antigen (PSA) for prostate cancer, and thyroglobulin for thyroid cancer), physical examination, imaging (e.g., computed tomography, magnetic resonance imaging, positron emission tomography (PET)), ultrasonography, endoscopy, biopsy, or cytology testing.
[0075] In some cases, the subject method, composition or system can be used to monitor the cancer of the experimenter with regular, semi-regular or irregular timetable.For example, the experimenter can be based on weekly, monthly, quarterly or annually using the cancer monitoring inspection of the subject method, composition or system.In some cases, the experimenter can carry out this type of inspection every 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months or more than 12 months.In some cases, can be based on the result of recent inspection, for example, in some cases, determine the interval between two consecutive inspections according to doctor's prescription or medical advice.
[0076] IV. Methods of evaluating organ transplantation Provided herein are methods for evaluating organ transplantation based on determining the absolute amount of free DNA molecules from transplanted tissue. Transplanted tissue as described herein is considered to be tissue of the subject of interest.
[0077] As provided herein, the method can utilize the correlation between the amount of free DNA molecules from transplanted tissue and the cell renewal rate in the transplanted tissue. Therefore, the cell renewal rate can be used as a criterion for evaluating organ transplantation.
[0078] V. Compositions for analyzing free DNA molecules from tissues Also provided herein are compositions for analyzing free DNA molecules from specific tissues (e.g., bone, liver, lung, brain, peritoneum, adrenal glands, skin, muscle, vagina, colon, bladder, breast, kidney, melanoma, ovary, pancreas, prostate, rectum, stomach, thyroid, or uterus).
[0079] The composition for determining the amount of free DNA molecules from the liver of an organism in a biological sample may include a pair of primers for amplifying liver-specific markers based on the methylation state of liver-specific markers. Liver-specific markers may include a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98% or 99% identity with SEQ ID NO: 1. In some cases, the pair of primers includes a primer including SEQ ID NO: 2 and a primer including SEQ ID NO: 3. In some cases, the composition also includes a detectably labeled probe including SEQ ID NO: 4 to detect liver-specific markers. In some cases, the composition also includes a primer including SEQ ID NO: 5 and a primer including SEQ ID NO: 6. In some cases, the composition also includes a detectably labeled probe including SEQ ID NO: 7 to detect liver-specific markers.
[0080] The compositions provided herein may include a pair of primers for amplifying liver-specific markers based on the methylation status of colon-specific markers. The colon-specific markers may include a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 8. In some cases, the pair of primers includes a primer including SEQ ID NO: 9 and a primer including SEQ ID NO: 10. In some cases, the composition also includes a detectably labeled probe including SEQ ID NO: 11 to detect colon-specific markers. In some cases, the composition also includes a primer including SEQ ID NO: 12 and a primer including SEQ ID NO: 13. In some cases, the composition also includes a detectably labeled probe including SEQ ID NO: 14 to detect colon-specific markers.
[0081] VI. Biological Samples The biological sample used in the method provided herein can include any tissue or material from a living or dead subject. The biological sample can be a free sample. The biological sample can include nucleic acid (e.g., DNA, e.g., genomic DNA or mitochondrial DNA or RNA) or a fragment thereof. The nucleic acid in the sample can be a free nucleic acid. The sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). The biological sample can be a body fluid, such as blood, plasma, serum, urine, vaginal fluid, hydrocele (e.g., testicular hydrocele), vaginal washing fluid, pleural effusion, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspiration fluid from different parts of the body (e.g., thyroid, breast), etc. Fecal samples can also be used. In various embodiments, most of the DNA in the biological sample (e.g., plasma sample obtained by centrifugation scheme) enriched with free DNA can be free (e.g., greater than 50%, 60%, 70%, 80%, 90%, 95% or 99% of the DNA can be free). Biological samples may be processed to physically disrupt tissue or cellular structures (e.g., centrifugation and / or cell lysis) to release intracellular components into a solution that may also contain enzymes, buffers, salts, detergents, etc., to prepare the sample for analysis.
[0082] The methods, compositions and systems provided herein can be used to analyze nucleic acid molecules in biological samples. Nucleic acid molecules can be cellular nucleic acid molecules, free nucleic acid molecules, or both. The free nucleic acids used by the methods provided herein can be nucleic acid molecules outside the cells of the biological sample. Free nucleic acid molecules can be present in various body fluids, for example, in blood, saliva, semen and urine. Due to cell death in various tissues, free DNA molecules can be produced, which may be caused by health conditions and / or diseases (for example, tumor invasion or growth, immune rejection after organ transplantation).
[0083] The free nucleic acid molecules used in the methods provided herein, for example, free DNA, can be present in plasma, urine, saliva or serum. Free DNA can occur naturally in the form of short fragments. Free DNA fragmentation can refer to the process in which high molecular weight DNA (such as DNA in the nucleus) is cut, broken or digested into short fragments when free DNA molecules are produced or released. In some cases, the methods, compositions and systems provided herein can be used to analyze cell nucleic acid molecules, for example, cell DNA from tumor tissue or cell DNA from white blood cells when the patient suffers from leukemia, lymphoma or myeloma. According to some examples of the present disclosure, samples taken from tumor tissue can be measured and analyzed.
[0084] VII. Subjects The methods, compositions and systems provided herein can be used to analyze samples from subjects, e.g., organisms, e.g., host organisms. The subject can be any human patient, such as a cancer patient, a patient at risk of cancer, or a patient with a family or personal history of cancer. In some cases, the subject is in a particular stage of cancer treatment. In some cases, the subject may have or be suspected of having cancer. In some cases, it is unknown whether the subject suffers from cancer.
[0085] The subject may have any type of cancer or tumor. In one example, the subject may have colon cancer or colorectal cancer. In another example, the subject may have colorectal cancer, or colon cancer and rectal cancer. In another example, the subject may have liver cancer, for example, hepatocellular carcinoma. Non-limiting examples of cancer may include, but are not limited to, adrenal cancer, anal cancer, basal cell carcinoma, bile duct cancer, bladder cancer, blood cancer, bone cancer, brain tumor, breast cancer, bronchial cancer, cardiovascular system cancer, cervical cancer, colon cancer, colorectal cancer, digestive system cancer, endocrine system cancer, endometrial cancer, esophageal cancer, eye cancer, gallbladder cancer, gastrointestinal tumors, hepatocellular carcinoma, kidney cancer, hematopoietic malignancies, laryngeal cancer, leukemia, liver cancer, lung cancer, lymphoma, melanoma, mesothelioma, Cancer of the muscular system, myelodysplastic syndrome (MDS), myeloma, nasal cavity cancer, nasopharyngeal cancer, nervous system cancer, lymphatic system cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, penis cancer, pituitary tumor, prostate cancer, rectal cancer, renal pelvis cancer, reproductive system cancer, respiratory system cancer, sarcoma, salivary gland cancer, skeletal system cancer, skin cancer, small intestine cancer, stomach cancer, testicular cancer, laryngeal cancer, thymus cancer, thyroid cancer, tumors, urinary system cancer, uterus cancer, vaginal cancer, or vulvar cancer. Lymphoma can be any type of lymphoma, including B cell lymphoma (e.g., diffuse large B cell lymphoma, follicular lymphoma, small lymphocytic lymphoma, mantle cell lymphoma, marginal zone B cell lymphoma, Burkitt's lymphoma, lymphoplasmacytic lymphoma, hairy cell leukemia or primary central nervous system lymphoma) or T cell lymphoma (e.g., precursor T lymphocyte lymphoma or peripheral T cell lymphoma). Leukemia can be any type of leukemia, including acute leukemia or chronic leukemia. The type of leukemia includes acute myeloid leukemia, chronic myeloid leukemia, acute lymphocytic leukemia, acute undifferentiated leukemia or chronic lymphocytic leukemia. In some cases, cancer patients do not suffer from a specific type of cancer. For example, in some cases, patients can suffer from non-breast cancer cancer.
[0086] Examples of cancer include cancers that cause solid tumors and cancers that do not cause solid tumors. In addition, any cancer mentioned herein can be a primary cancer (e.g., a cancer named after the part of the body in which it first began to grow) or a secondary or metastatic cancer (e.g., a cancer that originates from another part of the body).
[0087] The subject diagnosed by any of the methods described herein can be of any age, and can be an adult, infant, or child. In some cases, the subject is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99 years of age, or a range thereof (e.g., 2 to 20, 20 to 40, or 40 to 90). A particular class of patients that may benefit may be patients over 40 years of age. Another particular class of patients that may benefit may be pediatric patients. Furthermore, the subject diagnosed by any of the methods or compositions described herein can be male or female.
[0088] Any of the methods disclosed herein can also be performed on non-human subjects, such as laboratory or farm animals, or cell samples from organisms disclosed herein. Non-limiting examples of non-human subjects include dogs, goats, guinea pigs, hamsters, mice, pigs, non-human primates (e.g., gorillas, apes, orangutans, lemurs, or baboons), rats, sheep, cows, or zebrafish.
[0089] As described above, the subject methods, compositions and kits can be used for subjects at all stages of cancer treatment. The results of analyzing free nucleic acids in biological samples of subjects using the subject methods, compositions and kits can be used to guide the treatment plan of the subject. In some cases, drugs or therapies for treating or curing cancer in subjects may be needed. Exemplary treatment options may include chemotherapy, radiotherapy, surgical removal of tumor tissue, immunotherapy, targeted therapy, hormone therapy and stem cell therapy. In some cases, guidance on selecting different types of treatment options can be provided. In some non-limiting examples, the patient may have completed the treatment of the first cancer, for example, surgical removal of tumor tissue in the affected liver lobe, and the subject methods, compositions or kits can be used to perform routine monitoring tests on the patient to check whether there is recurrence or metastasis of liver cancer. In these cases, the test results can be used to guide whether the patient will need further treatment of cancer, and if liver cancer recurrence or metastasis to other tissues occurs, which treatment options can be adopted. In some cases, guidance on specific dosages or dosing regimens for treatment can be provided. For example, the amount of free nucleic acids from a certain tissue can be associated with the frequency / interval (e.g., daily, weekly, biweekly or monthly) of drug dosage or drug administration to be applied to the patient. In some cases, the results of the previous analysis can be used as the basis for evaluating and designing treatment plans and subsequent monitoring analyses.
[0090] VIII. Computer System Any method disclosed herein can be performed and / or controlled by one or more computer systems. In some examples, any step of the method disclosed herein can be performed and / or controlled by one or more computer systems as a whole, individually or sequentially. Any computer system mentioned herein can utilize any suitable number of subsystems. In some embodiments, the computer system includes a single computer device, wherein the subsystem can be a component of the computer device. In other embodiments, the computer system can include multiple computer devices, each of which is a subsystem with internal components. The computer system can include desktop computers and laptop computers, tablet computers, mobile phones and other mobile devices.
[0091] The subsystems may be interconnected via a system bus. Additional subsystems include a printer, a keyboard, a storage device, and a monitor coupled to a display adapter. Peripheral devices and input / output (I / O) devices coupled to an I / O controller may be connected to a computer system via any number of connections known in the art, such as an input / output (I / O) port (e.g., USB, FireWire®). For example, an I / O port or an external interface (e.g., Ethernet, Wi-Fi, etc.) may be used to connect a computer system to a wide area network, such as the Internet, a mouse input device, or a scanner. Interconnection via a system bus allows a central processor to communicate with each subsystem, and controls the execution of multiple instructions from a system memory or storage device (e.g., a fixed disk such as a hard disk or an optical disk), as well as information exchange between subsystems. System memory and / or storage devices may embody computer-readable media. Another subsystem is a data collection device, such as a camera, a microphone, an accelerometer, etc. Any data mentioned herein may be output from one component to another, and may be output to a user.
[0092] A computer system may include multiple identical components or subsystems connected together, for example, via external or internal interfaces. In some embodiments, a computer system, subsystem, or device may communicate via a network. In this case, a computer may be considered a client, and another computer may be considered a server, wherein each computer may be considered a part of the same computer system. Clients and servers may each include multiple systems, subsystems, or components.
[0093] The present disclosure provides a computer controlled system programmed to implement the method of the present disclosure. Fig.10 A computer system 101 is shown, which is programmed or otherwise configured to determine the absolute amount of free nucleic acid molecules from organism tissue as described herein. The computer system 101 can implement and / or adjust various aspects of the method provided in the present disclosure, for example, control the sequencing of nucleic acid molecules from biological samples, perform various steps of bioinformatics analysis of sequencing data (as described herein), integrate data collection, analysis and result reporting, and data management. The computer system 101 can be an electronic device of the user or a computer system remotely placed relative to the electronic device. The electronic device can be a mobile electronic device.
[0094] Computer system 101 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 105, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 101 also includes a memory or storage unit 110 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 115 (e.g., a hard disk), a communication interface 120 (e.g., a network adapter) for communicating with one or more other systems, and peripherals 125 (e.g., cache, other memory, data storage, and / or electronic display adapters). Memory 110, storage unit 115, interface 120, and peripherals 125 communicate with CPU 105 via a communication bus (solid lines) such as a motherboard. Storage unit 115 may be a data storage unit (or data repository) for storing data. Computer system 101 may be operatively coupled to a computer network ("network") 130 by means of communication interface 120. Network 130 may be the Internet, an internetwork and / or an extranet, or an intranet and / or an extranet in communication with the Internet. In some cases, network 130 is a telecommunications and / or data network. Network 130 may include one or more computer servers that may implement distributed computing, such as cloud computing. In some cases, network 130 may implement a peer-to-peer network with the help of computer system 101, which may enable devices coupled to computer system 101 to act as clients or servers.
[0095] The CPU 105 may execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory unit, such as the memory 110. The instructions may be directed to the CPU 105, which may then program or otherwise configure the CPU 105 to implement the methods of the present disclosure. Examples of operations performed by the CPU 105 may include fetching, decoding, executing, and writing back.
[0096] CPU 105 may be part of a circuit, such as an integrated circuit. One or more other components of system 101 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0097] Storage unit 115 can store files, such as drivers, libraries, and saved programs. Storage unit 115 can store user data, such as user preferences and user programs. In some cases, computer system 101 can include one or more additional data storage units external to computer system 101, such as located on a remote server that communicates with computer system 101 via an intranet or the Internet.
[0098] The computer system 101 can communicate with one or more remote computer systems via the network 130. For example, the computer system 101 can communicate with a user's remote computer system (e.g., a smart phone with an application installed that receives and displays sample analysis results sent from the computer system 101). Examples of remote computer systems include a personal computer (e.g., a portable PC), a tablet or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a phone, a smart phone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system 101 via the network 130.
[0099] The methods described herein may be implemented in the form of machine (e.g., computer processor) executable code stored on an electronic storage unit of computer system 101, such as memory 110 or electronic storage unit 115. Machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by processor 105. In some cases, the code may be retrieved from storage unit 115 and stored on memory 110 for access by processor 105. In some cases, electronic storage unit 115 may be excluded and machine executable instructions may be stored on memory 110.
[0100] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or may be compiled during runtime. The code may be provided in a programming language that may be selected so that the code can be executed in a precompiled or real-time compiled manner.
[0101] Aspects of the systems and methods provided herein, such as computer system 101, can be embodied in programming. Various aspects of the technology can be considered to be "products" or "articles of manufacture" that are carried on or embodied in the type of machine-readable media, usually in the form of machine (or processor) executable code and / or associated data. The machine executable code can be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media can include any or all tangible memories of a computer, processor, etc., or their related modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes communicate over the Internet or other various telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another computer or processor, for example, from a management server or host to a computer platform of an application server. Therefore, another medium that can carry software elements includes optical waves, radio waves, and electromagnetic waves, which are used, for example, over wired and optical fiber landline networks and over various air links on physical interfaces between local devices. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered as media that carry the software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable media" refer to any medium that participates in providing instructions to a processor for execution.
[0102] Thus, machine-readable media, such as computer executable code, may take many forms, including, but not limited to, tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer, and the like, such as may be used to implement the databases shown in the accompanying drawings, and the like. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier transmission media may take the form of electrical or electromagnetic signals or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example: a floppy disk, a diskette, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, a punched card tape, any other physical storage medium with a pattern of holes, a RAM, a ROM, a PROM and an EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0103] The computer system 101 may include or communicate with an electronic display 135 that includes a user interface (UI) 140 to provide, for example, results of sample analysis, such as, but not limited to, graphically displaying relative and / or absolute amounts of free nucleic acid from different tissues, control or reference amounts of free nucleic acid from certain tissues, comparisons between detected amounts and reference amounts, and readouts of the presence or absence of cancer metastasis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0104] The methods and systems of the present disclosure may be implemented by one or more algorithms. When executed by the central processing unit 105, the algorithms may be implemented by software. For example, the algorithms may control the sequencing of nucleic acid molecules from a sample, directly collect sequencing data, analyze sequencing data, or determine pathological classification based on analysis of sequencing data.
[0105] In some cases, such as Fig.11 As shown, a sample 202 can be obtained from a subject 201, such as a human subject. One or more methods as described herein can be performed on the sample 202, such as performing an assay. In some cases, the assay can include hybridization, amplification, sequencing, labeling, epigenetic modified bases, or any combination thereof. One or more results of the method can be input into a processor 204. One or more input parameters, such as sample identification, subject identification, sample type, reference, or other information can be input into the processor 204. One or more metrics of the assay can be input into the processor 204 so that the processor can generate a result, such as a pathological classification (e.g., diagnosis) or a treatment recommendation. The processor can send the result, input parameter, metric, reference, or any combination thereof to a display 205, such as a visual display or a graphical user interface. The processor 204 can (i) send the result, input parameter, metric, or any combination thereof to a server 207, (ii) receive the result, input parameter, metric, or any combination thereof from the server 207, (iii) or a combination thereof.
[0106] Aspects of the present disclosure may be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or a field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor includes a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will know and understand other ways and / or methods of implementing the embodiments described herein using hardware and combinations of hardware and software.
[0107] Any software component or functionality described in this application can be implemented as software code to be executed by a processor using any suitable computer language (e.g., Java, C, C++, C#, Objective-C, Swift) or scripting language (e.g., Perl or Python), using, for example, traditional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disk (CD) or DVD (digital versatile disk), flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.
[0108] Such programs may also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks that conform to various protocols including the Internet. In this way, a computer-readable medium may be created using a data signal encoded by such a program. Computer-readable media encoded with program code may be packaged with compatible devices or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard drive, a CD, or an entire computer system), and may exist on or within different computer products within a system or network. A computer system may include a monitor, a printer, or other suitable display for providing any results mentioned herein to a user.
[0109] Any method described herein can be performed in whole or in part with a computer system including one or more processors, which can be configured to perform steps. Therefore, embodiments can be directed to a computer system configured to perform the steps of any method described herein, wherein different components perform corresponding steps or corresponding step groups. Although presented in numbered steps, the steps of the method herein can be performed simultaneously or in different orders. In addition, a part of these steps can be used together with a part of other steps of other methods. In addition, all or part of the steps can be optional. In addition, any step of any method can be performed using a module, unit, circuit or other method for performing these steps.
[0110] Example The following examples further illustrate the described embodiments without limiting the scope of the present disclosure.
[0111] Example 1. Materials and methods This example describes various methods used in Examples 2-5.
[0112] Subjects The demographic data of the recruited subjects are listed in Table 2. All recruited subjects signed written consent.
[0113] Table 2. Demographic data of subjects analyzed in the study
[0114] Sample preparation For each subject, 10 mL of peripheral blood was collected into tubes containing EDTA. Blood samples were processed within 6 h after blood collection to separate plasma and buffy coat. DNA was extracted from plasma using the QIAamp DSP DNA Mini Kit (Qiagen) following the manufacturer's protocol. DNA extracted from 2 mL to 4 mL of plasma was subjected to two rounds of bisulfite treatment using the Epitect Plus Bisulfite Kit (Qiagen). Bisulfite-converted DNA was eluted in 50 μL of water for downstream analysis.
[0115] Identification of liver-specific and colon-specific methylation markers The methylation profile of the tissue of interest (i.e., liver or colon) was compared with that of other blood cells and tissues to mine tissue-specific methylation markers. The methylation profiles of different cell types were retrieved from the databases of the Roadmap Epigenome Project for lung, esophagus, small intestine, colon, pancreas, bladder, heart, and liver, and from the Blueprint Project for erythroblasts, neutrophils, B lymphocytes, and T lymphocytes.
[0116] The following standards for methylation markers were established.
[0117] 1. A CpG site was defined as hypermethylated in the target tissue if its methylation density was >50% in the target tissue (i.e., liver or colon) and <20% in other blood cells and tissues.
[0118] 2. Fragments with at least 5 hypermethylated CpG sites will be within differentially methylated regions (DMRs) to improve the signal-to-noise ratio of the latter and the analytical specificity of the methylation markers.
[0119] 3. DMR should be shorter than 166 bp because most circulating DNA molecules are short fragments and the peak size is 166 bp.
[0120] DMR of liver and colon Using the above criteria, one liver-specific marker and one colon-specific marker were identified. The liver-specific DMR is located in the exonic region of the protein tyrosine kinase 2β (PTK2B) gene on chromosome 8. The eight CpG sites within the liver-specific DMR are hypermethylated in the liver and hypomethylated in other tissues and blood cells ( Figure 7 ). PTK2B The gene is located on chromosome 8 and the genomic coordinates of the CpG site are shown in Figure 7 on the X-axis of . All eight CpG sites located within the DMR (region between the two vertical dashed lines) are hypermethylated for the liver compared to other tissues. The yellow highlighted region contains three CpG sites within the fluorescent probes assayed by digital PCR (see Table 3). The other CpG sites on each side of the highlighted region within the DMR are covered by the primers assayed by digital PCR. The colon-specific DMR is within the exonic region of the Sestrin 3 (SESN3) gene on chromosome 11. All six CpG sites within the colon-specific DMR are hypermethylated in colon tissue ( Figure 8 ), while it is hypomethylated in other tissues. SESN3 is located on chromosome 11, and the genomic coordinates of the CpG site are shown in Figure 8 All six CpG sites located within the DMR (the region between the two vertical dashed lines) are hypermethylated for the colon compared to other tissues. The three CpG sites within the yellow highlighted region are covered by fluorescent probes for the colon-specific methylation assay (see Table 3). The other CpG sites on each side of the highlighted region within the DMR are covered by primers for the digital PCR assay. Figure 7 and Figure 8 Individual results for liver, bladder, esophagus, heart, lung, pancreas, and small intestine are not shown8, and their averages are presented as "other tissues."
[0121] Design and performance of droplet digital PCR Two droplet digital PCR assays were developed to quantify methylated DNA molecules and non-methylated DNA molecules by each of the liver-specific methylation markers and the colon-specific methylation markers. The sequences of the primers and probes used for the assays are listed in Table 3 (the underlined nucleotides in the primers and probes are differentially methylated cytosines at CpG sites). The two droplet digital PCR assays can quantify methylation (from target tissues) and non-methylation (from non-target tissues) using probes labeled with FAM and VIC, respectively. The liver-specific marker is the PTK2B gene marker site (chr8: 27,183,116-27,183,176), and the colon-specific marker is the SESN3 gene marker site (chr11: 94,965,508-94,965,567).
[0122] For each sample, duplicate digital PCR assays were performed. A total volume of 20 μL of reaction mixture was prepared containing 8 μL of bisulfite converted DNA, forward and reverse primers at a final concentration of 450 nM each, a non-methylation-specific probe at a final concentration of 250 nM, and a methylation-specific probe for colon assays at a final concentration of 350 nM (liver assay) or 250 nM (colon assay). The reaction mixture was droplet generated before PCR reactions using a BioRad QX200 ddPCR droplet generator. Universal methylated DNA (CpGenome human methylated DNA from EMD Millipore) and universal non-methylated DNA (EpiTect non-methylated human control DNA from Qiagen) were run on each plate as positive and negative controls. The temperature profile was: 95 o C for 10 minutes; then 94 o 45 cycles of C for 15 sec; and 60 o C (liver test) or 56 o C (colonometry) for 1 minute; and finally at 98 o C for 10 minutes. After PCR was completed, the droplets of each sample were analyzed by a QX200 droplet reader, and the results were interpreted using QuantaSoft (version 1.7) software. The cutoff value of the positive fluorescent signal was determined with reference to the control. The number of methylated DNA sequences and unmethylated DNA sequences in each sample was calculated using the combined counts from duplicate wells followed by Poisson correction. The concentration of methylated DNA sequences or unmethylated DNA sequences in plasma was calculated as follows:
[0123] Where Cr represents the concentration of the target molecule in plasma (i.e., methylated DNA sequence or unmethylated DNA sequence), P represents the number of droplets containing the amplified signal of the target molecule (methylated DNA sequence or unmethylated DNA sequence), R represents the total number of droplets analyzed (with and without amplified signals), V d represents the average volume of the droplets (i.e., 0.9×10 -3 µL), Ve represents the plasma volume used for the experiment (i.e., 320 µL in the current example).
[0124] Table 3. Oligonucleotide sequences used for digital PCR assays
[0125] Analyzing DNA from different types of samples Formalin-fixed paraffin-embedded (FFPE) samples of 10 tissues (e.g., liver, lung, esophagus, stomach, small intestine, colon, pancreas, bladder, heart, and brain) were retrieved from the Department of Anatomy and Cytopathology, Prince of Wales Hospital, Hong Kong. These tissues were confirmed to be normal by histological examination. Buffy coat samples were collected from healthy subjects. DNA was extracted from FFPE tissues using the QIAamp DNAMini Kit (Qiagen). DNA from buffy coats was extracted using the QIAamp DNA Blood Mini Kit (Qiagen). 1 μg of cellular DNA was used for bisulfite conversion. The converted DNA was eluted in 20 μL of water and then diluted 50-fold for downstream analysis.
[0126] Measurement of donor-derived DNA in plasma of liver transplant recipients DNA extracted from liver tissue of donors and buffy coat of recipients was analyzed using the Illumina iScan system to determine the genotype information of donors and recipients. DNA extracted from 4 mL of plasma from each recipient was used for sequencing library preparation. Plasma DNA sequencing libraries were prepared with the KAPA Library Preparation Kit (KAPA Biosystems) according to the manufacturer's instructions. The indexed libraries were then multiplexed and sequenced using the Illumina HiSeq 2500 platform (75 × 2 cycles). At least 20 million paired-end reads were obtained for each sample. Paired-end reads were aligned to the unmasked repeat human reference genome (GRCh 37 / hg 19) using the Short Oligonucleotide Alignment Program 2 (SOAP2). Only paired-end reads were included, with both ends aligned to the same chromosome in the correct orientation and to a single position in the human genome. Paired-end reads with an insert size ≤ 600 bp were retrieved for analysis. If more than one pair of reads mapped to the same genomic location (i.e., duplicate reads), only one pair of reads was retained for subsequent analysis. A maximum of two nucleotide mismatches were allowed in any one member of the paired-end reads. The fractional concentration of circulating donor-specific DNA was determined by counting sequencing reads with single-nucleotide polymorphism (SNP) alleles that were homozygous in the recipient and heterozygous in the donor.
[0127] Example 2. Tissue Specificity of Liver and Colon Markers For liver-specific markers and colon-specific markers, DNA molecules derived from target tissues will be hypermethylated, while DNA molecules from non-target tissues will be hypomethylated. Therefore, in the liver assay, the percentage of total molecules is expressed as methylated as L%, and in the colon assay, the percentage of total molecules is expressed as methylated as C%. In order to confirm the specificity of liver markers and colon markers, DNA extracted from buffy coat samples and 10 normal tissues was analyzed using these two digital PCR assay groups. For each type of tissue, 4 samples from different individuals were included.
[0128] The mean L% for liver tissue was 67% (range: 57%-76%), and the mean L% for other tissue types was 0.6% (range: 0.0%-2.2%). The results for each tissue type are summarized in Table 4. These results demonstrate that the liver assay is able to specifically detect liver-derived DNA.
[0129] Table 4. Average fractional concentrations of liver-derived DNA (L%) and colon-derived DNA (C%)
[0130] The mean C% for colon tissue was 22% (range: 17%-33%). The mean C% for all other tissues was 1.2% (range: 0.1%-4.1%), indicating that the methylated sequences were specifically colonic in origin. The relatively low C% in colon tissue may be due to the heterogeneous cellular composition of colon tissue. The relatively low C% in colon tissue does not significantly hinder its clinical use when the same assay is used to compare levels in subjects with different disease states.
[0131] Example 3. Plasma Concentrations of Liver-Derived DNA in Liver Transplant Recipients The quantitative accuracy of the liver-specific assay was validated by analysis of plasma from liver transplant recipients. In these subjects, the fractional concentration of DNA from the transplanted liver could be accurately determined using next-generation sequencing based on the proportion of plasma DNA molecules carrying donor-specific alleles. Fourteen plasma samples collected from 13 patients who had received liver transplants were analyzed by liver-specific methylation markers and sequencing. A linear positive correlation was observed between the concentrations determined by the two methods (R = 0.99, P < 0.0001, Pearson correlation, Figure 2 ), indicating that liver-specific methylation markers can accurately reflect the concentration of liver-derived DNA in plasma.
[0132] As demonstrated herein, the percentage contribution of liver DNA concentration measured by liver-specific methylation markers correlates well with the results of donor-specific allele-based measurements. These results confirm the accuracy of liver-specific markers in reflecting liver-derived DNA concentrations in plasma.
[0133] Example 4. Increased liver-derived DNA in plasma of HCC patients The absolute and fractional concentrations of liver-derived DNA were determined by digital PCR targeting sequences of liver-specific methylation patterns in 40 HCC patients, 9 cirrhotic patients, 20 chronic HBV carriers, and 30 healthy subjects.
[0134] The median concentrations of liver-derived methylated sequences in healthy subjects, chronic HBV carriers, patients with cirrhosis, and patients with HCC were 40 copies / mL (interquartile range (IQR): 18-86), 122 copies / mL (IQR: 47-185), 118 copies / mL (IQR: 86-159), and 487 (IQR: 138-1151), respectively. Figure 3A ). The concentrations among the four groups were significantly different (P<0.001, Kruskal Wallis test). In post hoc analysis, the plasma concentration of liver-derived DNA in HCC patients was significantly higher than that in healthy subjects (P<0.001, Dunn's test) and chronic HBV carriers (P = 0.015, Dunn's test), but not higher than that in the cirrhosis group (P = 0.248, Dunn's test). There was no statistical difference in the concentrations among healthy subjects, chronic HBV carriers, and cirrhosis patients (P>0.05, Dunn's test).
[0135] The median fractional concentrations of liver-derived DNA in plasma were 1.4% (IQR: 0.94%-3.2%), 4.6% (IQR: 1.7%- 6.0%), 3.0% (interquartile range: 1.8%-7.3%), and 9.4% (IQR: 4.1%-16.0%) for healthy subjects, chronic HBV carriers, patients with cirrhosis, and patients with HCC, respectively. Figure 3B ). The graded concentrations among the four groups were significantly different (P < 0.001, Kruskal Wallis test). In post hoc analysis, the plasma concentration of liver-derived DNA was significantly higher in HCC patients than in healthy subjects (P < 0.001, Dunn's test), but not in chronic HBV carriers (P = 0.129, Dunn's test) and patients with cirrhosis (P = 0.592, Dunn's test). There were no statistical differences in the concentrations among healthy subjects, chronic HBV carriers, and patients with cirrhosis.
[0136] These results indicate that both the absolute and fractional concentrations of liver-derived DNA in plasma can distinguish HCC patients from non-HCC subjects (including healthy subjects, chronic HBV carriers, and patients with cirrhosis). To further determine whether the absolute or fractional concentrations are better for distinguishing HCC subjects from non-HCC subjects, a receiver operating characteristic (ROC) curve analysis was performed ( Fig. 9 The areas under the curve (AUC) for absolute concentration and fractional concentration were 0.82 and 0.78, respectively. The difference in AUC was statistically significant (P = 0.022, Delong test).
[0137] We further analyzed the correlation between the concentrations of liver-derived DNA in the plasma of HCC patients (absolute and graded) and the maximum size of the tumor (determined by computed tomography or after tumor resection). Interestingly, the maximum size of the tumor showed a stronger positive correlation with the absolute concentration (R = 0.74, P < 0.0001, Spearman correlation) than with the graded concentration (R = 0.56, P = 0.0002, Spearman correlation). Figure 4A and Figure 4B ). The concentration of liver-derived DNA was positively correlated with the maximum size of HCC patient tumors, suggesting that the amount of DNA released from the liver will reflect tumor burden.
[0138] Example 5. Analysis of liver-derived DNA and colon-derived DNA in CRC patients with and without liver metastasis The plasma concentrations of liver-derived DNA and colon-derived DNA were measured in 30 healthy subjects, 35 CRC patients without liver metastasis, and 27 CRC patients with liver metastasis. The median plasma concentrations of colon-derived DNA in the three groups were 0 copies / mL (IQR: 0-0), 4 copies / mL (IQR: 0-31), and 138 copies / mL (IQR: 0-6850), respectively ( Figure 5A ). The concentrations were significantly different among the three groups (P < 0.001, Kruskal Wallis test). In post hoc analysis, the concentrations in CRC patients with liver metastases and CRC patients without liver metastases were significantly higher than those in healthy subjects (P < 0.001 and P = 0.042, respectively, Dunn's test). There was no statistically significant difference between CRC patients with and without liver metastases (P = 0.079, Dunn's test).
[0139] The median fractional concentrations of colon-derived DNA in plasma were 0% (IQR: 0%-0%), 0.09% (IQR: 0%-1.1%), and 0.84% (IQR: 0%-49.5%) for healthy control subjects, CRC patients without liver metastasis, and CRC patients with liver metastasis, respectively ( Figure 5B ). The fractional concentrations were significantly different among the three groups (P < 0.001, Kruskal Wallis test). In post hoc analysis, CRC patients with liver metastases and CRC patients without liver metastases had significantly higher concentrations than healthy subjects (P < 0.001 and P = 0.041, respectively, Dunn's test). There was no statistically significant difference between CRC patients with and without liver metastases (P = 0.084, Dunn's test).
[0140] The median concentrations of liver-derived DNA in plasma were 40 copies / mL (IQR: 18-86), 23 copies / mL (IQR: 13-108), and 233 copies / mL (IQR: 56-2290) for healthy control subjects, CRC patients without liver metastasis, and CRC patients with liver metastasis, respectively ( Figure 5C ). The concentrations of the fractions were significantly different among the three groups (P < 0.001; Kruskal-Wallis test). In post hoc analysis, the concentrations were significantly higher in CRC patients with liver metastases than in CRC patients without liver metastases and in healthy controls (P < 0.001 and P < 0.001, respectively; Dunn's test). Interestingly, there was no significant difference between patients without liver metastases and healthy controls (P = 1.0; Dunn's test).
[0141] The median fractional concentrations of liver-derived DNA in plasma were 0.8% (IQR: 0.3%-2.8%), 1.4% (IQR: 0.9%-3.3%), and 3.1% (IQR: 1.5%-5.3%) for healthy control subjects, CRC patients without liver metastasis, and CRC patients with liver metastasis, respectively ( Figure 5D The concentrations of graded α were significantly different between the three groups (P < 0.001, Kruskal-Wallis test). In post hoc analysis, the concentrations of CRC patients with liver metastases were significantly higher than those of CRC patients without liver metastases (P < 0.003, Dunn's test). The concentrations of graded α in healthy subjects were not statistically different from those in CRC patients with liver metastases and those in CRC patients without liver metastases (P = 0.114 and P = 0.717, respectively).
[0142] Because significant differences in the absolute and fractionated concentrations of liver-derived DNA and colon-derived DNA in plasma were observed between CRC patients with and without liver metastases, ROC curve analysis was used to determine which parameter was most useful for distinguishing the two groups. The AUCs for the absolute and fractionated concentrations of liver-derived DNA were 0.85 and 0.75, respectively (P = 0.01, Delong's test), and the AUCs for the absolute and fractionated concentrations of colon-derived DNA were 0.69 and 0.69, respectively (P = 0.75, Delong's test) ( Figure 6 ).
[0143] In ROC analysis, the absolute concentration of liver-derived DNA was superior to the fractional concentration in distinguishing CRC patients with and without liver metastases (AUC: 0.85 vs. 0.75, P = 0.01, Figure 6 Without being bound by theory, a possible explanation is that in some patients with CRC metastases to the liver, the absolute concentrations of both liver-derived and colon-derived DNA increase. In some patients, despite an increase in the absolute concentration of liver-derived DNA, the fractionated concentration of the liver remains the same or decreases due to a greater increase in colon-derived DNA. Similarly, it was also shown that the absolute concentration of liver-derived DNA correlates better with tumor size than the fractionated concentration (R = 0.74 vs 0.56, Spearman correlation, Figure 4A-4B ).
[0144] Although preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Those skilled in the art will now expect multiple changes, modifications and substitutions without departing from the present disclosure. It should be understood that various alternatives of the embodiments of the present disclosure described herein can be used to implement the present disclosure. The appended claims are intended to define the scope of the present disclosure, and thus encompass methods and structures and their equivalents within the scope of these claims.
Claims
1. A method for measuring cytosine methylation in one or more differentially methylated regions (DMRs), the method comprising: (a) obtaining free DNA molecules from a first biological sample of a subject; and (b) assaying the free DNA molecules to measure, for each of the one or more DMRs, an amount of the free DNA molecules that include a first methylation state of a target sequence in the DMR; wherein (i) the one or more DMRs include a target sequence of one or both of Sestrin 3 (SESN3) and protein tyrosine kinase 2β (PTK2B); and (ii) the target sequence of each of the one or more DMRs includes one or more CpG sites.
2. The method of claim 1, wherein the determining comprises hybridizing the free DNA molecule comprising the target sequence with a probe.
3. The method of claim 1, wherein the determining comprises amplifying the free DNA molecules using one or more pairs of primers.
4. The method of claim 1, wherein the determining comprises converting non-methylated cytosine residues in the free DNA molecules into uracil via bisulfite.
5. The method of claim 1, wherein the determining comprises performing methylation-aware sequencing on cell-free DNA molecules from the first biological sample.
6. The method of claim 1, wherein the target sequence comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10 CpG methylation sites.
7. The method of claim 1, wherein the first methylation state comprises the methylation density of a single site within the target sequence, the distribution of methylated / unmethylated sites over a continuous region within the target sequence, or the methylation pattern or level of each single methylated site within the target sequence.
8. The method of claim 1, wherein the target sequence comprises a higher methylation density in a first tissue of the subject than in a second tissue of the subject.
9. The method of claim 8, wherein the target sequence comprises a methylation density greater than 50% in the first tissue.
10. The method of claim 1, wherein the target sequence comprises a polynucleotide sequence of PTK2B having at least 60% identity to SEQ ID NO:
1.
11. The method of claim 1, wherein the determining comprises (i) amplifying using primers comprising SEQ ID NO: 2, primers comprising SEQ ID NO: 3, or both, or (ii) using a detectably labeled probe comprising SEQ ID NO: 4 for detecting the target sequence.
12. The method of claim 1, wherein the amplifying comprises (i) amplifying using a primer comprising SEQ ID NO: 5, a primer comprising SEQ ID NO: 6, or both, or (ii) using a detectably labeled probe comprising SEQ ID NO: 7 for detecting the target sequence.
13. The method of claim 1, wherein the target sequence comprises a polynucleotide sequence of SESN3 having at least 60% identity to SEQ ID NO:
8.
14. The method of claim 1, wherein the determining comprises (i) amplifying using primers comprising SEQ ID NO: 9, primers comprising SEQ ID NO: 10, or both, or (ii) using a detectably labeled probe comprising SEQ ID NO: 11 for detecting the target sequence.
15. The method of claim 1, wherein the determining comprises (i) amplifying using primers comprising SEQ ID NO: 12, primers comprising SEQ ID NO: 13, or both, or (ii) using a detectably labeled probe comprising SEQ ID NO: 14 for detecting the target sequence.
16. The method of claim 1, wherein the target sequence of each of the one or more DMRs comprises an exon sequence.
17. The method of claim 1, wherein the one or more DMRs comprise a target sequence comprising one or more CpG sites of SEQ ID NO:
1.
18. The method of claim 1, wherein the one or more DMRs comprise a target sequence comprising one or more CpG sites of SEQ ID NO:
8.
19. The method of claim 1, wherein the measured amount of the free DNA molecules comprising the first methylation state is an absolute amount.
20. The method of claim 1, wherein the free DNA molecules comprising the first methylation state of the target sequence in the DMR comprise DNA molecules released from cancer metastasis sites.
21. A composition for determining the amount of free DNA molecules in a biological sample from the liver of an organism, comprising a pair of primers for amplifying a liver-specific marker based on its methylation level, wherein the liver-specific marker comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO:
1.
22. The composition of claim 21, wherein the pair of primers comprises a primer comprising SEQ ID NO: 2 and a primer comprising SEQ ID NO:
3.
23. The composition of claim 21 or 22, further comprising a detectably labeled probe comprising SEQ ID NO: 4 for detecting the liver-specific marker.
24. The composition according to any one of claims 21-23, further comprising a primer comprising SEQ ID NO: 5 and a primer comprising SEQ ID NO:
6.
25. The composition of any one of claims 21-24, further comprising a detectably labeled probe comprising SEQ ID NO: 7 for detecting the liver-specific marker.
26. A composition for determining the amount of free DNA molecules in a biological sample from the colon of an organism, comprising a pair of primers for amplifying a colon-specific marker based on its methylation level, wherein the colon-specific marker comprises a polynucleotide sequence having at least 60%, 70%, 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO:
8.
27. The composition of claim 26, wherein the pair of primers comprises a primer comprising SEQ ID NO: 9 and a primer comprising SEQ ID NO:
10.
28. The composition of claim 26 or 27, further comprising a detectably labeled probe comprising SEQ ID NO: 11 for detecting the colon-specific marker.
29. The composition of any one of claims 26-28, further comprising a primer comprising SEQ ID NO: 12 and a primer comprising SEQ ID NO:
13.
30. The composition of any one of claims 26-29, further comprising a detectably labeled probe comprising SEQ ID NO: 14 for detecting the colon-specific marker.
Citation Information
Patent Citations
Analysis of fragmentation patterns of cell-free DNA
US10453556B2
Methylation pattern analysis of tissues in a DNA mixture
US11062789B2
Detection of genetic or molecular aberrations associated with cancer
US8741811B2