Non-invasive determination of methylome of fetus or tumor from plasma

JP2025026936A5Pending Publication Date: 2025-07-25THE CHINESE UNIVERSITY OF HONG KONG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024200780
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-06-03
Filing Date
2024-11-18
Publication Date
2025-07-25

Smart Images

  • Figure 00000132_0000
    Figure 00000132_0000
  • Figure 00000132_0001
    Figure 00000132_0001
  • Figure 00000133_0000
    Figure 00000133_0000
Patent Text Reader

Abstract

To provide methods for determining DNA methylation patterns (methylomes).SOLUTION: A methylation profile can be deduced for fetal / tumor tissue based on a comparison of plasma methylation (or other sample with cell-free DNA) to a methylation profile of the mother / patient. A methylation profile can be determined for fetal / tumor tissue using tissue-specific alleles to identify DNA from the fetus / tumor when the sample has a mixture of DNA. A methylation profile can be used to determine copy number variations in genome of a fetus / tumor. Methylation markers for a fetus have been identified via various techniques. The methylation profile can be determined by determining a size parameter of a size distribution of DNA fragments, where reference values for the size parameter can be used to determine methylation levels. Additionally, a methylation level can be used to determine a level of cancer.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Field The present disclosure relates generally to determining the methylation pattern (methylome) of DNA, and more particularly to analyzing a biological sample (e.g., plasma) containing a mixture of DNA from separate genomes (e.g., from fetus and mother, or from tumor and normal cells) to determine the methylation pattern (methylome) of a small number of genomes. Uses of the determined methylome are also described. [Background technology]

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a PCT application claiming priority to U.S. Provisional Patent Application No. 61 / 830,571, entitled "Tumor Detection In Plasma Using Methylation Status And Copy Number," filed June 3, 2013, and to U.S. Provisional Patent Application No. 13 / 842,209, entitled "Non-Invasive Determination Of Methylome Of Fetus Or Tumor From Plasma," filed March 15, 2013, which claims the benefit of U.S. Provisional Patent Application No. 61 / 703,512, entitled "Method Of Determining The Whole Genome DNA Methylation Status Of The Placenta By Massively Parallel Sequencing Of Maternal Plasma," filed September 20, 2012, the entireties of which are incorporated herein by reference for all purposes.

[0003] background Embryonic and fetal development is a complex process that involves a series of highly regulated genetic and epigenetic phenomena. Cancer development is also a complex process that typically involves multiple genetic and epigenetic processes. Abnormalities in the epigenetic control of developmental processes have been implicated in infertility, spontaneous abortion, intrauterine growth abnormalities, and postnatal outcomes. DNA methylation is one of the most commonly studied epigenetic mechanisms. DNA methylation most commonly occurs in the context of the addition of a methyl group to the 5' carbon of cytosine residues between CpG dinucleotides. Cytosine methylation adds an additional layer of control to gene transcription and DNA function. For example, hypermethylation of gene promoters rich in CpG dinucleotides, termed CpG islands, is commonly associated with repression of gene function.

[0004] Despite the important role of epigenetic mechanisms in orchestrating developmental processes, human embryonic and fetal tissues are not readily available for analysis (as are tumors). Dynamic changes in such epigenetic processes in health and disease during the human prenatal period are virtually impossible to study. Extraembryonic tissues, especially the placenta, have become one of the primary tools for such investigations, as they can be obtained as part of prenatal diagnostic procedures or after birth. However, such tissues require invasive procedures.

[0005] The DNA methylation profile of the human placenta has intrigued researchers for decades. Human placenta displays numerous unique physiological features that are related to DNA methylation. At the global level, placental tissue is hypomethylated compared to most somatic tissues. At the gene level, the methylation status of given genomic loci is a specific feature unique to placental tissue. Both global and locus-specific methylation profiles show gestational age-dependent changes. Imprinted genes, i.e. genes whose expression is dependent on the parental origin of the allele, play important functions in the placenta. The placenta has been described as a pseudomalignant and hypermethylation of several tumor suppressor genes has been observed.

[0006] Studies of DNA methylation profiles in placental tissues have provided insight into the pathophysiology of pregnancy- or development-related disorders, e.g., preeclampsia and intrauterine growth restriction. Disturbances in genomic imprinting have been linked to developmental disorders, such as Prader-Willi syndrome and Angelman syndrome. Altered profiles of genomic imprinting and global DNA methylation in placental and fetal tissues have been observed in pregnancies resulting from assisted reproductive techniques (H Hiura et al.2012 Hum Reprod;27:2541-2548). Several environmental factors, such as maternal smoking (KE Haworth et al. 2013 Epigenomics;5:37-49), maternal dietary factors (X Jiang et al. 2012 FASEB J;26:3563-3574), and maternal metabolic conditions such as diabetes (N Hajj et al.,Diabetes.doi:10.2337 / db12-0289), have been associated with epigenetic abnormalities in the offspring. Summary of the Invention

[0007] overview Despite decades of efforts, no practical tools have been available to study fetal or tumor methylomes and to monitor their dynamic changes throughout pregnancy or the course of diseases such as malignancies. It is therefore desirable to provide methods to non-invasively analyze all or part of the fetal and tumor methylomes.

[0008] The embodiments provide systems, methods, and devices for determining and using the methylation profile of various tissues and samples. Examples are provided. The methylation profile can be estimated for fetal / tumor tissue based on a comparison of plasma methylation (or other samples with cell-free DNA, e.g., urine, saliva, genital effluent) with the methylation profile of the mother / patient. If the sample has a mixture of DNA, tissue-specific alleles can be used to identify DNA from the fetus / tumor and determine the methylation profile for that fetal / tumor tissue. The methylation profile can be used to determine copy number variations in the genome of the fetus / tumor. Methylation markers for the fetus have been identified by various techniques. The methylation profile can be determined by determining the size parameters of the size distribution of DNA fragments, and the methylation level can be determined using a reference value for the size parameter.

[0009] In addition, methylation levels can be used to determine the level of cancer. In the context of cancer, measuring changes in the methylome in plasma allows for cancer detection (e.g., for screening purposes), monitoring (e.g., detecting response following anti-cancer treatment and detecting cancer recurrence), and prognosis (e.g., for measuring the burden of cancer cells in the body, or for staging purposes, or for estimating the likelihood of death from disease or disease progression or metastatic processes).

[0010] A better understanding of the nature and advantages of embodiments of the present invention can be obtained with reference to the following detailed description and accompanying drawings. [Brief description of the drawings]

[0011] [Figure 1A-1] 1 shows Table 100 results of sequencing maternal blood, placenta, and maternal plasma according to an embodiment of the present invention. [Figure 1A-2] 1 shows Table 100 results of sequencing maternal blood, placenta, and maternal plasma according to an embodiment of the present invention. [Figure 1B]1 shows methylation density within a 1 Mb window of samples sequenced according to an embodiment of the present invention. [Figure 2A] Plots of β values ​​against methylation index are shown for: (A) maternal blood cells, (B) chorionic villus samples, and (C) term placenta tissue. [Figure 2B] Plots of β values ​​against methylation index are shown for: (A) maternal blood cells, (B) chorionic villus samples, and (C) term placenta tissue. [Figure 2C] Plots of β values ​​against methylation index are shown for: (A) maternal blood cells, (B) chorionic villus samples, and (C) term placenta tissue. [Figure 3A] Bar graphs showing the percentage of methylated CpG sites in plasma and blood cells from adult males and non-pregnant adult females: (A) autosomes, (B) X chromosome. [Figure 3B] Bar graphs showing the percentage of methylated CpG sites in plasma and blood cells from adult males and non-pregnant adult females: (A) autosomes, (B) X chromosome. [Figure 4A] Methylation density plots of corresponding loci in blood cell DNA and plasma DNA are shown: (A) non-pregnant adult female, (B) adult male. [Figure 4B] Methylation density plots of corresponding loci in blood cell DNA and plasma DNA are shown: (A) non-pregnant adult female, (B) adult male. [Figure 5A] Bar graphs of the percentage of methylated CpG sites among samples taken from pregnant women: (A) autosomes, (B) X chromosome. [Figure 5B] Bar graphs of the percentage of methylated CpG sites among samples taken from pregnant women: (A) autosomes, (B) X chromosome. [Figure 6] 1 shows a bar graph of methylation levels of different repeat classes of the human genome in maternal blood, placenta and maternal plasma. [Figure 7A] Circos plot 700 for the first batch of samples is shown. [Figure 7B] Circos plot 750 for third-phase samples is shown. [Figure 8A] 1 shows a plot of the comparative methylation density of maternal plasma DNA versus genomic tissue DNA for CpG sites surrounding informative single nucleotide polymorphisms. [Figure 8B] 1 shows a plot of the comparative methylation density of maternal plasma DNA versus genomic tissue DNA for CpG sites surrounding informative single nucleotide polymorphisms. [Figure 8C] 1 shows a plot of the comparative methylation density of maternal plasma DNA versus genomic tissue DNA for CpG sites surrounding informative single nucleotide polymorphisms. [Figure 8D] 1 shows a plot of the comparative methylation density of maternal plasma DNA versus genomic tissue DNA for CpG sites surrounding informative single nucleotide polymorphisms. [Figure 9] 9 is a flow chart illustrating a method 900 for determining a first methylation profile from a biological sample of an organism, according to an embodiment of the present invention. [Figure 10] 1 is a flow chart illustrating a method 1000 of determining a first methylation profile from a biological sample of an organism, according to an embodiment of the present invention. [Figure 11A] 13 shows graphs of the performance of a predictive algorithm using maternal plasma data and fractional fetal DNA concentration, according to an embodiment of the present invention. [Figure 11B] 13 shows graphs of the performance of a predictive algorithm using maternal plasma data and fractional fetal DNA concentration, according to an embodiment of the present invention. [Figure 12A] 12 is a table 1200 showing details of 15 selected genomic loci for methylation prediction according to an embodiment of the present invention. [Figure 12B] A graph 1250 is shown showing predicted categories of 15 selected genomic loci and their corresponding methylation levels in the placenta. [Figure 13] 13 is a flow chart of a method 1300 for detecting fetal chromosomal abnormalities from a biological sample of a female subject carrying at least one fetus. [Figure 14]14 is a flow chart of a method 1400 for identifying methylation markers by comparing a placental methylation profile with a maternal methylation profile according to an embodiment of the present invention. [Figure 15A] 15 is a table 1500 showing the performance of the DMR identification algorithm using first phase data with reference to 33 previously reported first phase markers. [Figure 15B] 15 is a table 1550 showing the performance of the DMR identification algorithm using third trimester data compared to placental samples obtained at delivery. [Figure 16] 16 is a table 1600 showing the number of loci predicted to be hypermethylated or hypomethylated based on direct analysis of maternal plasma bisulfite sequencing data. [Figure 17A] 17 is a plot 1700 showing the size distribution of maternal plasma DNA, non-pregnant female control plasma DNA, placental DNA and peripheral blood DNA. [Figure 17B] 17 is a plot 1750 of the size distribution and methylation profiles of maternal plasma, adult female control plasma, placental tissue and adult female control blood. [Figure 18A] 1 is a plot of methylation density and size of plasma DNA molecules, according to an embodiment of the invention. [Figure 18B] 1 is a plot of methylation density and size of plasma DNA molecules, according to an embodiment of the invention. [Figure 19A] A plot 1900 of methylation density and size of sequenced reads from adult non-pregnant females is shown. [Figure 19B] 1950 showing the size distribution and methylation characteristics of fetal-specific and maternal-specific DNA molecules in maternal plasma. [Figure 20] 2 is a flow chart of a method 2000 for estimating the methylation level of DNA in a biological sample of an organism according to an embodiment of the present invention. [Figure 21A] 2 is a table 2100 showing methylation densities of pre-operative plasma and tissue samples of hepatocellular carcinoma (HCC) patients. [Figure 21B]21 is a table 2150 showing the number of sequence reads and sequencing depth achieved for each sample. [Figure 22] Table 220 shows that methylation densities in the autosomes ranged from 71.2% to 72.5% in plasma samples from healthy controls. [Figure 23A] Methylation densities of buffy coat, tumor tissue, non-tumorous liver tissue, pre-surgery plasma, and post-surgery plasma from HCC patients are shown. [Figure 23B] Methylation densities of buffy coat, tumor tissue, non-tumorous liver tissue, pre-surgery plasma, and post-surgery plasma from HCC patients are shown. [Figure 24A] 2 is a plot 2400 showing methylation density of pre-operative plasma from HCC patients. [Figure 24B] 2 is a plot 2450 showing methylation density of post-surgery plasma from HCC patients. [Figure 25A] 2 shows the Z-scores of plasma DNA methylation density of pre- (plot 2500) and post-surgery (plot 2550) plasma samples of HCC patients using plasma methylome data of four healthy control patients as reference for chromosome 1. [Figure 25B] 2 shows the Z-scores of plasma DNA methylation density of pre- (plot 2500) and post-surgery (plot 2550) plasma samples of HCC patients using plasma methylome data of four healthy control patients as reference for chromosome 1. [Figure 26A] 2 is a table 2600 showing Z-score data for pre-op and post-op plasma. [Figure 26B] Circos plot 2620 showing the Z-scores of plasma DNA methylation density of pre- and post-operative plasma samples of HCC patients using four healthy control patients as reference for 1 Mb bins analyzed from the entire autosome. [Figure 26C] 26 is a table 2640 showing the distribution of Z-scores of 1 Mb bins of the whole genome in both pre-surgery and post-surgery plasma samples of HCC patients. [Figure 26D] 26 is a table 2660 showing methylation levels of tumor tissue and pre-operative plasma samples overlapping with some of the control plasma samples using CHH and CHG sequences. [Figure 27A] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27B] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27C] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27D] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27E] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27F] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27G] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Fig. 27H] 1 shows Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. [Figure 27I] 27 is a table 2780 showing the number of sequence reads and sequencing depth achieved per sample. [Figure 27J] 2 is a table 2790 showing the Z-score distribution of 1 Mb bins across the genome in plasma of patients with different malignancies. CL=lung adenocarcinoma; NPC=nasopharyngeal carcinoma; CRC=colorectal carcinoma; NE=neuroendocrine carcinoma; SMS=leiomyosarcoma. [Figure 28] 28 is a flow chart of a method 2800 of analyzing a biological sample from an organism to determine a classification of the level of cancer according to an embodiment of the invention. [Figure 29A] 29 is a plot 2900 showing the distribution of methylation densities in reference subjects, assuming the distribution follows a normal distribution. [Figure 29B] 29 is a plot 2950 showing the distribution of methylation density in cancer subjects, assuming that the distribution is normal and that the mean methylation level is two standard deviations below the cutoff. [Diagram 30] 3 is a plot 3000 showing the distribution of methylation density of plasma DNA in healthy subjects and cancer patients. [Diagram 31] Graph 3100 showing the distribution of differences in methylation density between the mean values ​​of plasma DNA of healthy subjects and tumor tissue of HCC patients. [Figure 32A] 32 is a table 3200 showing the effect of reducing sequencing depth when plasma samples contained 5% or 2% tumor DNA. [Figure 32B] 32 is a graph 3250 showing methylation densities of repeat and non-repeat regions in plasma from four healthy control subjects, as well as buffy coat, normal liver tissue, tumor tissue, pre-surgery plasma samples, and post-surgery plasma samples from HCC patients. [Diagram 33] 33 shows a block diagram of an exemplary computer system 3300 that can be used with systems and methods according to embodiments of the present invention. [Figure 34A] 1 shows the size distribution of plasma DNA in systemic lupus erythematosus (SLE) patient SLE04. [Figure 34B] Methylation analysis of plasma DNA obtained from SLE patient SLE04 (FIG. 34B) and HCC patient TBR36 (FIG. 34C) is shown. [Figure 34C] Methylation analysis of plasma DNA obtained from SLE patient SLE04 (FIG. 34B) and HCC patient TBR36 (FIG. 34C) is shown. [Diagram 35] 3 is a flowchart of a method 3500 for determining a classification of a level of cancer based on hypermethylation of CpG islands according to an embodiment of the present invention. [Diagram 36] 36 is a flow chart of a method 3600 of analyzing a biological sample of an organism using multiple chromosomal regions according to an embodiment of the invention. [Figure 37A] CNA analysis (inside to outside) of tumor tissue, non-bisulfite (BS) treated plasma DNA and bisulfite treated plasma DNA from patient TBR36 is shown. [Figure 37B]FIG. 13 is a scatter plot showing the association between Z-scores for detection of CNAs using bisulfite-treated and non-bisulfite-treated plasma of 1 Mb bins of patient TBR36. [Figure 38A] CNA analysis of tumor tissue, non-bisulfite (BS) treated plasma DNA and bisulfite treated plasma DNA (inside to outside) from patient TBR34 is shown. [Figure 38B] FIG. 13 is a scatter plot showing the association between Z-scores for detection of CNAs using bisulfite-treated and non-bisulfite-treated plasma of 1 Mb bins from patient TBR34. [Figure 39A] Circos plot showing CNA (inner ring) and methylation analysis (outer ring) of bisulfite-treated plasma of HCC patient TBR240. [Figure 39B] Circos plot showing CNA (inner ring) and methylation analysis (outer ring) of bisulfite-treated plasma of HCC patient TBR164. [Figure 40A] CNA analysis of pre- and post-treatment samples for patient TBR36 is shown. [Figure 40B] Methylation analysis of pre-treatment and post-treatment samples for patient TBR36 is shown. [Figure 41A] CNA analysis of pre- and post-treatment samples for patient TBR34 is shown. [Figure 41B] Methylation analysis of pre-treatment and post-treatment samples for patient TBR34 is shown. [Diagram 42] FIG. 1 shows an illustration of the diagnostic performance of genome-wide hypomethylation analysis with different numbers of sequenced reads. [Diagram 43] FIG. 1 shows ROC curves for genome-wide hypomethylation-based cancer detection using different bin sizes (50 kb, 100 kb, 200 kb and 1 Mb). [Figure 44A] The cumulative probability (CP) and diagnostic performance of the proportion of bins with abnormalities are shown. [Figure 44B] Diagnostic performance of plasma analysis for global hypomethylation, CpG island hypermethylation and CNAs is shown. [Diagram 45] A table is shown, including the results of global hypomethylation, CpG island hypermethylation and CNA in patients with hepatocellular carcinoma. [Figure 46] A table is shown, including results for global hypomethylation, CpG island hypermethylation and CNA in patients with cancers other than hepatocellular carcinoma. [Figure 47] Serial plasma methylation analysis of case TBR34 is shown. [Figure 48A] Circos plot showing CNAs (inner ring) and methylation changes (outer ring) in bisulfite-treated plasma DNA of HCC patient TBR36. [Figure 48B] 1 is a plot of methylation Z scores for regions with chromosomal gains and losses, as well as regions with no copy number changes, for HCC patient TBR36. [Figure 49A] Circos plot showing CNAs (inner ring) and methylation changes (outer ring) in bisulfite-treated plasma DNA of HCC patient TBR34. [Figure 49B] 1 is a plot of methylation Z scores for regions with chromosomal gains and losses, and regions with no copy number changes, for HCC patient TBR34. [Figure 50A] The results of plasma hypomethylation and CNA analysis for SLE patients SLE04 and SLE10 are shown. [Figure 50B] The results of plasma hypomethylation and CNA analysis for SLE patients SLE04 and SLE10 are shown. [Figure 51A] Zmeth analysis is shown for regions with and without CNA in the plasma of two HCC patients (TBR34 and TBR36). [Figure 51B] Zmeth analysis is shown for regions with and without CNA in the plasma of two HCC patients (TBR34 and TBR36). [Figure 51C] Zmeth analysis of CNA-bearing and CNA-free regions in plasma from two SLE patients (SLE04 and SLE10) is shown. [Fig. 51D] Zmeth analysis of CNA-bearing and CNA-free regions in plasma from two SLE patients (SLE04 and SLE10) is shown. [Figure 52A] Hierarchical clustering analysis is shown for plasma samples obtained from HCC patients, non-HCC cancer patients and healthy control patients using group A features for CNA, global methylation and CpG island methylation. [Figure 52B] Hierarchical clustering using group B features for CNA, global methylation, and CpG island methylation is shown. [Figure 53A] Hierarchical clustering analysis of plasma samples from HCC patients, non-HCC cancer patients and healthy control patients using the A group CpG island methylation signature. [Figure 53B] Hierarchical clustering analysis of plasma samples from HCC patients, non-HCC cancer patients and healthy control patients using group A global methylation density is shown. [Figure 54A] Hierarchical clustering analysis of plasma samples obtained from HCC patients, non-HCC cancer patients and healthy control patients using Group A overall CNAs is shown. [Figure 54B] Hierarchical clustering analysis of plasma samples from HCC patients, non-HCC cancer patients and healthy control patients using group B CpG island methylation density. [Figure 55A] Hierarchical clustering analysis is shown for plasma samples obtained from HCC patients, non-HCC cancer patients and healthy control patients. [Figure 55B] Hierarchical clustering analysis of plasma samples from HCC patients, non-HCC cancer patients and healthy control patients using group B global methylation density. [Figure 56] The average methylation density (red dots) of 1 Mb bins among 32 healthy subjects is shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] definition A "methylome" provides a measure of the amount of DNA methylation at multiple sites or loci within a genome. The methylome can correspond to the entire genome, a substantial portion of the genome, or a relatively small portion of the genome. A "fetal methylome" refers to the methylome of a fetus in a pregnant woman. The fetal methylome can be determined using various fetal tissues or sources of fetal DNA, such as cell-free fetal DNA in placental tissue and maternal plasma. A "tumor methylome" refers to the methylome of a tumor in an organism (e.g., a human). The tumor methylome can be determined using cell-free tumor DNA in tumor tissue or maternal plasma. The fetal methylome and tumor methylome are examples of methylomes of interest. Other examples of methylomes of interest are the methylomes of organs (e.g., the methylomes of brain cells, bone, lung, heart, muscle, kidney, etc.) that can contribute DNA to bodily fluids (e.g., plasma, serum, sweat, saliva, urine, genital secretions, semen, stools fluid, diarrheal fluid, cerebrospinal fluid, gastrointestinal secretions, pancreatic secretions, intestinal secretions, sputum, tears, breast aspirates, and thyroid gland, etc.). The organ may be a transplanted organ.

[0013] A "plasma methylome" is a methylome determined from the plasma or serum of an animal (e.g., a human). The plasma methylome is an example of a cell-free methylome, since plasma and serum contain cell-free DNA. The plasma methylome is also an example of a mixed methylome, since it is a mixture of fetal / maternal methylomes or tumor / patient methylomes. A "placental methylome" can be determined from a chorionic villus sample (CVS) or a placental tissue sample (e.g., obtained after birth). A "cellular methylome" refers to a methylome determined from a patient's cells (e.g., blood cells). The methylome of blood cells is referred to as the blood cell methylome (or blood methylome).

[0014] A "site" refers to a single site, which may be a single base position or a group of related base positions (e.g., CpG sites). A "locus" refers to a region that contains multiple sites. A locus may contain only one site, making it equivalent to a site within the sequence.

[0015] The "methylation index" of each genomic site (e.g., CpG site) refers to the ratio of sequence reads that show methylation at the site to the total number of reads that cover the site. The "methylation density" of a region is the number of reads at sites in the region that show methylation divided by the total number of reads that cover sites in the region. A site may have a specific characteristic, for example, it may be a CpG site. Thus, the "CpG methylation density" of a region is the number of reads that show CpG methylation divided by the total number of reads that cover CpG sites in the region (e.g., a specific CpG site, a CpG site in a CG island, or a larger region). For example, the methylation density of each 100 kb bin in the human genome can be determined from the total number of unconverted cytosines (corresponding to methylated cytosines) at CpG sites after bisulfite treatment as a proportion of the total CpG sites covered by sequence reads that map to the 100 kb region. This analysis can also be performed for other bin sites (e.g., 50 kb or 1 Mb, etc.). The region can be the entire genome or a chromosome or a portion of a chromosome (e.g., a chromosome arm). The methylation index of a CpG site is the same as the methylation density of the region if the region only contains that CpG site. "Proportion of methylated cytosine" refers to the number of cytosine sites ("C") that are shown to be methylated (e.g., unconverted after bisulfite conversion) relative to the total number of analyzed cytosine residues (i.e., including cytosines outside the CpG sequence in the region). The methylation index, methylation density and proportion of methylated cytosine are examples of "methylation levels".

[0016] "Methylation profile" (also referred to as methylation status) includes information related to DNA methylation of a region. Information related to DNA methylation may include, but is not limited to, the methylation index of CpG sites, the methylation density of CpG sites within a region, the distribution of CpG sites across adjacent regions, the pattern or level of methylation of individual CpG sites within a region containing two or more CpG sites, and non-CpG methylation. The methylation profile of a substantial portion of a genome may be considered equivalent to a methylome. "DNA methylation" in mammalian genomes typically refers to the addition of a methyl group to the 5' carbon of a cytosine residue between CpG dinucleotides (i.e., 5-methylcytosine). DNA methylation may occur at cytosines within other sequences, such as CHG and CHH (H is adenine, cytosine, or thymine). Cytosine methylation may be in the form of 5-hydroxymethylcytosine. Non-cytosine methylation, such as N6-methyladenine, has also been reported.

[0017] "Tissue" refers to any cell. Different types of tissues can refer to different types of cells (e.g., liver, lung, or blood), but also tissues obtained from different organisms (mother vs. fetus) or normal vs. tumor cells. "Biological sample" refers to any sample obtained from a subject (e.g., a human, e.g., a pregnant woman, a cancer patient or suspected cancer, an organ transplant recipient, or a subject suspected of having a disease hypothesis related to an organ (e.g., the heart in a myocardial infarction, or the brain in a stroke)) and containing one or more nucleic acid molecules of interest. Biological samples can be body fluids such as blood, plasma, serum, urine, vaginal fluid, uterine or vaginal flushing fluid, multiple body fluids, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, etc. Stool samples can also be used.

[0018] The term "level of cancer" may refer to whether cancer is present, the stage of cancer, the size of the tumor, the presence or absence of metastasis, the total tumor burden in the body, and / or other measures of the severity of cancer. The level of cancer may be a number or other characteristic. The level may be zero. The level of cancer also includes premalignant or precancerous states associated with a mutation or several mutations. The level of cancer may be used in various ways. For example, screening may determine whether cancer is present in a person not previously known to have cancer. Evaluation may survey a person diagnosed with cancer to monitor the progression of cancer over time, study the effectiveness of treatment, or determine a prognosis. In one embodiment, a prognosis may represent the probability that a patient will die from cancer, or that the cancer will progress after a certain period or time, or that the cancer will metastasize. Detection may mean "screening" or may mean checking whether a person with characteristics suggestive of cancer (e.g., symptoms or other positive tests) has cancer.

[0019] Detailed Description Epigenetic mechanisms play an important role in embryonic and fetal development. However, human embryonic and fetal tissues (including placental tissue) are not readily available (US Pat. No. 6,927,028). Certain embodiments address this issue by analyzing samples containing cell-free fetal DNA molecules present in the maternal circulation. The fetal methylome can be estimated in a variety of ways. For example, the maternal plasma methylome can be compared to the cellular methylome (from the mother's blood cells), and the differences have been shown to correlate with the fetal methylome. As another example, fetal-specific alleles can be used to determine the methylation of the fetal methylome at specific loci. Furthermore, fragment size can be used as an indicator of methylation rate, as a correlation between size and methylation rate has been shown.

[0020] In one embodiment, genome-wide bisulfite sequencing is used to analyze the methylation profile (part or all of the methylome) of maternal plasma DNA at single nucleotide resolution. By utilizing polymorphic differences between the mother and fetus, the fetal methylome can be constructed from maternal blood samples. In another implementation, polymorphic differences were used, but differences between the plasma methylome and blood cell methylome can be used.

[0021] In another embodiment, by utilizing single nucleotide variation and / or copy number abnormality between tumor genome and non-tumor genome, and sequencing data obtained from plasma (or other sample), tumor methylation profiling can be performed in samples of patients suspected or diagnosed with cancer. The difference in methylation level in the plasma sample of a test individual compared with the plasma methylation level of a healthy control or healthy control group can allow the test individual to be identified as having cancer. Furthermore, methylation profile can serve as a signature to reveal the type of cancer, such as, for example, from which organ the patient developed and from which organ metastasis occurred.

[0022] The non-invasive nature of this approach allowed for serial assessment of fetal and maternal plasma methylomes from maternal blood samples taken in the first, third and postpartum trimesters. Pregnancy-related changes were observed. The approach can also be adapted to samples obtained during the second trimester. Fetal methylomes estimated from maternal plasma during pregnancy resembled placental methylomes. Imprinted genes and differentially methylated regions were identified from maternal plasma data.

[0023] We have developed an approach that offers the possibility of identifying biomarkers or directly testing pregnancy-related pathologies by studying fetal methylome non-invasively, continuously and comprehensively. The embodiments can also be used to study tumor methylome non-invasively, continuously and comprehensively to screen or detect whether a subject has cancer, to monitor malignancies in cancer patients, and for prognosis. The embodiments can be applied to any cancer type, including but not limited to lung cancer, breast cancer, colorectal cancer, prostate cancer, nasopharyngeal cancer, gastric cancer, testicular cancer, skin cancer (e.g., melanoma), cancer occurring in the nervous system, bone cancer, ovarian cancer, liver cancer (e.g., hepatocellular carcinoma), hematopoietic tumors, pancreatic cancer, endometrioma, renal cancer, cervical cancer, bladder cancer, etc.

[0024] A description of methods for determining the methylome or methylation profile is first discussed, followed by a description of various methylomes (e.g., fetal methylome, tumor methylome, maternal or patient methylome, and mixed methylomes (e.g., obtained from plasma)). Next, the determination of the fetal methylation profile is described, either using fetal specific markers or by comparing the mixed methylation profile to the cellular methylation profile. Fetal methylation markers are determined by comparing the methylation profiles. The relationship between size and methylation is discussed. The use of the methylation profile to detect cancer is also provided.

[0025] I. Methylome determination A myriad of approaches have been used to interrogate the placental methylome, but each approach has limitations. For example, sodium bisulfite, a chemical that changes unmethylated cytosine residues to uracil and leaves methylated cytosines unchanged, converts differences in cytosine methylation into gene sequence differences for further matching. The gold standard method to study cytosine methylation is based on treating tissue DNA with sodium bisulfite and then directly sequencing individual clones of bisulfite-converted DNA molecules. After analyzing multiple clones of DNA molecules, the pattern and quantitative characteristics of cytosine methylation per CpG site can be obtained. However, cloned bisulfite sequencing is a low-throughput and very laborious procedure that cannot be easily applied on a genome-wide scale.

[0026] Methylation-sensitive restriction enzymes, which typically digest unmethylated DNA, provide an inexpensive approach to study DNA methylation. However, data from such studies are limited to loci that have the enzyme recognition motif, and the results are non-quantitative. Immunoprecipitation of DNA bound by anti-methylated cytosine antibodies can be used to interrogate large segments of the genome, but tends to be biased toward loci with high density methylation, as the antibodies bind more strongly to such regions. Microarray-based approaches rely on a priori design of interrogation probes and the hybridization efficiency between the probes and target DNA.

[0027] To globally survey the methylome, in some embodiments, massively parallel sequencing (MPS) is used to provide genome-wide information and quantitative assessment of per-nucleotide and per-allele methylation levels. Recently, genome-wide MPS after bisulfite conversion has become feasible (R Lister et al 2008 Cell; 133: 523-536).

[0028] Of the few published studies that have applied genome-wide bisulfite sequencing to investigate the human methylome (R Lister et al. 2009 Nature; 462: 315-322; L Laurent et al. 2010 Genome Res; 20: 320-331; Y Li et al. 2010 PLoS Biol; 8: e1000533; and M Kulis et al. 2012 Nat Genet; 44: 1236-1242), two studies focused on embryonic stem cells and fetal fibroblasts (R Lister et al. 2009 Nature; 462: 315-322; L Laurent et al. 2010 Genome Res; 20: 320-331). Both of these studies analyzed DNA from cell lines.

[0029] A. Genome-wide bisulfite sequencing Certain embodiments can overcome the above problems and allow for comprehensive, non-invasive, and continuous interrogation of the fetal methylome. In one embodiment, genome-wide bisulfite sequencing was used to analyze cell-free fetal DNA molecules present in the circulation of pregnant women. Due to the low abundance and fragmentation of plasma DNA molecules, a high-resolution fetal methylome could be constructed from maternal plasma to continuously observe changes as the pregnancy progresses. Given the strong interest in noninvasive prenatal testing (NIPT), the embodiments can provide a powerful novel means for discovering fetal biomarkers or serve as a straightforward platform for achieving NIPT of fetal- or pregnancy-related diseases. Data obtained from genome-wide bisulfite sequencing of various samples from which the fetal methylome can be obtained are provided below. In one embodiment, this technology can be applied to methylation profiling in pregnancies complicated with preeclampsia, or intrauterine growth restriction, or preterm birth. In such complicated pregnancies, this technology can be used continuously due to its non-invasive nature to allow monitoring and / or prognosis and / or response to treatment.

[0030] FIG. 1A shows the results of maternal blood, placenta, and maternal plasma sequencing in Table 100 according to an embodiment of the present invention. In one embodiment, whole genome sequencing was performed on bisulfite converted DNA libraries prepared with methylated DNA library adapters (Illumina) (R Lister et al. 2008 Cell; 133: 523-536) of blood cells from blood samples collected in the first trimester, CVS, placental tissue collected at the end of term, maternal plasma samples collected in the first and third trimesters, and postpartum. Blood cell and plasma DNA samples from one adult male and one adult non-pregnant female were also analyzed. A total of 9.5 billion pairs of raw sequence reads were generated in this study. The sequencing coverage of each sample is shown in Table 100.

[0031] Sequence reads that were uniquely mappable to the human reference genome reached average haploid genome coverage of 50x, 34x, and 28x in first-, third-, and postpartum maternal plasma samples, respectively. Coverage of CpG sites within the genome ranged from 81% to 92% in samples obtained from the gestational periods. Sequence reads spanning CpG sites reached average haploid coverage of 33x / strand, 23x / strand, and 19x / strand in first-, third-, and postpartum maternal plasma samples, respectively. Bisulfite conversion efficiency for all samples was greater than 99.9% (Table 100).

[0032] In Table 100, the ambiguous percentage (labeled "a") refers to the percentage of reads that are mapped to both the Watson and Crick strands of the reference human genome. The lambda conversion percentage refers to the percentage of unmethylated cytosines in the internal lambda DNA control that are converted to "thymine" residues by bisulfite modification. H is generally equal to A, C, or T. "a" refers to reads that can be mapped to a specific genomic locus but cannot be assigned to the Watson or Crick strand. "b" refers to read pairs that have identical start and end coordinates. For "c", lambda DNA was spiked into each sample prior to bisulfite conversion. The lambda conversion percentage refers to the percentage of cytosine nucleotides that remain as cytosines after bisulfite conversion and is used as an indicator of the percentage of successful bisulfite conversion. "d" refers to the number of cytosine nucleotides present in the reference human genome that remain as cytosine sequences after bisulfite conversion.

[0033] During bisulfite modification, unmethylated cytosines are converted to uracil and then converted to thymine after PCR amplification, while methylated cytosines remain unchanged (M Frommer et al. 1992 Proc Natl Acad Sci USA;89:1827-31). Thus, after sequencing and alignment, the methylation status of individual CpG sites can be inferred from the number of methylated sequence reads "M" (methylated) and the number of unmethylated sequence reads "U" (unmethylated) at cytosine residues within the CpG sequence. Using bisulfite sequencing data, the entire methylome of maternal blood, placenta and maternal plasma was constructed. The average methylated CpG density (also called methylation density MD) of a particular locus in maternal plasma can be calculated using the following formula:

[0034]

number

[0035] where M is the number of methylated reads at CpG sites within the locus, and U is the number of unmethylated reads at CpG sites within the locus. If there are two or more CpG sites within the locus, then U corresponds to the total number of those sites.

[0036] B. Various Methods As described above, methylation profiling can be performed using massively parallel sequencing (MPS) of bisulfite-converted plasma DNA. MPS of bisulfite-converted plasma DNA can be performed in a random or shotgun fashion. The depth of sequencing can vary depending on the size of the region of interest.

[0037] In another embodiment, the region of interest in the bisulfite-converted plasma DNA can be first captured using a process based on liquid-phase or solid-phase hybridization, followed by MPS. Massively parallel sequencing can be performed using a sequencing-by-synthesis platform (e.g., Illumina), a sequencing-by-ligation platform (e.g., Life Technologies' SOLiD platform), a semiconductor-based sequencing system (e.g., Life Technologies' Ion Torrent or Ion Proton platform), or a single molecule sequencing system (e.g., Helicos system or Pacific Biosciences system or nanopore-based sequencing system). Nanopore-based sequencing includes, for example, nanopores constructed using lipid bilayers and protein nanopores, as well as solid-state nanopores (e.g., graphene-based nanopores). Since the selected single molecule sequencing platform allows the direct elucidation of the methylation status of DNA molecules (e.g. N6-methyladenine, 5-methylcytosine and 5-hydroxymethylcytosine) without bisulfite conversion (BA Flusberg et al. 2010 Nat Methods; 7: 461-465; J Shim et al. 2013 Sci Rep; 3:1389. doi: 10.1038 / srep01389), the use of such a platform allows the analysis of the methylation status of sample DNA (e.g. plasma DNA) that has not been converted by bisulfite.

[0038] In addition to sequencing, other techniques can be used. In one embodiment, methylation profiling can be performed by methylation-specific PCR or methylation-sensitive restriction enzyme digestion followed by PCR or ligase chain reaction followed by PCR. In yet another embodiment, the PCR is in the form of single molecule or digital PCR (B Vogelstein et al. 1999 Proc Natl Acad Sci USA; 96: 9236-9241). In yet another embodiment, the PCR can be real-time PCR. In other embodiments, the PCR can be multiplex PCR.

[0039] II. Methylome analysis Some embodiments can use whole genome bisulfite sequencing to determine the methylation profile of plasma DNA. The methylation profile of the fetus can be determined by sequencing maternal plasma DNA samples as described below. Thus, fetal DNA molecules (and fetal methylomes) can be non-invasively accessed throughout the pregnancy, and changes can be continuously monitored as the pregnancy progresses. Due to the comprehensiveness of the sequencing data, it was possible to study the maternal plasma methylome on a genome-wide scale at single nucleotide resolution.

[0040] Since the genomic location information of the sequenced reads was known, these data allowed the study of the methylome or the global methylation level of any region of interest in the genome, allowing comparisons between different genetic elements. Furthermore, multiple sequence reads covered each CpG site or locus. A description of some of the metrics used to measure the methylome is provided below.

[0041] A. Methylation of plasma DNA molecules DNA molecules are present in human plasma at low concentrations and in fragmented form, typically with lengths similar to mononucleosomal units (YMD Lo et al. 2010 Sci Transl Med; 2: 61ra91; and YW Zheng at al. 2012 Clin Chem; 58: 549-558). Despite these limitations, the genome-wide bisulfite sequencing pipeline was able to analyze the methylation of plasma DNA molecules. In yet another embodiment, since the selected single molecule sequencing platform allows the methylation state of DNA molecules to be elucidated directly without bisulfite conversion (BA Flusberg et al. 2010 Nat Methods; 7: 461-465; J Shim et al. 2013 Sci Rep; 3:1389. doi: 10.1038 / srep01389), the use of such a platform allows plasma DNA that is not converted by bisulfite to be used to measure plasma DNA methylation levels or to measure the plasma methylome. Such a platform can detect N6-methyladenine, 5-methylcytosine, and 5-hydroxymethylcytosine, which may result in improved results (e.g., improved sensitivity or specificity) associated with different biological functions of different forms of methylation. Such improved results may be useful when applying the embodiments to the detection or monitoring of certain disorders (e.g., pre-eclampsia) or certain cancer types.

[0042] Bisulfite sequencing can also distinguish different forms of methylation. In one embodiment, bisulfite sequencing can include an additional step that can distinguish 5-methylcytosine from 5-hydroxymethylcytosine. One such approach is oxidative bisulfite sequencing (oxBS-seq), which can reveal the location of 5-methylcytosine and 5-hydroxymethylcytosine at single base resolution (MJ Booth et al. 2012 Science; 336: 934-937; MJ Booth et al. 2013 Nature Protocols; 8: 1841-1851). In bisulfite sequencing, 5-methylcytosine and 5-hydroxymethylcytosine cannot be distinguished because both are read as cytosine. On the other hand, in oxBS-seq, the specific oxidation of 5-hydroxymethylcytosine to 5-formylcytosine by treatment with potassium perruthenate (KRuO4) followed by conversion of the newly formed 5-formylcytosine to uracil using bisulfite conversion allows 5-hydroxymethylcytosine to be distinguished from 5-methylcytosine. Thus, 5-methylcytosine reads can be obtained from one oxBS-seq run, and 5-hydroxymethylcytosine levels are estimated by comparing the results of bisulfite sequencing. In another embodiment, 5-methylcytosine can be distinguished from 5-hydroxymethylcytosine by using Tet-assisted bisulfite sequencing (TAB-seq) (M Yu et al. 2012 Nat Protoc; 7: 2159-2170). TAB-seq can identify 5-hydroxymethylcytosine at single-base resolution and determine its abundance at each modification site. The method involves β-glucosyltransferase-mediated protection (glucosylation) of 5-hydroxymethylcytosine and recombinant mouse Tet1 (mTet1)-mediated oxidation of 5-methylcytosine to 5-carboxylcytosine.After subsequent bisulfite treatment and PCR amplification, both cytosine and 5-carboxylcytosine (derived from 5-methylcytosine) are converted to thymine (T), while 5-hydroxymethylcytosine is read as C.

[0043] FIG. 1B shows methylation density within a 1 Mb window of a sample sequenced according to an embodiment of the present invention. Plot 150 is a Circos plot showing methylation density in maternal plasma and genomic DNA within a 1 Mb window across the genome. From outside to inside: chromosome diagrams can be in a pter-qter orientation in a clockwise direction: (centromeres are shown in red), maternal blood (red), placenta (yellow), maternal plasma (green), shared reads in maternal plasma (blue), and fetal-specific reads in maternal plasma (purple). The overall CpG methylation levels (i.e., density levels) of maternal blood cells, placenta, and maternal plasma can be found in Table 100. The methylation levels of maternal blood cells are usually higher than that of placenta across the entire genome.

[0044] B. Comparison of Bisulfite Sequencing with Other Methods Massively parallel bisulfite sequencing was used to study the placental methylome. In addition, the placental methylome was studied using an oligonucleotide array platform (Illumina) that covers approximately 480,000 CpG sites in the human genome (M Kulis et al. 2012 Nat Genet; 44: 1236-1242; and C Clark et al. 2012 PLoS One; 7: e50233). In one embodiment using beadchip-based genotyping and methylation analysis, genotyping was performed using Illumina HumanOmni2.5-8 genotyping arrays according to the manufacturer's protocol. Genotypes were called using the GenCall algorithm in Genome Studio Software (Illumina). The call rate was greater than 99%. For microarray-based methylation analysis, genomic DNA (500–800 ng) was treated with sodium bisulfite for the Illumina Infinium Methylation assay using the Zymo EZ DNA Methylation Kit (Zymo Research, Orange, CA, USA) according to the manufacturer's recommendations.

[0045] Methylation assays were performed on 4 μl of bisulfite converted genomic DNA (50 ng / μl) according to the Infinium HD Methylation Assay protocol. Hybridized bead chips were scanned on an Illumina iScan instrument. DNA methylation data were analyzed by GenomeStudio (v2011.1) Methylation Module (v1.9.0) software, including normalization to an internal standard and background subtraction. The methylation index of individual CpG sites is represented by the beta value (β), which is calculated using the ratio of fluorescence intensity between methylated and unmethylated alleles:

[0046]

number

[0047] For CpG sites represented on the array and sequenced to at least 10-fold coverage, the beta values ​​obtained by the array were compared to the methylation index determined by sequencing of the same sites. The beta value was expressed as the ratio of the intensity of the methylated probe to the combined intensity of the methylated and unmethylated probes covering the same CpG site. The methylation index for each CpG site refers to the ratio of methylated reads to the total number of reads covering that CpG.

[0048] Figures 2A-C show plots of β values ​​determined by Illumina Infinium HumanMethylation 450K BeadChip arrays against methylation indices determined by genome-wide bisulfite sequencing for corresponding CpG sites interrogated by both platforms: (A) maternal blood cells, (B) chorionic villus samples, and (C) term placental tissue. The data obtained from both platforms were highly concordant, with Pearson correlation coefficients of 0.972, 0.939, and 0.954 for maternal blood cells, CVS, and term placental tissue, respectively, with R 2 The values ​​were 0.945, 0.882 and 0.910.

[0049] Furthermore, the sequencing data were compared with those reported by Chu et al., who investigated the methylation profile of 12 paired CVS and maternal blood cell DNA samples using an oligonucleotide array covering approximately 27,000 CpG sites (T Chu et al. 2011 PLoS One; 6: e14723). Correlation data between CVS and maternal blood cell DNA, and the sequencing results of each of the 12 paired samples in the previous study, showed that the average Pearson coefficient (0.967) and R 2 (0.935), and the average Pearson coefficient (0.943) and R 2A mean mean score of (0.888) was obtained. Among the CpG sites represented on both arrays, our data were highly correlated with published data. The percentage of non-CpG methylation was less than 1% in maternal blood cells, CVS and placental tissues (Table 100). These results were consistent with the current finding that a significant amount of non-CpG methylation was mainly restricted to pluripotent cells (R Lister et al. 2009 Nature; 462: 315-322; L Laurent et al. 2010 Genome Res; 20: 320-331).

[0050] C. Comparison of plasma and blood methylomes in non-pregnant subjects Figures 3A and 3B show bar graphs of the percentage of methylated CpG sites in plasma and blood cells taken from adult males and non-pregnant adult females: (A) autosomes, (B) X chromosome. The diagrams show the similarity between the plasma methylome and blood methylome of males and non-pregnant females. The overall percentage of methylated CpG sites in plasma samples from males and non-pregnant females was similar to the corresponding blood cell DNA (Table 100 and Figures 2A and 2B).

[0051] Next, locus-specific correlations of the methylation profiles of plasma and blood cell samples were studied. The methylation density of each 100 kb bin in the human genome was determined by determining the total number of unconverted cytosines at CpG sites as a percentage of all CpG sites covered by sequence reads mapped to the 100 kb region. Methylation densities were highly concordant between plasma samples and corresponding blood cell DNA of male and female samples.

[0052] Figures 4A and 4B show plots of methylation density of corresponding loci in blood cell DNA and plasma DNA: (A) non-pregnant adult females, (B) adult males. Pearson correlation coefficients and R 2 The Pearson correlation coefficient and R 2The values ​​were 0.953 and 0.908, respectively. These data are consistent with previous findings based on genotypic evaluation of plasma DNA molecules of allogeneic hematopoietic stem cell transplant recipients, which showed that hematopoietic cells are the main source of DNA in human plasma (YW Zheng at al. 2012 Clin Chem; 58: 549-558).

[0053] D. Global methylation levels of the methylome We next examined DNA methylation levels in maternal plasma DNA, maternal blood cells, and placental tissue to determine methylation levels in repetitive, non-repeat, and global regions.

[0054] Figure 5A and Figure 5B show bar graphs of the percentage of methylated CpG sites among samples taken from pregnant women: (A) autosomes, (B) X chromosome. The total percentage of methylated CpGs was 67.0% and 68.2% in first and third trimester maternal plasma samples, respectively. Unlike the results obtained from non-pregnant individuals, these percentages were lower than those in first trimester maternal blood cell samples, but higher than those in CVS and term placental tissue samples (Table 100). Notably, the percentage of methylated CpGs in post-delivery maternal plasma samples was 73.1%, which was similar to the blood cell data (Table 100). These trends were observed in CpGs distributed across all autosomes and the X chromosome, spanning both non-repetitive and multiple classes of repetitive regions of the human genome.

[0055] Both repetitive and non-repetitive regions in the placenta were found to be hypomethylated compared to maternal blood cells. These results were consistent with findings in the literature that the placenta is hypomethylated compared to other tissues (e.g., peripheral blood cells).

[0056] Of the sequenced CpG sites in blood cell DNA from pregnant women, non-pregnant women, and adult men, 71%-72% were methylated (Table 100 in Figure 1). These data are comparable to the 68.4% of CpG sites in blood mononuclear cells reported by Y Li et al. 2010 PLoS Biol; 8: e1000533. Consistent with previous reports of the hypomethylated nature of placental tissue, 55% and 59% of CpG sites were methylated in CVS and term placental tissue, respectively (Table 100).

[0057] Figure 6 shows a bar graph of methylation levels of different repeat classes of the human genome in maternal blood, placenta and maternal plasma. Repeat classes are defined by the UCSC genome browser. Data shown are from first trimester samples. Unlike earlier data showing that the hypomethylated nature of placental tissue was mainly observed in certain repeat classes within the genome (B Novakovic et al. 2012 Placenta; 33: 959-970), here we show that the placenta was in fact hypomethylated in most classes of genomic elements relative to blood cells.

[0058] E. Methylome Similarities In some embodiments, the same platform can be used to determine the methylomes of placental tissue, blood cells and plasma. Thus, direct comparison of the methylomes of these biological sample types was possible. The high level of similarity between the methylomes of blood cells and plasma in men and non-pregnant women, as well as between maternal blood cells and postpartum maternal plasma samples, further confirmed that hematopoietic cells are the main source of DNA in human plasma (YW Zheng at al. 2012 Clin Chem; 58: 549-558).

[0059] Similarities are evident in terms of the overall percentage of methylated CpGs in the genome, as well as the high correlation of methylation densities between corresponding loci in blood cell DNA and plasma DNA. However, the overall percentage of methylated CpGs in first and third trimester maternal plasma samples was decreased when compared to maternal blood cell data or post-delivery maternal plasma samples. The decrease in methylation levels during pregnancy was due to the hypomethylated nature of fetal DNA molecules present in maternal plasma.

[0060] The methylation profile in post-delivery maternal plasma samples was reversed to more closely resemble that of maternal blood cells, suggesting that fetal DNA molecules were cleared from the maternal circulation. Calculation of fetal DNA concentration based on fetal SNP markers indeed showed that the concentration changed from 33.9% pre-delivery to only 4.5% in post-delivery samples.

[0061] F. Other Applications The embodiments have been successful in constructing DNA methylomes through MPS analysis of plasma DNA. The ability to determine placental or fetal methylomes from maternal plasma provides a non-invasive method for determining, detecting and monitoring abnormal methylation profiles associated with pregnancy-associated conditions such as pre-eclampsia, intrauterine growth restriction, preterm birth, etc. For example, detection of disease-specific abnormal methylation signatures allows screening, diagnosis and monitoring of such pregnancy-associated conditions. Measurement of maternal plasma methylation levels allows screening, diagnosis and monitoring of such pregnancy-associated conditions. In addition to direct application to the study of pregnancy-associated conditions, the above approach can be applied to other medical areas for plasma DNA analysis. For example, cancer methylomes can be determined from the plasma DNA of cancer patients. Cancer methylome analysis from plasma as described herein may be a synergistic scientific technology for cancer genome analysis from plasma (KCA Chan at al. 2013 Clin Chem; 59: 211-224 and RJ Leary et al. 2012 Sci Transl Med; 4:162ra154).

[0062] For example, the determination of the methylation level of plasma samples can be used for cancer screening. If the methylation level of plasma samples shows abnormal levels compared to healthy controls, cancer can be suspected. Further confirmation and evaluation of the cancer type or tissue origin of cancer can then be performed by determining the plasma signature of methylation at different genomic loci, or by plasma genomic analysis to detect tumor-associated copy number abnormalities, chromosomal translocations and single nucleotide mutations. Indeed, in one embodiment of the present invention, methylome and genomic profiling of cancer in plasma can be performed simultaneously. Alternatively, radiological and imaging tests (e.g., computed tomography, magnetic resonance imaging, positron emission tomography) or endoscopic tests (e.g., upper gastrointestinal endoscopy or colonoscopy) can be used to further test individuals suspected of having cancer based on plasma methylation level analysis.

[0063] In cancer screening or detection, the determination of methylation levels of plasma samples (or other biological samples) can be used in conjunction with other modalities for cancer screening or detection, such as prostate specific antigen measurement (e.g., for prostate cancer), carcinoembryonic antigen (e.g., for colorectal, gastric, pancreatic, lung, breast, medullary thyroid cancer), alpha-fetoprotein (e.g., for liver cancer or germ cell tumors), CA125 (e.g., for ovarian and breast cancer) and CA19-9 (e.g., for pancreatic cancer).

[0064] Additionally, other tissues can be sequenced to obtain cellular methylomes. For example, liver tissue can be analyzed to determine liver-specific methylation patterns, which may be used to identify liver pathology. Other tissues that can be analyzed include brain cells, bone, lung, heart, muscle, and kidney, etc. The methylation profile of various tissues can change over time, for example, as development, aging, disease progression (e.g., inflammation or cirrhosis of the liver or autoimmune processes (e.g., systemic lupus erythematosus)), or as a result of treatment (e.g., treatment with demethylating agents such as 5-azacytidine and 5-azadeoxycytidine). The dynamic nature of DNA methylation makes such analysis potentially very valuable in monitoring physiological and pathological processes. For example, disease progression in organs that contribute plasma DNA can be detected when detecting changes in an individual's plasma methylome compared to baseline values ​​obtained when the individual was healthy.

[0065] Also, the methylome of transplanted organs can be determined from the plasma DNA of organ transplant recipients. The graft methylome analysis from plasma described in the present invention can be a synergistic technology for graft genome analysis from plasma (YW Zheng at al, 2012; YMD Lo at al. 1998 Lancet; 351: 1329-1330; and TM Snyder et al. 2011 Proc Natl Acad Sci USA; 108: 6229-6234). Since plasma DNA is generally considered a marker of cell death, an increase in the plasma level of DNA released from a transplanted organ can be used as a marker for increased cell death from that organ, for example, a rejection episode or other pathological process involving that organ (e.g., infection or abscess). In the event that anti-rejection therapy is successfully initiated, the plasma level of DNA released by the transplanted organ is expected to decrease.

[0066] III. Determining the fetal or tumor methylome using SNPs As mentioned above, in non-pregnant healthy individuals, the plasma methylome corresponds to the blood methylome. However, in pregnant women, these methylomes are different. Fetal DNA molecules circulate in maternal plasma in a background of mostly maternal DNA (YMD Lo et al. 1998 Am J Hum Genet; 62: 768-775). Therefore, in pregnant women, the plasma methylome is mostly a mixture of the placental methylome and the blood methylome. Therefore, it is possible to extract the placental methylome from plasma.

[0067] In one embodiment, single nucleotide polymorphism (SNP) differences between the mother and fetus are used to identify fetal DNA molecules in maternal plasma. The aim is to identify SNP loci where the mother is homozygous but the fetus is heterozygous; fetal-specific alleles can be used to determine which DNA fragments are of fetal origin. Genomic DNA from maternal blood cells was analyzed using a SNP genotyping array, Illumina HumanOmni2.5-8. Meanwhile, for SNP loci where the mother is heterozygous and the fetus is homozygous, maternal-specific SNP alleles can be used to determine which plasma DNA fragments are of maternal origin. The methylation level of such DNA fragments reflects the methylation level of the associated genomic region in the mother.

[0068] A. Correlation of fetal-specific read methylation and the placental methylome Loci with two distinct alleles, where the amount of one allele (B) was significantly less than the amount of the other allele (A), were identified from the results of sequencing the biological samples. Reads covering the B allele were considered fetal specific (fetal-specific reads). Reads covering the A allele were shared by the mother and fetus (shared reads), since the mother was determined to be homozygous for A and the fetus was determined to be heterozygous for A / B.

[0069] In one pregnancy case analyzed, which was used to illustrate some of the aspects of the present invention, the pregnant mother was found to be homozygous at 1,945,516 loci on autosomes. The maternal plasma DNA sequencing reads covering these SNPs were examined. Reads carrying non-maternal alleles were detected at 107,750 loci, which were considered to be informative loci. For each informative SNP, the allele that was not derived from the mother was called the fetal-specific allele, and the other alleles were called the shared allele.

[0070] The fractional fetal / tumor DNA concentration (also referred to as fetal DNA fraction) in maternal plasma can be determined. In one embodiment, the fractional fetal DNA concentration (f) in maternal plasma is determined by the formula:

[0071]

number

[0072] where p is the number of sequenced reads with fetal-specific alleles and q is the number of sequenced reads with alleles shared between the mother and fetus (YMD Lo et al. 2010 Sci Transl Med; 2: 61ra91). The fetal DNA percentages in the first trimester, third trimester and postpartum maternal plasma samples were found to be 14.4%, 33.9% and 4.5%, respectively. The number of reads aligned to the Y chromosome was also used to calculate the fetal DNA percentage. Based on the Y chromosome data, the results were 14.2%, 34.9% and 3.7% in the first trimester, third trimester and postpartum maternal plasma samples, respectively.

[0073] By separately analyzing fetal-specific or shared sequence reads, embodiments show that circulating fetal DNA molecules were even more hypomethylated than background DNA molecules. Comparison of methylation densities of corresponding loci in both first and third trimester fetal-specific maternal plasma reads and placental tissue data revealed a high level of correlation. These data provide evidence at the genome level that the placenta is the main source of fetal-derived DNA molecules in maternal plasma, and represent a major advancement compared to previous evidence based on information from selected loci.

[0074] The methylation density of each 1 Mb region in the genome was determined using fetal-specific or shared reads covering CpG sites adjacent to informative SNPs. The fetal-specific and non-fetal-specific methylomes constructed from maternal plasma sequence reads can be displayed, for example, in a Circos plot (M Krzywinski et al. 2009 Genome Res; 19: 1639-1645). The methylation density per 1 Mb bin in maternal blood cells and placental tissue samples was also determined.

[0075] FIG. 7A shows a Circos plot 700 of a first trimester sample. FIG. 7B shows a Circos plot 750 of a third trimester sample. Plots 700 and 750 show methylation density per 1 Mb bin. The chromosome diagram (outermost ring) is in a clockwise direction with pter-qter orientation (centromeres shown in red). The second outermost track shows the number of CpG sites in the corresponding 1 Mb region. The scale of the red bar shown is up to 20,000 sites per 1 Mb bin. The methylation density of the corresponding 1 Mb region is shown in the other tracks based on the color scheme shown in the center.

[0076] In the first trimester sample (FIG. 7A), from inside to outside, the tracks are chorionic villus specimen, fetal-specific reads in maternal plasma, maternal-specific reads in maternal plasma, combined fetal and non-fetal reads in maternal plasma, and maternal blood cells. In the third trimester sample (FIG. 7B), the tracks are term placental tissue, fetal-specific reads in maternal plasma, maternal-specific reads in maternal plasma, combined fetal and non-fetal reads in maternal plasma, post-delivery maternal plasma, and maternal blood cells (obtained from the first trimester blood sample). It can be seen that in both the first and third trimester plasma samples, the fetal methylome was more hypomethylated than the methylation state of the non-fetal specific methylome.

[0077] The global methylation profile of the fetal methylome closely resembled that of the CVS or placental tissue samples. Conversely, the DNA methylation profile of shared reads in plasma, which were primarily maternal DNA, closely resembled that of maternal blood cells. We then performed a systematic locus-by-locus comparison of the methylation densities of maternal plasma DNA reads and maternal or fetal tissues. We determined the methylation density of CpG sites that were present as informative SNPs on the same sequence reads and were covered by at least five maternal plasma DNA sequence reads.

[0078] 8A-8D show plots of methylation density comparisons of genomic tissue DNA to maternal plasma DNA for CpG sites surrounding informative single nucleotide polymorphisms. FIG. 8A shows the methylation density of fetal-specific reads in first trimester maternal plasma samples compared to the methylation density of reads in CVS samples. As shown, the values ​​of fetal-specific reads are in good agreement with the values ​​of CVS reads.

[0079] Figure 8B shows the methylation density of fetal-specific reads in third trimester maternal plasma samples compared to the methylation density of reads in term placental tissue. Again, these sets of densities are in good agreement, indicating that a fetal methylation profile can be obtained by analyzing reads with fetal-specific alleles.

[0080] Figure 8C shows the methylation density of shared reads in the first trimester maternal plasma sample compared to the methylation density of reads in maternal blood cells. Considering that the majority of shared reads are of maternal origin, these two sets of values ​​are in good agreement. Figure 8D shows the methylation density of shared reads in the third trimester maternal plasma sample compared to the methylation density of reads in maternal blood cells.

[0081] For fetal-specific reads in maternal plasma, the Spearman correlation coefficient between first-trimester maternal plasma and CVS was 0.705 (P<2.2*e-16); the Spearman correlation coefficient between third-trimester maternal plasma and term placental tissue was 0.796 (P<2.2*e-16) (Figures 8A and 8B). Similar comparisons were made for shared reads in maternal plasma and maternal blood cell data. For first-trimester plasma samples, the Pearson correlation coefficient was 0.653 (P<2.2*e-16), and for third-trimester plasma samples, the Pearson correlation coefficient was 0.638 (P<2.2*e-16) (Figures 8C and 8D).

[0082] B. Fetal Methylome In one embodiment, to construct a fetal methylome from maternal plasma, sequence reads that covered at least one informative fetal SNP site and contained at least one CpG site within the same read were selected. Reads that showed fetal-specific alleles were included in constructing the fetal methylome. Reads that showed shared alleles (i.e., non-fetal specific alleles) were included in constructing the non-fetal specific methylome, which is mainly composed of maternal DNA molecules.

[0083] Fetal-specific reads from first-trimester maternal plasma samples covered 218,010 CpG sites on autosomes. The corresponding figures for third-trimester and postpartum maternal plasma samples were 263,611 and 74,020, respectively. On average, shared reads covered those CpG sites 33.3-fold, 21.7-fold, and 26.3-fold, respectively. Fetal-specific reads from first-trimester, third-trimester, and postpartum maternal plasma samples covered those CpG sites 3.0-fold, 4.4-fold, and 1.8-fold, respectively.

[0084] Since fetal DNA is a minority population in maternal plasma, the coverage of those CpG sites by fetal-specific reads was proportional to the fetal DNA percentage of the sample. In the first trimester maternal plasma sample, the overall percentage of methylated CpGs among fetal reads was 47.0%, while the overall percentage of methylated CpGs in shared reads was 68.1%. In the third trimester maternal plasma sample, the percentage of methylated CpGs among fetal reads was 53.3%, while the percentage of methylated CpGs in shared reads was 68.8%. These data showed that fetal-specific reads in maternal plasma were less methylated than shared reads in maternal plasma.

[0085] C. Method Using the above techniques, tumor methylation profiles can also be determined. Methods for determining the methylation profiles of fetuses and tumors are described below.

[0086] 9 is a flow chart showing a method 900 for determining a first methylation profile from a biological sample of an organism according to an embodiment of the present invention. The method 900 can construct an epigenetic map of a fetus from a methylation profile of maternal plasma. The biological sample includes cell-free DNA that includes a mixture of cell-free DNA from a first tissue and a second tissue. By way of example, the first tissue can be obtained from a fetus, a tumor, or a transplanted organ.

[0087] A plurality of DNA molecules are analyzed from the biological sample at block 910. Analysis of the DNA molecules may include determining the location of the DNA molecule within the genome of the organism, determining the genotype of the DNA molecule, and determining whether the DNA molecule is methylated at one or more sites.

[0088] In one embodiment, DNA molecules are analyzed using sequence reads of the DNA molecules, where the sequencing recognizes methylation. Thus, the sequence reads include the methylation state of the DNA molecules from the biological sample. The methylation state can include whether a particular cytosine residue is 5-methylcytosine or 5-hydroxymethylcytosine. The sequence reads can be obtained from various sequencing methods, PCR methods, arrays, and other suitable techniques for identifying the sequence of a fragment. The methylation state of the site of the sequence read can be obtained as described herein.

[0089] At block 920, a plurality of first loci are identified where the first genome of the first tissue is heterozygous for each first allele and each second allele, and the second genome of the second tissue is homozygous for each first allele. For example, fetal-specific reads can be identified at the plurality of first loci. Alternatively, tumor-specific reads can be identified at the plurality of first loci. Tissue-specific reads can be identified from sequencing reads where a percentage of sequence reads of the second allele falls within a particular range, for example, a range of about 3% to 25%, thereby indicating that DNA fragments from the heterozygous genome at the locus are a minority population and DNA fragments from the homozygous genome at the locus are a majority population.

[0090] At block 930, DNA molecules located at one or more sites of each first locus are analyzed. A number of DNA molecules that are methylated at the site and correspond to each second allele of the locus are determined. There may be more than one site per locus. For example, SNPs may indicate that the fragment is fetal specific and that the fragment may have multiple sites where the methylation status is determined. The number of reads at each methylated site may be determined, and the total number of methylated reads at the locus may be determined.

[0091] A locus can be defined by a specific number of sites, a specific set of sites, or the specific size of the region surrounding the mutation that contains tissue-specific alleles.A locus can have only one site.A site can have a specific property, for example, it can be a CpG site.Determining the number of unmethylated reads is the same and is included in determining methylation status.

[0092] At block 940, a methylation density is calculated for each of the first loci based on the number of DNA molecules that are methylated at one or more sites at the locus and correspond to each second allele at the locus. For example, the methylation density of CpG sites corresponding to the locus can be determined.

[0093] At block 950, a first methylation profile of the first tissue is generated from the methylation densities of the first loci. The first methylation profile may correspond to a particular site, e.g., a CpG site. The methylation profile may be for all loci that have fetal-specific alleles, or just some of those loci.

[0094] IV. Use of Plasma Methylome and Blood Methylome Differences Above, it was shown that fetal-specific reads from plasma correlate with the placental methylome. Because the maternal component of the maternal plasma methylome is mainly contributed by blood cells, the difference between the plasma methylome and the blood methylome can be used to determine the placental methylome of all loci, not just the location of fetal-specific alleles. The difference between the plasma methylome and the blood methylome can also be used to determine the methylome of the tumor.

[0095] A. Method 10 is a flow chart showing a method 1000 of determining a first methylation profile from a biological sample of an organism according to an embodiment of the present invention. The biological sample (e.g., plasma) includes cell-free DNA comprising a mixture of cell-free DNA from a first tissue and from a second tissue. The first methylation profile corresponds to the methylation profile of the first tissue (e.g., fetal tissue or tumor tissue). The method 1200 may allow for the inference of differentially methylated regions from maternal plasma.

[0096] At block 1010, a biological sample is received. The biological sample may simply be received at a device (e.g., a sequencing device). The biological sample may be in a form obtained from an organism or may be in a processed form, for example, the sample may be plasma extracted from a blood sample.

[0097] At block 1020, a second methylation profile is obtained that corresponds to the DNA of the second tissue. The second methylation profile may be read from memory, having been previously determined. The second methylation profile may be determined from the second tissue (e.g., a different sample that contains only or primarily cells of the second tissue). The second methylation profile may correspond to a cellular methylation profile and may be obtained from cellular DNA. As another example, the second profile may be determined from a plasma sample taken prior to pregnancy or prior to the development of cancer, since the plasma methylome of a non-pregnant individual without cancer closely resembles the methylome of blood cells.

[0098] The second methylation profile may provide a methylation density at each of a plurality of loci in the genome of the organism. The methylation density at a particular locus corresponds to the proportion of DNA in the second tissue that is methylated. In one embodiment, the methylation density is a CpG methylation density, where the CpG sites associated with the locus are used to determine the methylation density. If there is one site at the locus, the methylation density may be the same as the methylation index. The methylation density also corresponds to the unmethylation density, since the values ​​of the methylation density and the unmethylation density are complementary.

[0099] In one embodiment, the second methylation profile is obtained by performing methylation-aware sequencing of cellular DNA from a sample of an organism. One example of methylation-aware sequencing involves treating the DNA with sodium bisulfite, followed by DNA sequencing. In another example, methylation-aware sequencing can be performed without sodium bisulfite using single molecule sequencing platforms that allow the methylation state of DNA molecules (e.g., N6-methyladenine, 5-methylcytosine and 5-hydroxymethylcytosine) to be directly elucidated without bisulfite conversion (AB Flusberg et al. 2010 Nat Methods; 7: 461-465; J Shim et al. 2013 Sci Rep; 3:1389. doi: 10.1038 / srep01389); or by immunoprecipitation of methylated cytosine (e.g., by using antibodies against methylcytosine or by using methylated DNA binding proteins or peptides) (LG Acevedo et al. 2011 Epigenomics; 3: 93-101) followed by sequencing; or by the use of methylation-sensitive restriction enzymes followed by sequencing. In another embodiment, non-sequencing methods are used, such as arrays, digital PCR and mass spectrometry.

[0100] In another embodiment, the second methylation density of the second tissue can be obtained in advance from a control sample of the subject or from another subject. The methylation density from another subject can serve as a reference methylation profile with reference methylation density. The reference methylation density can be determined from a number of samples, where the average level (or other statistical value) of different methylation densities at a locus can be used as the reference methylation density at that locus.

[0101] At block 1030, a cell-free methylation profile is determined from the cell-free DNA of the mixture. The cell-free methylation profile provides a methylation density at each of a plurality of loci. The cell-free methylation profile can be determined by obtaining sequence reads from sequencing the cell-free DNA and using the sequence reads to obtain methylation information. The cell-free methylation profile can be determined similarly to a cellular methylome.

[0102] In block 1040, the percentage of cell-free DNA from the first tissue in the biological sample is determined. In one embodiment, the first tissue is fetal tissue and the corresponding DNA is fetal DNA. In another embodiment, the first tissue is tumor tissue and the corresponding DNA is tumor DNA. The percentage can be determined in various ways, for example, using fetal-specific alleles or tumor-specific alleles. The percentage can also be determined using copy number, for example, as described in U.S. Patent Application Publication No. 13 / 801,748, entitled "Mutational Analysis Of Plasma DNA For Cancer Detection," filed March 13, 3013 (incorporated by reference).

[0103] At block 1050, a plurality of loci for determining the first methylome is identified. These loci may correspond to each locus used to determine the cell-free methylation profile and the second methylation profile. Thus, the plurality of loci may be coincident. It is possible that more loci can be used to determine the cell-free methylation profile and the second methylation profile.

[0104] In some embodiments, loci that are hypermethylated or hypomethylated in the second methylation profile can be identified, for example, using maternal blood cells. To identify hypermethylated loci in maternal blood cells, CpG sites with a methylation index of X% or more (e.g., 80%) can be scanned from one end of the chromosome. Then, the next CpG site in the downstream region (e.g., within 200 bp downstream) can be searched. If the immediately downstream CpG site has a methylation index of X% or more (or other specified amount), the first CpG site and the second CpG site can be grouped. The grouping can be continued until there are no other CpG sites in the next downstream region; or until the immediately downstream CpG site has a methylation index of less than X%. The region of grouped CpG sites can be reported as hypermethylated in maternal blood cells if the region contains at least five directly adjacent hypermethylated CpG sites. Similar analyses can be performed to search for loci that are hypomethylated in maternal blood cells for CpG sites with a methylation index of 20% or less. Methylation densities of secondary methylation profiles can be calculated for short-listed loci and used to estimate primary methylation profiles (e.g., placental tissue methylation densities) of corresponding loci, e.g., from maternal plasma bisulfite sequencing data.

[0105] At block 1060, a first methylation profile of the first tissue is determined by calculating a differential parameter comprising the difference between the methylation density of the second methylation profile and the methylation density of the cell-free methylation profile for each of the plurality of loci, the difference being determined as a percentage.

[0106] In one embodiment, a first methylation density (D) of a locus in a first (e.g., placental) tissue is estimated using the formula:

[0107]

number

[0108] Wherein, mbc indicates the methylation density of the second methylation profile at the locus (e.g., the short listed locus determined in the bisulfite sequencing data of maternal blood cells); mp indicates the methylation density of the corresponding locus in the bisulfite sequencing data of maternal plasma; f indicates the percentage of cell-free DNA from the first tissue (e.g., fetal DNA concentration fraction), and CN indicates the copy number at the locus (e.g., higher amplification value or lower number of deletions compared to normal). If there is no amplification or deletion in the first tissue, CN can be 1. In trisomy (i.e., duplication of the region in tumor or fetus), CN is 1.5 (because the increase is from 2 copies to 3 copies), and monosomy has 0.5. Higher amplification can increase by a value of 0.5. In this example, D can correspond to a differential parameter.

[0109] In block 1070, the first methylation density is transformed to obtain a corrected first methylation density of the first tissue. The transformation may provide a fixed difference between the differential parameter and the actual methylation profile of the first tissue. For example, the values ​​may differ by a fixed constant or by a slope. The transformation may be linear or non-linear.

[0110] In one embodiment, the distribution of the estimated value D was found to be lower than the actual methylation level of placental tissue. For example, the estimated value can be linearly transformed using data obtained from CG islands, which are genomic segments with over-represented CpG sites. The genomic locations of CG islands used in this study were obtained from the UCSC Genome Browser database (NCBI build 36 / hg18) (PA Fujita et al. 2011 Nucleic Acids Res; 39: D876-882). For example, CG islands can be defined as genomic segments with GC content of 50% or more, genomic length of more than 200bp, and measured / predicted CpG number ratio of more than 0.6 (M Gardiner-Garden et al 1987 J Mol Biol; 196: 261-282).

[0111] In one implementation, CpG islands with at least four CpG sites and an average read depth of 5 or more per CpG site in the sequenced samples may be included to obtain a linear conversion formula. After determining the linear relationship between the methylation density of CpG islands in the CVS or term placenta and the estimated value D, the predicted value was determined using the following formula: First period forecast value = D x 1.6 + 0.2 Third stage predicted value = D x 1.2 + 0.05

[0112] B. Fetal Example As described above, by using method 1000, the methylation status of the placenta can be estimated from maternal plasma. Circulating DNA in plasma is mainly derived from hematopoietic cells. The proportion of cell-free DNA contributed by other unknown internal organs is unknown. Furthermore, placenta-derived cell-free DNA accounts for approximately 5-40% of the total DNA in maternal plasma, with an average value of approximately 15%. Therefore, it can be assumed that the methylation level in maternal plasma is equal to the existing background methylation + the contribution of the placenta during pregnancy, as described above.

[0113] The maternal plasma methylation level, MP, can be determined using the following formula:

[0114]

number

[0115] where BKG is the background DNA methylation level in plasma obtained from blood cells and internal organs, PLN is the methylation level of the placenta, and f is the fractional fetal DNA concentration in maternal plasma.

[0116] In one embodiment, the placental methylation level can be theoretically estimated by the following formula:

[0117]

number

[0118] Equation (1) and equation (2) are equal when CN is equal to 1, D is equal to PLN, and BKG is equal to mbc. In another embodiment, the fractional fetal DNA concentration can be assumed or set to a particular value (e.g., as part of the assumption that a minimum value f exists).

[0119] Methylation levels in maternal blood were considered to represent background methylation in maternal plasma.In addition to loci that were hyper- or hypomethylated in maternal blood cells, we further sought to predict the predictive approach by focusing on defined regions with clinical relevance, e.g., CG islands in the human genome.

[0120] The average methylation density of a total of 27,458 CG islands (NCBI Build36 / hg18) on the autosomes and X chromosomes was obtained from the sequencing data of maternal plasma and placenta. Only CG islands with ≥10 covered CpG sites and an average read depth ≥5 per covered CpG site in all analyzed samples, including placenta, maternal blood, and maternal plasma, were selected. As a result, 26,698 CG islands (97.2%) remained as valid, and their methylation levels were estimated using the plasma methylation data and the fetal DNA fractional concentration according to the formula above.

[0121] It was noted that the distribution of estimated PLN values ​​was lower than the actual methylation levels of CpG islands in placental tissue. Therefore, in one embodiment, the estimated PLN values, or simply the estimated values ​​(D), were used as an arbitrary unit to estimate the methylation levels of CpG islands in the placenta. After transformation, the linearly estimated values ​​and their distributions were more similar to the actual dataset. The transformed estimates were named methylation predictive values ​​(MPVs) and were then used to predict the methylation levels of loci in the placenta.

[0122] In this example, CG islands were classified into three categories based on their methylation density in the placenta: low (≤0.4), medium (>0.4-<0.8) and high (≥0.8). Using the estimation equation, the MPV of the same set of CG islands was calculated, and then the values ​​were used to classify the CG islands into three categories with the same cutoff. By comparing the actual and estimated datasets, it was found that 75.1% of the short-listed CG islands could be accurately matched to the same category in the tissue data by their MPV. Approximately 22% of the CG islands were assigned to groups with one level of difference (high vs. medium, or medium vs. low), and less than 3% were completely misclassified (high vs. low) (Figure 12A). The overall classification performance was also determined: 86.1%, 31.4% and 68.8% of CG islands with methylation densities of ≥0.4, ≥0.4-<0.8 and ≥0.8 in the placenta were accurately estimated as "low", "medium" and "high" (Figure 12B).

[0123] 11A and 11B show graphs of the performance of a prediction algorithm using maternal plasma data and fetal DNA fractional concentration according to an embodiment of the present invention. Fig. 11A is a graph 1100 showing the accuracy of CG islet classification using MPV-corrected classification (estimated category exactly matches the actual dataset); one level of difference (estimated category differs by one level from the actual dataset); and misclassification (estimated category is opposite to the actual dataset). Fig. 11B is a graph 1150 showing the percentage of CG islets classified into each estimated category.

[0124] If maternal background methylation is low in each genomic region, the presence of hypermethylated placenta-derived DNA in the circulating blood increases the overall plasma methylation level to a degree that depends on the fetal DNA concentration fraction. If the released fetal DNA is fully methylated, a significant change can be observed. Conversely, if maternal background methylation is high, the degree of change in plasma methylation level becomes more significant when hypomethylated fetal DNA is released. Thus, the estimation scheme may be more practical when it is estimated for loci whose methylation levels are known to differ between maternal background and placenta, especially for hypermethylated loci and hypomethylated markers in the placenta.

[0125] FIG. 12A is a table 1200 showing details of 15 selected genomic loci for methylation prediction according to an embodiment of the present invention. To validate the method, 15 previously studied differentially methylated genomic loci were selected. The methylation levels of the selected regions were estimated and compared with the 15 previously studied differentially methylated loci (RWK Chiu et al. 2007 Am J Pathol; 170: 941-950; SSC Chim et al. 2008 Clin Chem; 54: 500-511; SSC Chim et al. 2005 Proc Natl Acad Sci USA; 102: 14753-14758; DWY Tsui et al. 2010 PLoS One; 5: e15069).

[0126] FIG. 12B is a graph 1250 showing the estimated categories of 15 selected genomic loci and their corresponding methylation levels in the placenta. The estimated methylation categories are: low, 0.4 or less; medium, >0.4 to <0.8; high, 0.8 or more. Table 1200 and graph 1300 show that their methylation levels in the placenta can be accurately estimated, with some exceptions (RASSF1A, CGI009, CGI137, and VAPA). Of these four markers, only CGI009 showed significant discrepancies with the actual dataset. The others were only slightly misclassified.

[0127] In table 1200, "1" refers to the estimated value (D) calculated by the following formula:

[0128]

number

[0129] where f is the fractional fetal DNA concentration. Label "2" refers to methylation predictive value (MPV), which refers to a linearly transformed estimate using the formula: MPV=D×1.6+0.25. Label "3" refers to classification cutoffs for the estimates: low, 0.4 or less; medium, >0.4 to <0.8; high, 0.8 or more. Label "4" refers to classification cutoffs for the actual placenta dataset: low, 0.4 or less; medium, >0.4 to <0.8; high, 0.8 or more. Label "5" indicates that placental status refers to the methylation status of the placenta compared to the methylation status of maternal blood cells.

[0130] C. Calculation of the Fetal DNA Fractional Concentration In one embodiment, the percentage of fetal DNA from the first tissue can be the Y chromosome of the male fetus. The percentage of Y chromosome sequences in the maternal plasma sample (Y chromosome rate) was a composite of the Y chromosome reads from the male fetus and the number of maternal (female) reads that misaligned to the Y chromosome (RWK Chiu et al. 2011 BMJ; 342: c7401). Thus, the relationship between the Y chromosome rate and the fetal DNA concentration fraction (f) in the sample can be given by:

[0131]

number

[0132] In the formula, %chrY male (Y chromosome percentage (male)) refers to the percentage of reads that aligned to the Y chromosome in plasma samples containing 100% male DNA; %chrY female (Y chromosome fraction (female)) refers to the percentage of reads that aligned to the Y chromosome in plasma samples containing 100% female DNA.

[0133] The Y chromosome percentage can be determined from reads that align to the Y chromosome without mismatches from a sample obtained from a woman carrying a male fetus, e.g., the reads are from a bisulfite converted sample. male Values ​​can be obtained from bisulfite sequencing of two adult male plasma samples. female Values ​​can be obtained from bisulfite sequencing of two non-pregnant adult female plasma samples.

[0134] In other embodiments, the fetal DNA percentage can be determined from fetal-specific alleles on autosomes. As another example, epigenetic markers can be used to determine the fetal DNA percentage. Other methods of determining fetal DNA percentage can be used.

[0135] D. Methods for Determining Copy Number Using Methylation The placental genome is more hypomethylated than the maternal genome. As mentioned above, the methylation of the plasma of a pregnant woman depends on the concentration fraction of fetal DNA from the placenta in the maternal plasma. Thus, through the analysis of the methylation density of chromosomal regions, it is possible to detect differences in the contribution of fetal tissue to maternal plasma. For example, in a pregnant woman carrying a trisomic fetus (e.g., a fetus carrying trisomy 21 or trisomy 18 or trisomy 13), the fetus contributes an additional amount of DNA from the trisomic chromosome to the maternal plasma when compared to the disomic chromosome. In this situation, the plasma methylation density of the trisomic chromosome (or any chromosomal region with amplification) will be lower than that of the disomic chromosome. The degree of difference can be predicted by mathematical calculations by considering the fetal DNA concentration fraction in the plasma sample. The higher the fetal DNA concentration fraction in the plasma sample, the greater the difference in methylation density between the trisomic chromosome and the disomic chromosome. For regions with deletions, the methylation density will be higher.

[0136] One example of a deletion is Turner syndrome, where a female fetus has only one copy of the X chromosome. In this situation, in a pregnant woman carrying a fetus with Turner syndrome, the methylation density of the X chromosome in plasma DNA will be higher than in the situation of the same pregnant woman carrying a female fetus with a normal number of X chromosomes. In one embodiment of this strategy, maternal plasma can be first analyzed (e.g., using MPS or PCR-based techniques) for the presence or absence of Y chromosome sequences. If Y chromosome sequences are present, the fetus can be classified as male, and no further analysis is required. On the other hand, if Y chromosome sequences are absent in maternal plasma, the fetus can be classified as female. In this situation, the methylation density of the X chromosome in maternal plasma can then be analyzed. A higher than normal X chromosome methylation density indicates that the fetus is at high risk of having Turner syndrome. This approach can also be applied to other sex chromosome aneuploidies. For example, in a fetus with XYY, the methylation density of the Y chromosome in maternal plasma will be lower than that of a normal XY fetus with a similar level of fetal DNA in maternal plasma. As another example, in fetuses with Klinefelter syndrome (XXY), Y chromosome sequences are present in maternal plasma, but the methylation density of the X chromosome in maternal plasma is lower than that of normal XY fetuses with similar levels of fetal DNA in maternal plasma.

[0137] From the above considerations, the plasma methylation density (MP) of disomic chromosomes Non-aneu (non-differential)) is

[0138]

number

[0139] where BKG is the background DNA methylation level in blood cells and plasma from internal organs, PLN is the methylation level of the placenta, and f is the fetal DNA fraction in maternal plasma.

[0140] Plasma methylation density of trisomy chromosomes (MP Aneu (different number)) is

[0141]

number

[0142] where 1.5 corresponds to the copy number CN and the addition of one more chromosome is an increase of 50%. Diff ) becomes:

[0143]

number

[0144] In one embodiment, the comparison of the methylation density of a potentially aneuploid chromosome (or chromosomal region) to the overall methylation density of one or more other putatively non-aneuploid chromosomes or genomes can be used to effectively normalize the fetal DNA concentration in a plasma sample. The comparison can be by calculation of a parameter (including, for example, a ratio or difference) between the methylation densities of the two regions to obtain a normalized methylation density. The comparison can eliminate the dependency of the obtained methylation level (e.g., determined as a parameter from the two methylation densities).

[0145] If the methylation density of a potentially aneuploid chromosome is not normalized to the methylation density of one or more other chromosomes or to other parameters reflecting the fractional concentration of fetal DNA, the fractional concentration will be the main factor influencing the methylation density in plasma. For example, the plasma methylation density of chromosome 21 of a pregnant woman carrying a trisomy 21 fetus with a fractional fetal DNA concentration of 10% will be the same as that of a pregnant woman carrying a euploid fetus with a fractional fetal DNA concentration of 15%, but the normalized methylation density will show a difference.

[0146] In another embodiment, the methylation densities of potentially aneuploid chromosomes can be normalized to fractional fetal DNA concentration. For example, the following formula can be applied to normalize the methylation density:

[0147]

number

[0148] During the ceremony, M.P. Normalized (Normalized) is the methylation density normalized using the fractional fetal DNA concentration in plasma, MP non-normalized (non-normalized) is the measured methylation density, BKG is the background methylation density obtained from maternal blood cells or tissues, PLN is the methylation density in placental tissues, and f is the fetal DNA fraction. The methylation densities of BKG and PLN can be based on previously established reference values ​​from maternal blood cells and placental tissues obtained from normal pregnant women. Various genetic and epigenetic approaches can be used to determine the fetal DNA fraction in plasma samples, for example, by measuring the proportion of sequence reads from the Y chromosome using massively parallel sequencing or PCR on non-bisulfite converted DNA.

[0149] In one implementation, the normalized methylation densities of potentially aneuploid chromosomes can be compared to a reference group of pregnant women carrying euploid fetuses. The mean and SD of the normalized methylation densities of the reference group can be determined. The normalized methylation densities of the test cases can then be expressed as a Z-score indicating the number of SDs from the mean of the reference group according to the formula:

number

[0150] During the ceremony, M.P. Normalized where is the normalized methylation density of the test cases, Mean is the mean of the normalized methylation densities of the reference cases, and SD is the standard deviation of the normalized methylation densities of the reference cases. A cutoff (e.g., a Z score less than -3) can be used to classify whether a chromosome is significantly hypomethylated, thereby determining the aneuploid status of the sample.

[0151] In another embodiment, MP Diff may be used as a normalized methylation density. In such an embodiment, the PLN may be estimated, for example, using method 1000. In some implementations, the reference methylation density (which may be normalized using f) may be determined from the methylation levels of non-aneuploid regions. For example, the Mean may be determined from one or more chromosomal regions of the same sample. The cutoff may be determined according to f, or may simply be set to a level sufficient enough that there is a minimum concentration.

[0152] Thus, the comparison of the methylation level of a region to a cutoff can be accomplished in various ways. The comparison can include normalization (e.g., as described above), which can be performed equally to the methylation level or cutoff value, depending on how those values ​​are defined. Thus, whether the determined methylation level of a region is statistically different from a reference level (determined from the same sample or another sample) can be determined in various ways.

[0153] The above analysis can be applied to the analysis of chromosomal regions, which may include whole chromosomes or parts of chromosomes (e.g., adjacent or disjointed subregions of chromosomes). In one embodiment, potentially aneuploid chromosomes can be divided into several bins. The bins can be of the same or different size. The methylation density of each bin can be normalized to the concentration fraction of the sample, or to the methylation density of one or more putative non-aneuploid chromosomes or to the overall methylation density of the genome. The normalized methylation density of each bin can then be compared to a reference group to determine whether it is significantly hypomethylated. The proportion of bins that are significantly hypomethylated can then be determined. A cutoff (e.g., more than 5%, 10%, 15%, 20% or 30% of the bins are significantly hypomethylated) can be used to classify the aneuploid status of the case.

[0154] When testing for amplifications or deletions, the methylation density can be compared to a reference methylation density, which can be specific to the particular region being tested. Because methylation can vary from region to region, particularly depending on the size of the region (e.g., smaller regions show greater variation), each region can have a different reference methylation density.

[0155] As described above, one or more pregnant women each carrying a euploid fetus can be used to define the normal range of methylation density of the region of interest, or the difference in methylation density between two chromosomal regions. The normal range of PLN can also be determined (e.g., by direct measurement or as an estimate by method 1000). In other embodiments, the ratio between two methylation densities can be used, for example, the methylation densities of potentially aneuploid and non-aneuploid chromosomes can be used in the analysis instead of their difference. This methylation analysis approach can be combined with a sequence read counting approach (RWK Chiu et al. 2008 Proc Natl Acad Sci USA;105:20458-20463) and an approach involving size analysis of plasma DNA (US Patent Publication No. 2011 / 0276277) to determine or confirm aneuploidy. Sequence read counting approaches used in combination with methylation analysis can be performed using random sequencing (RWK Chiu et al. 2008 Proc Natl Acad Sci USA;105:20458-20463; DW Bianchi DW et al. 2012 Obstet Gynecol 119:890-901) or targeted sequencing (AB Sparks et al. 2012 Am J Obstet Gynecol 206:319.e1-9; B Zimmermann et al. 2012 Prenat Diagn 32:1233-1241; GJ Liao et al. 2012 PLoS One;7:e38154).

[0156] The use of BKG can account for the difference in background between samples. For example, one woman may have a different BKG methylation level than another woman, but the difference between BKG and PLN can be used across samples in such a situation. The cutoff for different chromosomal regions can be different, for example, when the methylation density of one region of the genome is different from another region of the genome.

[0157] This approach can be generalized to detect any chromosomal abnormality, including deletions and amplifications in the fetal genome. Furthermore, the resolution of this analysis can be adjusted to the desired level, for example, the genome can be divided into bins of 10Mb, 5Mb, 2Mb, 1Mb, 500kb, 100kb. Thus, this technology can also be used to detect subchromosomal duplications or deletions. Thus, this technology allows for non-invasively obtaining the molecular karyotype of the prenatal fetus. When used in this way, this technology can be used in combination with non-invasive prenatal testing methods based on molecular counting (A Srinivasan et al. 2013 Am J Hum Genet;92:167-176; SCY Yu et al. 2013 PLoS One 8: e60968). In other embodiments, the size of the bins may not be the same. For example, the size of the bins can be adjusted so that each bin contains the same number of CpG dinucleotides. In this case, the physical size of the bins will be different.

[0158] The above formula can be rewritten as follows to apply to various types of chromosomal abnormalities:

[0159]

number

[0160] Where CN represents the number of copy number changes in the affected region. CN is equal to 1 for a gain of one copy of a chromosome, 2 for a gain of two copies of a chromosome, and -1 for a loss of one of two homologous chromosomes (e.g., detection of fetal Turner syndrome, where a female fetus is missing one of the X chromosomes, resulting in an XO karyotype). This formula does not need to be modified when changing the size of the bin. However, sensitivity and specificity may decrease when a smaller bin size is used, since fewer CpG dinucleotides (or other nucleotide combinations that show differential methylation between fetal and maternal DNA) will be present in smaller bins, which increases the stochastic variation in the measurement of methylation density. In one embodiment, the number of required reads can be determined by analyzing the coefficient of variation of methylation density and the desired sensitivity level.

[0161] To demonstrate the feasibility of this approach, plasma samples from nine pregnant women were analyzed. Five pregnant women, each carrying a euploid fetus, and the other four, each carrying a trisomy 21 (T21) fetus. Three of the five euploid pregnant women were randomly selected to form a reference group. The remaining two euploid pregnancy cases (Eu1 and Eu2) and four T21 cases (T21-1, T21-2, T21-3, and T21-4) were analyzed using this approach to test for potential T21 status. Plasma DNA was bisulfite converted and sequenced using an Illumina HiSeq2000 platform. In one embodiment, the methylation density of individual chromosomes was calculated. The difference in methylation density between chromosome 21 and the average value of the other 21 autosomes was then determined to obtain the normalized methylation density (Table 1). The mean and SD of the reference group were used to calculate the Z scores for the six test cases.

[0162] [Table 1]

[0163] In another embodiment, the genome is divided into bins of 1 Mb, and the methylation density of each bin of 1 Mb is determined. The methylation density of all bins on the potentially aneuploid chromosomes can be normalized using the median methylation density of all bins located on the presumed non-aneuploid chromosomes. In one implementation, for each bin, the difference in methylation density from the median of non-aneuploid bins can be calculated. The Z-score of these values ​​can be calculated using the mean and standard deviation values ​​of the reference group. The percentage of bins that show hypomethylation (Table 2) can be determined and compared to the cutoff percentage.

[0164] [Table 2]

[0165] This DNA methylation-based approach for detecting fetal chromosomal or intrachromosomal abnormalities can be used in combination with approaches based on molecular counting, such as by sequencing (RWK Chiu et al. 2008 Proc Natl Acad Sci USA; 105: 20458-20463) or digital PCR (YMD Lo et al. 2007 Proc Natl Acad Sci USA; 104: 13116-13121), or DNA molecular size measurement (US Patent Application Publication No. 2011 / 0276277). Such combinations (e.g., DNA methylation + molecular counting, or DNA methylation + size measurement, or DNA methylation + molecular counting + size measurement) may have synergistic effects that are advantageous in clinical settings, such as improved sensitivity and / or specificity. For example, the number of DNA molecules that need to be analyzed, for example by sequencing, can be reduced without adversely affecting diagnostic accuracy. This feature allows such tests to be performed more economically. As another example, for a given number of DNA molecules analyzed, the combined approach allows fetal chromosomal or intrachromosomal abnormalities to be detected in a lower concentration fraction of fetal DNA.

[0166] 13 is a flow chart of a method 1300 for detecting chromosomal abnormalities from a biological sample of an organism. The biological sample includes cell-free DNA, including a mixture of cell-free DNA from a first tissue and a second tissue. The first tissue can be obtained from a fetus or a tumor, and the second tissue can be obtained from a pregnant woman or a patient.

[0167] At block 1310, a plurality of DNA molecules from the biological sample are analyzed. Analysis of the DNA molecules may include determining the location of the DNA molecule in the genome of the organism and determining whether the DNA molecule is methylated at one or more sites. The analysis may be performed by obtaining sequence reads from methylation-aware sequencing, so that the analysis may be performed only on data previously obtained from the DNA. In other embodiments, the analysis may include actual sequencing or other effective steps to obtain data.

[0168] Determining the location can include mapping (e.g., via sequence reads) of the DNA molecule to a portion of the human genome, for example, to a particular region. In some implementations, a read can be ignored if it does not map to the region of interest.

[0169] At block 1320, for each of the plurality of sites, the respective number of DNA molecules that are methylated at the site is determined. In one embodiment, the sites are CpG sites, and may be only certain CpG sites selected using one or more criteria described herein. The number of DNA that is methylated is equivalent to the determination of the number that is unmethylated if normalization is performed using the total number of DNA molecules analyzed at the particular site (e.g., the total number of sequence reads).

[0170] In block 1330, a first methylation level of the first chromosomal region is calculated based on the respective number of DNA molecules that are methylated at the sites in the first chromosomal region. The first chromosomal region may be of any size (e.g., the sizes described above). The methylation level may account for the total number of DNA molecules aligned to the first chromosomal region, for example, as part of a normalization procedure.

[0171] The first chromosome region may be of any size (e.g., an entire chromosome) and may consist of discrete subregions (i.e., the subregions are separate from one another). The methylation level of each subregion may be determined and then combined (e.g., as an average or median) to determine the methylation level of the first chromosome region.

[0172] In block 1340, the first methylation level is compared to a cutoff value. The cutoff value may be a reference methylation level or may be related to a reference methylation level (e.g., a particular distance from a normal level). The cutoff value may be determined from other pregnant female subjects carrying fetuses that do not have a chromosomal abnormality in the first chromosome region, from samples of individuals without cancer, or from a genetic locus of an organism that is known not to be associated with aneuploidy (i.e., a disomy region).

[0173] In one embodiment, the cutoff value may be defined as having a difference from a reference methylation level of the following formula:

[0174]

number

[0175] Where BKG is the female background (or the mean or median value from other subjects), f is the concentration fraction of cell-free DNA from the first tissue, and CN is the copy number to be examined. CN is an example of a scaling factor that corresponds to the type of abnormality (deletion or duplication). A cutoff of CN=1 can be used to first examine all amplifications, and then additional cutoffs can be used to determine the extent of amplification. The cutoff value can be based on the concentration fraction of cell-free DNA from the first tissue to determine the expected level of methylation of the locus (e.g., in the absence of copy number abnormalities).

[0176] In block 1350, a classification of the abnormality of the first chromosome region is determined based on the comparison. A statistically significant difference in the level may indicate an increased risk of the fetus having a chromosomal abnormality. In various embodiments, the chromosomal abnormality may be trisomy 21, trisomy 18, trisomy 13, Turner syndrome, or Klinefelter syndrome. Other examples are intrachromosomal deletion, intrachromosomal duplication, or DiGeorge syndrome.

[0177] V. Marker Determination As mentioned above, certain parts of the fetal genome are differentially methylated from the maternal genome. These differences may be generalized throughout pregnancy. The regions of differential methylation can be used to identify DNA fragments of fetal origin.

[0178] A. Methods for determining DMR from placental and maternal tissues Placenta has tissue-specific methylation signature. Fetal-specific DNA methylation markers have been developed for maternal plasma detection and non-invasive prenatal diagnostic applications based on differentially methylated loci between placental tissue and maternal blood cells (SSC Chim et al. 2008 Clin Chem; 54: 500-511; EA Papageorgiou et al 2009 Am J Pathol; 174: 1609-1618; and T Chu et al. 2011 PLoS One; 6: e14723). An embodiment is provided that searches for such differentially methylated regions (DMRs) genome-wide.

[0179] 14 is a flow chart of a method 1400 for identifying methylation markers by comparing a placental methylation profile to a maternal methylation profile (e.g., determined from blood cells) according to an embodiment of the invention. Method 1400 may also be used to determine tumor markers by comparing a tumor methylation profile to a methylation profile corresponding to healthy tissue.

[0180] At block 1410, the placental methylome and blood methylome are obtained. The placental methylome can be determined from a placental sample (e.g., CVS or term placenta). It is understood that the methylome may include the methylation density of only a portion of the genome.

[0181] In block 1420, a region containing a certain number of sites (e.g., 5 CpG sites) is identified for which a sufficient number of reads have been obtained. In one embodiment, the identification starts from one end of each chromosome to locate the first 500 bp region containing at least 5 suitable CpG sites. A CpG site may be considered suitable if the site is covered by at least 5 sequence reads.

[0182] A placental methylation index and a blood methylation index are calculated for each site in block 1430. For example, a methylation index was calculated for all eligible CpG sites within each 500 bp region separately.

[0183] In block 1440, the methylation indices were compared between maternal blood cells and placenta samples to determine whether the sets of indices differed from each other. For example, the methylation indices were compared between maternal blood cells and CVS or term placenta using, for example, a Mann-Whitney test. For example, a P value of 0.01 or less was considered statistically significantly different, although other values ​​may be used if a lower number reduces false positive regions.

[0184] In one embodiment, if the number of suitable CpG sites was less than 5 or if the Mann-Whitney test was not significant, the 500 bp region was moved 100 bp downstream. Regions continued to be moved downstream until the Mann-Whitney test was significant for the 500 bp region. Then the next 500 bp region was considered. If the next region was found to show statistical significance by the Mann-Whitney test, it was added to the current region unless the combined flanking region was 1,000 bp or less.

[0185] In block 1450, adjacent regions that were statistically significantly different (e.g., by Mann-Whitney test) may be merged. Note that there is a difference between the methylation indices of the two samples. In one embodiment, adjacent regions are merged if they are within a certain distance (e.g., 1,000 bp) of each other and they show similar methylation characteristics. In one implementation, the similarity of methylation characteristics between adjacent regions may be defined using any of the following: (1) showing the same trend in placental tissue with respect to maternal blood cells (e.g., both regions are more methylated in placental tissue than in blood cells); (2) having a difference in methylation density of less than 10% in adjacent regions in placental tissue; and (3) having a difference in methylation density of less than 10% in adjacent regions in maternal blood cells.

[0186] At block 1460, the methylation density of the maternal blood cell DNA and blood methylome from a placenta sample (e.g., CVS or term placental tissue) in the region is calculated. Methylation density can be determined as described herein.

[0187] In block 1470, putative DMRs are determined where the total placental methylation density and the total blood methylation density are statistically significantly different for all sites within the region. In one embodiment, all suitable CpG sites within the combined region are compared to χ 2 Get examined. 2 Tests were performed to assess whether the number of methylated cytosines as a percentage of methylated and unmethylated cytosines among all relevant CpG sites within the merged region was statistically significantly different between maternal blood cells and placental tissues. 2 In the test, a P value of 0.01 or less can be considered statistically significantly different. 2 The merged segments that showed significance upon testing were considered as putative DMRs.

[0188] At block 1480, loci where maternal blood cell DNA had a methylation density above the high cutoff or below the low cutoff were identified. In one embodiment, loci where maternal blood cell DNA had a methylation density below 20% or above 80% were identified. In other embodiments, bodily fluids other than maternal blood may be used, such as, but not limited to, saliva, uterine or cervical washings from the female reproductive tract, tears, sweat, saliva, and urine.

[0189] The key to successful development of fetal-specific DNA methylation markers in maternal plasma may be that the methylation status of maternal blood cells is as highly methylated as possible or as unmethylated as possible. This may reduce (e.g., minimize) the chance of having maternal DNA molecules that interfere with the analysis of placenta-derived fetal DNA molecules that show the opposite methylation profile. Thus, in one embodiment, candidate DMRs were selected by further sorting. Candidate hypomethylated loci were those that showed 20% or less methylation density in maternal blood cells and at least 20% higher methylation density in placental tissue. Candidate hypermethylated loci were those that showed 80% or more methylation density in maternal blood cells and at least 20% lower methylation density in placental tissue. Other percentages may be used.

[0190] At block 1490, DMRs were then identified among some loci where the placental methylation density significantly differed from the blood methylation density by comparing the difference to a threshold. In one embodiment, the threshold is 20%, so that the methylation density differed from the methylation density of maternal blood cells by at least 20%. Thus, the difference between the placental methylation density and the blood methylation density at each identified locus can be calculated. The difference can be a simple subtraction. In other embodiments, scaling factors and other functions can be used to determine the difference (e.g., the difference can be the result of a function applied to a simple subtraction).

[0191] In one implementation, using this method, 11,729 hypermethylated loci and 239,747 hypomethylated loci were identified in first-trimester placenta samples. The top 100 hypermethylated loci are listed in Appendix Table S2A. The top 100 hypomethylated loci are listed in Appendix Table S2B. Tables S2A and S2B list the chromosome, start and end positions, size of the region, methylation density in maternal blood, methylation density in placenta samples, P-values ​​(all very small), and methylation differences. Positions correspond to the reference genome hg18, which can be found at hgdownload.soe.ucsc.edu / goldenPath / hg18 / chromosomes.

[0192] We identified 11,920 hypermethylated and 204,768 hypomethylated loci from third-trimester placenta samples. The top 100 hypermethylated loci from third trimester are listed in Table S2C, and the top 100 hypomethylated loci are listed in Table S2D. We validated the enumeration of first-trimester candidates with 33 loci previously reported to be differentially methylated between maternal blood cells and first-trimester placental tissues. 79% of the 33 loci were identified as DMRs using our algorithm.

[0193] FIG. 15A is a table 1500 showing the performance of the DMR identification algorithm using first-phase data with reference to 33 previously reported first-phase markers. In the table, "a" indicates that loci 1-15 were previously reported in (RWK Chiu et al. 2007 Am J Pathol; 170:941-950 and SSC Chim et al. 2008 Clin Chem; 54:500-511); loci 16-23 were previously reported in (KC Yuen, thesis 2007, The Chinese University of Hong Kong, Hong Kong); and loci 24-33 were previously reported in (EA Papageorgiou et al. 2009 Am J Pathol; 174:1609-1618). "b" indicates that these data were obtained from the above publications. "c" indicates that the methylation density of maternal blood cells and chorionic villus specimens and their differences were observed from sequencing data based on genomic coordinates generated in this study but provided by the original study. "d" indicates that data for the locus was identified using an embodiment of method 1400 on bisulfite sequencing data without reference to the above publications by Chiu et al (2007), Chim et al (2008), Yuen (2007) and Papageorgiou et al (2009). The full length of the locus included previously reported genomic regions, but generally spanned larger regions. "e" indicates that the candidate DMRs were classified as true positive (TP) or false positive (FN) based on the need to observe a difference of more than 0.20 between the methylation density of the corresponding genomic coordinates of the DMR in maternal blood cells and chorionic villus specimens.

[0194] FIG. 15B is a table 1550 showing the performance of the DMR identification algorithm using third trimester data compared to placental samples obtained at delivery. "a" indicates that the same enumeration of 33 loci as described in FIG. 17A was used. "b" indicates that the 33 loci were previously identified from early pregnancy samples, and therefore may not be applicable to third trimester data. Thus, the bisulfite sequencing data generated in this study on term placental tissue was reviewed based on the genomic coordinates provided by the original study. Differences in methylation density between maternal blood cells and term placental tissue of more than 0.20 were used to determine whether the loci were indeed true DMRs in the third trimester. "c" indicates that the data on the loci were identified using method 1400 on the bisulfite sequencing data without reference to the publications by Chiu et al (2007), Chim et al (2008), Yuen (2007) and Papageorgiou et al (2009). The total length of locus included previously reported genomic regions, but generally spanned larger regions. "d" indicates that the candidate DMRs containing loci suitable for differential methylation in the third trimester were classified as true positive (TP) or false positive (FN) based on the need to observe a difference of more than 0.20 between the methylation density of the corresponding genomic coordinates of the DMR in maternal blood cells and term placenta tissue. For loci that were not suitable for differential methylation in the third trimester, their absence in the enumeration of DMRs or the presence of DMRs containing the locus but showing a methylation difference of less than 0.20 were considered as true negative (TN) DMRs.

[0195] B. DMRs derived from maternal plasma sequencing data If the fetal DNA concentration fraction of the sample is also available, placental tissue DMRs should be directly identifiable from maternal plasma DNA bisulfite sequencing data. Because the placenta may be the major source of fetal DNA in maternal plasma (SSC Chim et al. 2005 Proc Natl Acad Sci USA 102, 14753-14758), in this study we showed that the methylation status of fetal-specific DNA in maternal plasma correlates with the placental methylome.

[0196] Thus, instead of using a placental sample, embodiments of method 1400 may be performed using the plasma methylome to determine a predicted placental methylome. Thus, methods 1000 and 1400 may be combined to determine DMR. Using method 1000, predictive values ​​for the placental methylation profile may be determined and used in method 1400. In this analysis, examples also focus on loci that were 20% or less or 80% or more methylated in maternal blood cells.

[0197] In one implementation, to estimate loci that were hypermethylated in placental tissues compared to maternal blood cells, loci were selected that showed 20% or less methylation in maternal blood cells and 60% or more methylation by predicted values ​​with at least a 50% difference between blood cell methylation density and predicted values. To estimate loci that were hypomethylated in placental tissues compared to maternal blood cells, loci were selected that showed 80% or more methylation in maternal blood cells and 40% or less methylation by predicted values ​​with at least a 50% difference between blood cell methylation density and predicted values.

[0198] FIG. 16 is a table 1600 showing the number of loci predicted to be hypermethylated or hypomethylated based on direct analysis of maternal plasma bisulfite sequencing data. "N / A" means not applicable. "a" indicates that the search for hypermethylated loci started from the list of loci showing less than 20% methylation density in maternal blood cells. "b" indicates that the search for hypomethylated loci started from the list of loci showing more than 80% methylation density in maternal blood cells. "c" indicates that bisulfite sequencing data obtained from chorionic villus specimens was used to validate the first trimester maternal plasma data and term placental tissue was used to validate the third trimester maternal plasma data.

[0199] As shown in Table 1600, the majority of the non-invasively estimated loci showed the expected methylation patterns in tissue and overlapped with the DMRs mined from tissue data and shown in the previous section. The appendix lists the DMRs identified from plasma. Table S3A lists the top 100 loci estimated to be hypermethylated from the bisulfite sequencing data of first trimester maternal plasma. Table S3B lists the top 100 loci estimated to be hypomethylated from the bisulfite sequencing data of first trimester maternal plasma. Table S3C lists the top 100 loci estimated to be hypermethylated from the bisulfite sequencing data of third trimester maternal plasma. Table S3B lists the top 100 loci estimated to be hypomethylated from the bisulfite sequencing data of third trimester maternal plasma.

[0200] C. Gestational changes in the placental and fetal methylome The overall percentage of methylated CpGs in the CVS was 55%, whereas the overall percentage of methylated CpGs in the term placenta was 59% (Table 100 in Figure 1). Although more hypomethylated DMRs could be identified from the CVS than the term placenta, the number of hypermethylated DMRs was similar in the two tissues. Thus, it was clear that the CVS was more hypomethylated than the term placenta. This pregnancy trend was also evident in the maternal plasma data. The percentage of methylated CpGs among fetal-specific reads was 47.0% in first trimester maternal plasma, whereas it was 53.3% in third trimester maternal plasma. The number of identified hypermethylated loci was similar in the first (1,457 loci) and third trimester (1,279 loci) maternal plasma samples, but there were substantially more hypomethylated loci in the first trimester samples (21,812 loci) than in the third trimester samples (12,677 loci) (Table 1600 in Figure 16).

[0201] D. Use of markers Differentially methylated markers, or DMRs, are useful in some embodiments. The presence of such markers in maternal plasma indicates and confirms the presence of fetal or placental DNA. This confirmation can be used as a quality control for non-invasive prenatal testing. DMRs serve as general fetal DNA markers in maternal plasma and may have advantages over markers that rely on genotype differences between mother and fetus, such as polymorphism-based markers or Y-chromosome-based markers. DMRs are general fetal markers that are useful for all pregnancies. Polymorphism-based markers are only applicable to some pregnancies in which the fetus inherits the marker from its father and the mother does not carry the marker in her genome. Furthermore, fetal DNA concentration can be measured in maternal plasma samples by quantifying DNA molecules derived from those DMRs. Knowing the characteristics of DMRs expected in normal pregnancy, pregnancy-related complications, especially pregnancy-related complications involving placental tissue changes, can be detected by observing the deviation of maternal plasma DMR or methylation characteristics from the characteristics expected in normal pregnancy. Pregnancy-related complications involving placental tissue changes include, but are not limited to, fetal chromosomal aneuploidies, such as trisomy 21, preeclampsia, intrauterine growth restriction, and preterm birth.

[0202] E. Marker-Based Kits The embodiments may provide compositions and kits for carrying out the methods described herein and other applicable methods. The kits may be used to carry out assays that analyze fetal DNA (e.g., cell-free fetal DNA in maternal plasma). In one embodiment, the kits may include at least one oligonucleotide useful for specific hybridization with one or more loci identified herein. The kits may also include at least one oligonucleotide useful for specific hybridization with one or more reference loci. In one embodiment, placental hypermethylation markers are measured. The test locus may be methylated DNA in maternal plasma, and the reference locus may be methylated DNA in maternal plasma. Similar kits may be constructed to analyze tumor DNA in plasma.

[0203] In some examples, the kit may include at least two oligonucleotide primers that can be used to amplify at least one section of a target locus (e.g., a locus in the appendix) and a reference locus. Instead of or in addition to primers, the kit may include a labeled probe for detecting DNA fragments corresponding to the target locus and the reference locus. In various embodiments, one or more oligonucleotides in the kit correspond to a locus in the appendix table. Typically, the kit also provides instructions to guide the user in analyzing the test sample and evaluating the physiological function or pathological state in the test subject.

[0204] In various embodiments, a kit is provided for analyzing fetal DNA in a biological sample containing a mixture of fetal DNA and DNA from a female subject carrying a fetus.The kit may include one or more oligonucleotides that specifically hybridize to at least a section of the genomic region listed in Tables S2A, S2B, S2C, S2D, S3A, S3B, S3C, and S3D.Thus, any number of oligonucleotides from all of these tables or are from only one table can be used.An oligonucleotide can function as a primer and can be constructed as a primer pair that corresponds to a specific region in a table.

[0205] VI. Relationship between size and methylation density Plasma DNA molecules are known to exist in the circulating blood in the form of short molecules, with the majority of molecules having a length of approximately 160 bp (YMD Lo et al. 2010 Sci Transl Med; 2: 61ra91, YW Zheng at al. 2012 Clin Chem; 58: 549-558). Interestingly, our data revealed a correlation between the methylation status and size of plasma DNA molecules. Thus, plasma DNA fragment length is associated with DNA methylation levels. The characteristic size properties of plasma DNA molecules suggest that the majority are associated with mononucleosomes, which may result from enzymatic degradation during apoptosis.

[0206] Circulating DNA is naturally fragmented. Specifically, circulating fetal DNA is shorter than maternal DNA in maternal plasma samples (KCA Chan et al. 2004 Clin Chem; 50: 88-92). Since paired-end alignment allows for size analysis of bisulfite-treated DNA, it is possible to directly assess whether a correlation exists between the size of plasma DNA molecules and their respective methylation levels. This was examined in maternal plasma and control plasma samples from non-pregnant adult women.

[0207] Paired-end sequencing (including whole molecule sequencing) of both ends of each DNA molecule was used to analyze each sample in this study. By aligning the pair of terminal sequences of each DNA molecule to the reference human genome and recording the genomic coordinates of the extreme ends of the sequenced reads, the length of the sequenced DNA molecule can be determined. Plasma DNA molecules are naturally fragmented into small molecules, and plasma DNA sequencing libraries are typically created without any fragmentation step. Therefore, the length estimated by sequencing was representative of the size of the original plasma DNA molecule.

[0208] In a previous study, the size characteristics of fetal and maternal DNA molecules in maternal plasma were determined (YMD Lo et al. 2010 Sci Transl Med; 2: 61ra91). It was shown that plasma DNA molecules had a size similar to mononucleosomes and fetal DNA molecules were shorter than maternal DNA molecules. In this study, the relationship of the methylation status of plasma DNA molecules to their size was determined.

[0209] A. Results Figure 17A is a plot 1700 showing the size distribution of maternal plasma DNA, non-pregnant female control plasma DNA, placental DNA and peripheral blood DNA. In the maternal samples and non-pregnant female control plasma, these two bisulfite treated plasma samples showed the same characteristic size distribution as previously reported (YMD Lo et al. 2010 Sci Transl Med; 2: 61ra91), with the most abundant overall sequences of 166-167 bp in length and a 10 bp periodicity of DNA molecules shorter than 143 bp.

[0210] FIG. 17B is a plot 1750 of the size distribution and methylation profile of maternal plasma, adult female control plasma, placental tissue, and adult female control blood. For DNA molecules having the same size and containing at least one CpG site, their average methylation density was calculated. The association between the size of the DNA molecules and their methylation density was then plotted. Specifically, for sequenced reads covering at least one CpG site, the average methylation density of each fragment length ranging from 50 bp up to 180 bp was determined. Interestingly, the methylation density increased with plasma DNA size, maximizing at approximately 166-167 bp. However, this pattern was not observed in placental DNA samples and control blood DNA samples fragmented using the sonication system.

[0211] Figure 18 shows plots of methylation density and size of plasma DNA molecules. Figure 18A is a plot 1800 for first trimester maternal plasma. Figure 18B is a plot 1850 for third trimester maternal plasma. Data for all sequenced reads covering at least one CpG site are represented by the blue curve 1805. Data for reads that also contained fetal-specific SNP alleles are represented by the red curve 1810. Data for reads that also contained maternal-specific SNP alleles are represented by the green curve 1815.

[0212] The reads that contained fetal-specific SNP alleles were considered to be from fetal DNA molecules. The reads that contained maternal-specific SNP alleles were considered to be from maternal DNA molecules. In general, DNA molecules with high methylation density were longer in size. This trend was observed in both fetal and maternal DNA molecules in both the first and third trimesters. The overall size of fetal DNA molecules was shorter than that of maternal DNA molecules reported previously.

[0213] FIG. 19A shows a plot 1900 of methylation density and size of sequenced reads from adult non-pregnant women. Plasma DNA samples obtained from adult non-pregnant women also showed the same association between size and methylation state of DNA molecules. Genomic DNA samples, on the other hand, were fragmented by a sonication step prior to MPS analysis. As shown in plot 1900, data obtained from blood cells and placental tissue specimens did not show the same trend. Since cell fragmentation is artificial, no association between size and density is expected. Since naturally fragmented DNA molecules in plasma show a size dependency, it can be hypothesized that the lower the methylation density, the more likely the molecules are cut into smaller fragments.

[0214] Figure 19B is a plot 1950 showing the size distribution and methylation profile of fetal-specific and maternal-specific DNA molecules in maternal plasma.Fetal-specific and maternal-specific plasma DNA molecules also show the same correlation between fragment size and methylation level.The fragment length of both placenta-derived and maternal circulating cell-free DNA increases with methylation level.Furthermore, the distributions of their methylation status do not overlap with each other, suggesting that the phenomenon occurs regardless of the original fragment length of the source of circulating DNA molecules.

[0215] B. Method Thus, the size distribution can be used to estimate the overall methylation rate of the plasma sample. This methylation measurement can then be followed during pregnancy, cancer monitoring, or during treatment by serial measurements of the size distribution of plasma DNA, according to the correlation shown in Figures 18A and 18B. The methylation measurement can also be used to look for increased or decreased DNA release from an organ or tissue of interest. For example, one can specifically look for DNA methylation signatures that are unique to a particular organ (e.g., the liver) and measure the concentration of these signatures in the plasma. Since DNA is released into the plasma as cells die, increased levels can mean increased cell death or cell damage in that particular organ or tissue. Decreased levels from a particular organ can mean that treatments to combat injury or pathological processes in that organ are under control.

[0216] Figure 20 is a flow chart of a method 2000 for estimating the methylation level of DNA in a biological sample of an organism according to an embodiment of the present invention. The methylation level of a specific region of a genome or the entire genome can be estimated. If a specific region is desired, DNA fragments derived only from that specific region can be used.

[0217] In block 2010, the amount of DNA fragments corresponding to various sizes is measured. For each size of the plurality of sizes, the amount of the plurality of DNA fragments from the biological sample corresponding to the size can be measured. For example, the number of DNA fragments having a length of 140 bases can be measured. The amount can be recorded as a histogram. In one embodiment, the size of each of the plurality of nucleic acids from the biological sample is measured, which can be performed individually (e.g., by single molecule sequencing of the entire molecule or only the ends of the molecule) or collectively (e.g., by electrophoresis). The size can correspond to a range. Thus, the amount can be the amount of DNA fragments having a size within a particular range. If paired-end sequencing is performed, the DNA fragments that are located (aligned) to a particular region (determined by paired sequence reads) can be used to determine the methylation level of the region.

[0218] In block 2020, a first value of a first parameter is calculated based on the amount of DNA fragments at a plurality of sizes. In one embodiment, the first parameter provides a statistical measure of the size profile (e.g., a histogram) of DNA fragments in the biological sample. The parameter may be referred to as a size parameter because it is determined from the sizes of a plurality of DNA fragments.

[0219] The first parameter can be a parameter of various forms. One parameter is the percentage of DNA fragments of a certain size or size range to the total DNA fragments or to DNA fragments of another size or range. Such a parameter is the number of DNA fragments of a certain size divided by the total number of fragments, which can be obtained from a histogram (any data structure that gives the absolute or relative count of fragments of a certain size). As another example, the parameter can be the number of fragments of a certain size or within a certain range divided by the number of fragments of another size or range. The division can serve as a normalization function to account for the different number of DNA fragments analyzed for different samples. The normalization can be achieved by analyzing the same number of DNA fragments for each sample, which effectively gives the same result as division by the total number of fragments analyzed. Further examples of parameters and size analysis can be found in US Patent Application Publication No. 13 / 789,553 (incorporated by reference for all purposes).

[0220] In block 2030, the first size value is compared with a reference size value. The reference size value can be calculated from the DNA fragment of the reference sample. To determine the reference size value, the methylation profile of the reference sample can be calculated and quantified, and the value of the first size parameter can be calculated and quantified as well. Thus, when the first size value is compared with the reference size value, the methylation level can be determined.

[0221] At block 2040, a methylation level is estimated based on the comparison. In one embodiment, it is determined whether a first value of a first parameter is above or below a reference size value, which may determine whether a methylation level of a current sample is above or below a methylation level for a reference size value. In another embodiment, the comparison is accomplished by inputting the first value into a calibration function. The calibration function may effectively compare the first value to a calibration value (a set of reference size values) by identifying a point on a curve that corresponds to the first value. The estimated methylation level is then provided as an output value of the calibration function.

[0222] Therefore, size parameters can be calibrated to methylation levels. For example, methylation levels can be measured and associated with the particular size parameters of the sample. Data points from different samples can then be fitted to the calibration function. In some implementations, different calibration functions can be used for different DNA subsets. Thus, there can be some forms of calibration based on prior knowledge of the relationship between methylation and size of particular DNA subsets. For example, the calibration of fetal DNA and maternal DNA can be different.

[0223] As shown above, the placenta is more hypomethylated compared to maternal blood, and therefore fetal DNA is smaller due to its lower degree of methylation. Thus, the average fragment size (or other statistics) of the sample can be used to estimate methylation density. Since fragment size can be measured using paired-end sequencing rather than methylation-aware sequencing, which can be technically more complex, this approach can be cost-effective when used clinically. This approach can be used to monitor methylation changes associated with pregnancy progression or pregnancy-related diseases such as pre-eclampsia, preterm birth, and fetal disorders (e.g., disorders caused by chromosomal or genetic abnormalities or intrauterine growth retardation).

[0224] In another embodiment, this approach can be used to detect and monitor cancer. For example, successful treatment of cancer will change the methylation profile in plasma or another body fluid measured using this size-based approach toward that of a healthy individual without cancer. Conversely, if cancer is progressing, the methylation profile in plasma or another body fluid will deviate from that of a healthy individual without cancer.

[0225] In summary, hypomethylated molecules were shorter than hypermethylated molecules in plasma. The same trend was observed for both fetal and maternal DNA molecules. As DNA methylation is known to affect nucleosome packing, our data suggests that hypomethylated DNA molecules are probably less densely packed with histones and therefore more susceptible to enzymatic degradation. Meanwhile, the data shown in Figure 18A and Figure 18B also showed that the size distributions of fetal and maternal DNA do not completely separate from each other, even though fetal DNA is much more hypomethylated than maternal reads. In Figure 19B, it can be seen that the methylation levels of fetal-specific and maternal-specific reads are different from each other, even in the same size category. This observation suggests that the hypomethylation state of fetal DNA is not the only factor explaining its relative shortness to maternal DNA.

[0226] VII. Imprinting status of the locus Fetal DNA molecules that share the same genotype as the mother but have a different epigenetic signature can be detected in maternal plasma (LLM Poon et al. 2002 Clin Chem; 48: 35-41). To show that the sequencing approach is sensitive in capturing fetal DNA molecules in maternal plasma, the same strategy was applied to the detection of imprinted fetal alleles in maternal plasma samples. Two genomic imprinted regions were identified: H19 (chromosome 11: 1,977,419-1,977,821, NCBI Build36 / hg18) and MEST (chromosome 7: 129,917,976-129,920,347, NCBI Build36 / hg18). Both of them contain informative SNPs to distinguish maternal and fetal sequences. For the maternally expressed gene H19, the mother was homozygous (A / A) and the fetus was heterozygous (A / C) for the intraregional SNP rs2071094 (chromosome 11:1,977,740). One of the A maternal alleles was fully methylated and the other unmethylated. However, in the placenta, the A allele was unmethylated, whereas the paternally inherited C allele was fully methylated. Two methylated reads with the C genotype corresponding to the imprinted paternal allele from the placenta were detected in maternal plasma.

[0227] MEST, also known as PEG1, is a paternally expressed gene. Both the mother and fetus were heterozygous (A / G) for the SNP rs2301335 (chromosome 7:129,920,062) within the imprinted locus. In maternal blood, the G allele was methylated, whereas the A allele was unmethylated. The methylation pattern in the placenta was reversed, with the maternal A allele being methylated and the paternal G allele being unmethylated. Three unmethylated G alleles from the father were detectable in maternal plasma. In contrast, VAV1, a non-imprinted locus on chromosome 19 (chromosome 19:6,723,621-6,724,121), did not show any allele methylation pattern in tissue and plasma DNA samples.

[0228] Therefore, the methylation status can be used to determine which DNA fragments are of fetal origin. For example, detection of the A allele alone in maternal plasma cannot be used as a fetal marker when the mother is GA heterozygous. However, when identifying the methylation status of A molecules in plasma, methylated A molecules are fetal specific, while unmethylated A molecules are maternal specific, or vice versa.

[0229] We then focused on loci that have been reported to show genomic imprinting in placental tissues. Based on the list of loci reported by Woodfine et al. (2011 Epigenetics Chromatin; 4: 1), we further selected loci that contained SNPs within imprinting control regions. Four loci met the criteria, these were H19, KCNQ10T1, MEST and NESP.

[0230] For the maternal blood cell sample reads of H19 and KCNQ10T1, these maternal reads were homozygous for the SNPs, with approximately equal proportions of methylated and unmethylated reads. CVS and term placental tissue samples revealed that the fetus was heterozygous for both loci, with each allele exclusively methylated or unmethylated, i.e., monoallelic methylation. In the maternal plasma samples, paternally inherited fetal DNA molecules were detected at both loci. In H19, the paternally inherited molecules were represented by sequenced reads that contained fetal-specific alleles and were methylated. In KCNQ10T1, the paternally inherited molecules were represented by sequenced reads that contained fetal-specific alleles and were unmethylated.

[0231] Meanwhile, the mother was heterozygous for both MEST and NESP. For MEST, both the mother and fetus were GA heterozygous for the SNP. However, the methylation status of the CpG adjacent to the SNP was opposite in the mother and fetus, as revealed by the Watson strand data of maternal blood cells and placental tissue. The A allele was unmethylated in the mother's DNA, but methylated in the fetal DNA. For MEST, the maternal allele was methylated. Thus, it can be determined that the fetus had the A allele inherited from its mother (methylated in CVS) and the mother had the A allele inherited from her father (unmethylated in maternal blood cells). Interestingly, in the maternal plasma samples, all four molecular groups could be easily distinguished (e.g., each of the two alleles in the mother and each of the two alleles in the fetus). Thus, by combining genotypic information with the methylation status at imprinted loci, fetal DNA molecules inherited from the mother could be easily distinguished from background maternal DNA molecules (LLM Poon et al. 2002 Clin Chem; 48: 35-41).

[0232] This approach can be used to detect uniparental disomy. For example, if the father of the fetus is known to be homozygous for the G allele, the inability to detect the unmethylated G allele in maternal plasma indicates the absence of a paternal allele contribution. Furthermore, under such circumstances, if both the methylated G allele and the methylated A allele are detected in the plasma of the pregnant woman, it suggests that the fetus has maternal heterodisomy, i.e., inherited two different alleles from the mother without inheritance from the father. Alternatively, if both the methylated A allele (the fetal allele inherited from the mother) and the unmethylated A allele (the maternal allele inherited from the maternal grandfather) are detected in maternal plasma without the unmethylated G allele (the paternal allele that would have been inherited by the fetus), it suggests that the fetus has maternal isodisomy, i.e., inherited two identical alleles from the mother and none from the father.

[0233] For NESP, the mother was GA heterozygous at the SNP, while the fetus was homozygous for the G allele. In NESP, the paternal allele was methylated. In maternal plasma samples, the methylated, paternally inherited fetal G allele could be easily distinguished from the unmethylated background maternal G allele.

[0234] VIII. Cancer / Donor Some embodiments can be used to detect, screen, monitor (e.g., for recurrence, remission, or response to treatment (e.g., presence or absence)), stage, classify (e.g., to aid in selection of the most appropriate treatment), and determine prognosis of cancer using methylation analysis of circulating plasma / serum DNA.

[0235] Cancer DNA is known to show aberrant DNA methylation (JG Herman et al. 2003 N Engl J Med; 349: 2042-2054). For example, compared to non-cancer cells, CpG island promoters of genes (e.g., tumor suppressor genes) are hypermethylated, while CpG sites in the gene body are hypomethylated. If the methylation profile of cancer cells can be reflected in the methylation profile of tumor-derived plasma DNA molecules using the methods described herein, it is expected that the global methylation profile in plasma will be different between individuals with cancer when compared to healthy individuals without cancer, or when compared to those who have been cured of cancer. The type of difference in methylation profile can be in terms of quantitative differences in methylation density of genome and / or methylation density of genomic sections. For example, due to the universally hypomethylated nature of DNA from cancer tissues (Gama-Sosa MA et al. 1983 Nucleic Acids Res; 11: 6883-6894), a decrease in methylation density in the plasma methylome or genomic compartments would be observed in the plasma of cancer patients.

[0236] Quantitative changes in methylation profiles should also be reflected among plasma methylome data. For example, plasma DNA molecules derived from genes that are hypermethylated only in cancer cells will show hypermethylation in the plasma of cancer patients when compared to plasma DNA molecules derived from the same gene but in healthy control samples. Since aberrant methylation occurs in the majority of cancers, the methods described herein can be applied to the detection of any form of malignant tumors with aberrant methylation, including but not limited to those in the lung, breast, colorectum, prostate, nasopharynx, stomach, testis, skin, nervous system, bone, ovary, liver, blood system tissues, pancreas, uterus, kidney, bladder, lymphatic tissues, etc. The malignant tumors may be of various histological subtypes, such as carcinoma, adenocarcinoma, sarcoma, fibroadenocarcinoma, neuroendocrine, and undifferentiated carcinoma.

[0237] On the other hand, it is expected that tumor-derived DNA molecules can be distinguished from background non-tumor-derived DNA molecules, because the overall short size characteristic of tumor-derived DNA is accentuated in DNA molecules from loci with tumor-associated abnormal hypomethylation, which will have an additional effect on the size of the DNA molecules. Also, tumor-derived plasma DNA molecules can be distinguished from background non-tumor-derived plasma DNA molecules using multiple characteristics associated with tumor DNA, such as, but not limited to, single nucleotide mutations, copy number gains and losses, translocations, inversions, abnormal hyper- or hypomethylation, and size characteristics. Since all of these changes can occur independently, the combination of these features may provide additional advantages for sensitive and specific detection of cancer DNA in plasma.

[0238] A. Size and Cancer The size of tumor-derived DNA molecules in plasma also resembles the size of a mononucleosomal unit and is shorter than background non-tumor-derived DNA molecules that are simultaneously present in the plasma of cancer patients. Size parameters have been shown to be associated with cancer, as described in U.S. Patent Application Publication No. 13 / 789,553 (incorporated by reference for all purposes).

[0239] Because both fetal and maternal DNA in plasma showed an association between molecular size and methylation status, tumor-derived DNA molecules are expected to show the same trends: hypomethylated molecules, for example, will be shorter than hypermethylated molecules in the plasma of cancer patients or subjects undergoing cancer testing.

[0240] B. Methylation density in various tissues in cancer patients In this example, plasma and tissue samples from hepatocellular carcinoma (HCC) patients were analyzed. Blood samples were collected from HCC patients before and one week after surgical resection of the tumor. Plasma and buffy coats were collected after centrifugation of the blood samples. Resected tumors and adjacent non-tumorous liver tissues were collected. DNA samples extracted from plasma and tissue samples were analyzed using massively parallel sequencing with and without prior bisulfite treatment. Plasma DNA from four healthy individuals without cancer was also analyzed as controls. Bisulfite treatment of DNA samples converts unmethylated cytosine residues to uracil. In downstream polymerase chain reaction and sequencing, these uracil residues behave as thymidine. Meanwhile, bisulfite treatment does not convert methylated cytosine residues to uracil. After massively parallel sequencing, the sequencing reads were analyzed using Methy-Pipe (P Jiang, et al. Methy-Pipe: An integrated bioinformatics data analysis pipeline for whole genome methylome analysis, paper presented at the IEEE International Conference on Bioinformatics and Biomedicine Workshops, Hong Kong, 18 to 21 December 2010) to determine the methylation status of cytosine residues at all CG dinucleotide positions (i.e., CpG sites).

[0241] FIG. 21A is a table 2100 showing the methylation density of preoperative plasma and tissue samples of HCC patients. The CpG methylation density of a region of interest (e.g., CpG site, promoter, or repeat region, etc.) refers to the percentage of reads that show CpG methylation relative to the total number of reads that cover the CpG dinucleotides of the genome. The methylation density of buffy coat and non-tumorous liver tissues was similar. The overall methylation density of tumor tissues based on data from all autosomes was 25% lower than that of buffy coat and non-tumorous liver tissues. Hypomethylation was consistent across all individual chromosomes. The methylation density of plasma was between the values ​​of non-malignant and cancer tissues. This observation is consistent with the fact that both cancer and non-cancerous tissues contribute to the circulating DNA of cancer patients. It has also been shown that the hematopoietic system is the main source of circulating DNA in individuals without active malignant conditions (YYN Lui, et al. 2002 Clin Chem; 48: 421-7). Therefore, plasma samples from four healthy controls were also analyzed. The number of sequence reads and sequencing depth achieved per sample are shown in Table 2150 of FIG. 21B.

[0242] FIG. 22 is a table 220 showing that methylation density in autosomes ranged from 71.2% to 72.5% in plasma samples from healthy controls. These data show the expected levels of DNA methylation in plasma samples obtained from individuals without a source of tumor DNA. In cancer patients, tumor tissue also sheds DNA into the circulation (KCA Chan et al. 2013 Clin Chem; 59: 211-224); RJ Leary et al. 2012 Sci Transl Med; 4: 162ra154). Due to the hypomethylated nature of HCC tumors, the presence of both tumor- and non-tumor-derived DNA in the patient's pre-op plasma results in a decrease in methylation density when compared to the plasma levels of healthy controls. Indeed, the methylation density of the pre-op plasma samples was between that of the tumor tissue and that of the healthy control plasma. The reason is that the methylation level of plasma DNA in cancer patients is influenced by the degree of abnormal methylation (in this case hypomethylation) in tumor tissue and the concentration fraction of tumor-derived DNA in circulating blood. The lower methylation density of tumor tissue and the higher concentration fraction of tumor-derived DNA in circulating blood result in a lower methylation density of plasma DNA in cancer patients. Most tumors have been reported to show global hypomethylation (JG Herman et al. 2003 N Engl J Med; 349: 2042-2054; MA Gama-Sosa et al. 1983 Nucleic Acids Res; 11: 6883-6894). Therefore, the current findings seen in HCC samples should be applicable to other types of tumors.

[0243] In one embodiment, the methylation density of plasma DNA can be used to determine the fractional concentration of tumor-derived DNA in plasma / serum samples when the methylation level of tumor tissue is known. The methylation level (e.g., methylation density) of tumor tissue can be obtained when a tumor sample is available or a tumor biopsy is available. In another embodiment, information on the methylation level of tumor tissue can be obtained from a survey of methylation levels in a group of tumors of similar type, and this information (e.g., average or median levels) is applied to the patient analyzed using the science and technology described in the present invention. The methylation level of tumor tissue can be determined by analysis of the patient's tumor tissue or inferred from analysis of tumor tissue of other patients with the same or similar cancer type. Methylation of tumor tissue can be determined using a range of methylation recognition platforms, including but not limited to massively parallel sequencing, single molecule sequencing, microarrays (e.g., oligonucleotide arrays), or mass spectrometry (e.g., Epityper analysis, Sequenom, Inc.). In some embodiments, such analysis may be preceded by procedures sensitive to the methylation status of DNA molecules, such as, but not limited to, cytosine immunoprecipitation and methylation-recognizing restriction enzyme digestion. If the methylation level of the tumor is known, the fractional concentration of tumor DNA in the plasma of cancer patients can be calculated after plasma methylome analysis.

[0244] The association between plasma methylation level (P), tumor DNA concentration fraction (f), and tumor tissue methylation level (TUM) can be described as P=BKG×(1-f)+TUM×f, where BKG is the background DNA methylation level in plasma from blood cells and other internal organs. For example, the global methylation density of all autosomes was shown to be 42.9% in tumor biopsy tissue obtained from this HCC patient (i.e., TUM value in this case). The average methylation density of plasma samples obtained from four healthy controls (i.e., BKG value in this case) was 71.6%. The plasma methylation density of pre-operative plasma was 59.7%. Using these values, f is estimated to be 41.5%.

[0245] In another embodiment, the methylation level of tumor tissue can be estimated non-invasively based on plasma methylome data when the concentration fraction of tumor-derived DNA in plasma samples is known. The concentration fraction of tumor-derived DNA in plasma samples can be determined by other genetic analyses, such as genome-wide analysis of allelic deletions (GAAL) and analysis of single nucleotide mutations, as previously described (US Patent Application Publication No. 13 / 308,473; KCA Chan et al. 2013 Clin Chem; 59: 211-24). The calculation is based on the same associations described above, except that in this embodiment, the value of f is known and the value of TUM is unknown. The estimation can be performed for the whole genome or for parts of the genome, similar to the data observed in the context of determining placental tissue methylation levels from maternal plasma data.

[0246] In another embodiment, the intra-bin variation or characteristics of methylation density can be used to distinguish subjects with cancer from those without cancer. The resolution of methylation analysis can be further increased by dividing the genome into bins of a certain size (e.g., 1 Mb). In such an embodiment, the methylation density of each 1 Mb bin was calculated for collected samples, such as buffy coat, resected HCC tissue, non-tumorous liver tissue adjacent to tumor, and plasma collected before and after tumor resection. In another embodiment, the size of the bins may not be constant. In some implementations, the bins themselves may vary in size, but the number of CpG sites is constant within each bin.

[0247] Figures 23A and 23B show methylation densities of buffy coat, tumor tissue, non-tumorous liver tissue, pre-surgery plasma, and post-surgery plasma of HCC patients. Figure 23A is a plot 2300 of the results for chromosome 1. Figure 23B is a plot 2350 of the results for chromosome 2.

[0248] In most of the 1 Mb window, the methylation density of the buffy coat and the non-tumorous liver tissue adjacent to the tumor was similar, while the methylation density of the tumor tissue was lower. The methylation density of the preoperative plasma is between that of the tumor tissue and the non-malignant tissue. The methylation density of the matched genomic region in the tumor tissue can be estimated using the methylation data of the preoperative plasma and the tumor DNA concentration fraction. The method is the same as above using the methylation density values ​​of all autosomes. The described tumor methylation estimation can also be performed using this higher resolution plasma DNA methylation data. Other bin sizes such as 300 kb, 500 kb, 2 Mb, 3 Mb, 5 Mb or more than 5 Mb can also be used. In one embodiment, the size of the bins may not be constant. In one implementation, the bins themselves may vary in size, but the number of CpG sites is constant within each bin.

[0249] C. Comparison of plasma methylation density between cancer patients and healthy controls As shown in FIG. 2100, the methylation density of pre-surgery plasma DNA was lower than that of non-malignant tissues in cancer patients. This may be due to the presence of DNA from tumor tissues that was hypomethylated. This lower plasma DNA methylation density can potentially be used as a biomarker for cancer detection and monitoring. In cancer monitoring, when cancer progresses, the amount of cancer-derived DNA in plasma increases over time. In this example, the increase in the amount of circulating cancer-derived DNA in plasma leads to a further decrease in plasma DNA methylation density at the genome-wide level.

[0250] Conversely, if the cancer is responding to treatment, the amount of cancer-derived DNA in plasma decreases over time. In this example, the decrease in the amount of cancer-derived DNA in plasma leads to an increase in plasma DNA methylation density. For example, when a lung cancer patient with an epidermal growth factor receptor mutation is treated with a targeted therapy (e.g., tyrosine kinase inhibition), an increase in plasma DNA methylation density indicates a response. The subsequent emergence of tumor clones resistant to tyrosine kinase inhibition is associated with a decrease in plasma DNA methylation density, indicating recurrence.

[0251] Plasma methylation density measurements can be performed serially and the rate of change of such measurements can be calculated and used to predict or correlate with clinical progression or remission or prognosis. At selected genomic loci (e.g., promoter regions of several tumor suppressor genes) that are hypermethylated in cancer tissue but hypomethylated in normal tissue, the association between cancer progression and favorable response to treatment is opposite to the pattern described above.

[0252] To demonstrate the feasibility of this approach, we compared DNA methylation densities of plasma samples taken from cancer patients before and after surgical resection of their tumors with plasma DNA from four healthy control patients.

[0253] Table 2200 shows the DNA methylation density of each autosome and the combined values ​​of autosomes for all pre- and post-surgery plasma samples of cancer patients and four healthy control patients.For all chromosomes, the methylation density of pre-surgery plasma DNA samples is lower than that of post-surgery samples and plasma samples from four healthy subjects.The difference in plasma DNA methylation density between pre-surgery samples and post-surgery samples provides supporting evidence that the lower methylation density in pre-surgery plasma samples was due to the presence of DNA from HCC tumors.

[0254] Restoration of DNA methylation density in post-surgery plasma samples to levels similar to those in healthy control plasma samples suggested that the majority of tumor-derived DNA was lost upon surgical resection of the source (i.e., tumor). These data suggest that the methylation density of pre-surgery plasma, determined using data available from large genomic regions (e.g., entire autosomes or individual chromosomes), may have lower methylation levels than healthy control samples, allowing identification (i.e., diagnosis or screening) of test cases with cancer.

[0255] The data from pre-surgery plasma also showed that plasma methylation levels can be used to monitor tumor burden and thus prognosticate and monitor cancer progression in patients, with even lower methylation levels than those in post-surgery plasma. Reference values ​​can be determined from plasma of healthy controls or individuals at risk for cancer but who do not currently have cancer. Individuals at risk for HCC include those with chronic hepatitis B or C infection, those with hemochromatosis, and those with liver cirrhosis.

[0256] Plasma methylation density values ​​above (e.g., lower than) a predefined cutoff based on the reference value can be used to assess whether the plasma of a non-pregnant woman has tumor DNA. To detect the presence of hypomethylated circulating tumor DNA, the cutoff can be defined as lower than the 5th or 1st percentile of the values ​​of the control population, or based on the number of standard deviations below the mean methylation density value of the control, e.g., 2 or 3 standard deviations (SD), or based on the determination of the median multiple (MoM). For hypermethylated tumor DNA, the cutoff can be defined as higher than the 95th or 99th percentile of the values ​​of the control population, or based on the number of standard deviations above the mean methylation density value of the control, e.g., 2 or 3 SD, or based on the determination of the median multiple (MoM). In one embodiment, the control population is age-matched to the test subject. Age matching does not have to be strict and can be performed in age bands (e.g., 30-40 years old for a test subject aged 35 years old).

[0257] We then compared the methylation densities of 1 Mb bins between plasma samples from cancer patients and four control patients. For illustration purposes, results for chromosome 1 are shown.

[0258] Figure 24A is a plot 2400 showing the methylation density of pre-surgery plasma obtained from an HCC patient. Figure 24B is a plot 2450 showing the methylation density of post-surgery plasma obtained from an HCC patient. The blue dots represent the results of the control patients and the red dots represent the results of the HCC patient plasma samples.

[0259] As shown in Figure 24A, the methylation density of pre-surgery plasma obtained from HCC patients was lower than that of control patients in most bins. Similar patterns were observed in other chromosomes. As shown in Figure 24B, the methylation density of post-surgery plasma obtained from HCC patients was similar to that of control patients in most bins. Similar patterns were observed in other chromosomes.

[0260] To assess whether the test subject has cancer, the test subject's results are compared with the values ​​of the reference group. In one embodiment, the reference group is composed of some healthy subjects. In another embodiment, the reference group can be composed of subjects with non-malignant conditions (e.g., chronic hepatitis B infection or liver cirrhosis). The difference in methylation density between the test subject and the reference group can then be quantified.

[0261] In one embodiment, the reference range can be obtained from the value of the control group. The deviation in the result of the test subject from the upper or lower limit of the reference group can then be used to determine whether the subject has a tumor. This amount is influenced by the concentration fraction of tumor-derived DNA in plasma and the difference in methylation levels between malignant and non-malignant tissues. A higher concentration fraction of tumor-derived DNA in plasma leads to a larger difference in methylation density between the test plasma sample and the control. A larger degree of difference in the methylation levels of malignant and non-malignant tissues is also associated with a larger difference in methylation density between the test plasma sample and the control. In yet another embodiment, different reference groups are selected for test subjects of different age ranges.

[0262] In another embodiment, the mean and SD of the methylation density of four control patients were calculated for each 1 Mb bin. Then, for the corresponding bin, the difference between the methylation density of the HCC patients and the mean of the control patients was calculated. In one embodiment, this difference was then divided by the SD of the corresponding bin to determine the Z-score. In other words, the Z-score represents the difference in methylation density between the test plasma sample and the control plasma sample, expressed as the number of SDs from the mean of the control patients. A Z-score of greater than 3 for a bin indicates that the plasma DNA of the HCC patients is more highly methylated than the control patients in that bin by greater than 3 SDs, while a Z-score of less than -3 for a bin indicates that the plasma DNA of the HCC patients is more hypomethylated than the control patients in that bin by greater than 3 SDs.

[0263] Figures 25A and 25B show the Z-scores of plasma DNA methylation density for pre-surgery (plot 2500) and post-surgery (plot 2550) plasma samples of HCC patients using plasma methylome data of four healthy control patients as reference for chromosome 1. Each dot represents the results of one 1 Mb bin. Black dots represent bins with Z-scores between -3 and 3. Red dots represent bins with Z-scores less than -3.

[0264] Figure 26A is a table 2600 showing the Z-score data for pre-op and post-op plasma. The majority of bins (80.9%) on chromosome 1 in the pre-op plasma sample had a Z-score less than -3, indicating that the pre-op plasma DNA of HCC patients was significantly less methylated than the pre-op plasma DNA of control patients. Conversely, the number of red dots was substantially reduced in the post-op plasma sample (8.3% of bins on chromosome 1), suggesting that the majority of tumor DNA was removed from the circulation due to surgical resection of the source of circulating tumor DNA.

[0265] FIG. 26B is a Circos plot 2620 showing the Z-scores of plasma DNA methylation density for pre- and post-surgery plasma samples from an HCC patient using four healthy control patients as references for 1 Mb bins analyzed from all autosomes. The outermost ring shows symbols for human autosomes. The middle ring shows data for the pre-surgery plasma sample. The innermost ring shows data for the post-surgery plasma sample. Each dot represents the results of one 1 Mb bin. Black dots represent bins with Z-scores between -3 and 3. Red dots represent bins with Z-scores less than -3. Green dots represent bins with Z-scores greater than 3.

[0266] FIG. 26C is a table 2640 showing the distribution of Z-scores of 1 Mb bins of the whole genome in both pre- and post-surgery plasma samples of HCC patients. The results show that the pre-surgery plasma DNA of HCC patients was less methylated in the majority of regions in the whole genome (85.2% of 1 Mb bins) than the pre-surgery plasma DNA of the controls. In contrast, the majority of regions in the post-surgery plasma samples (93.5% of 1 Mb bins) did not show significant hyper- or hypomethylation compared to the controls. These data indicate that the majority of tumor DNA, which is primarily hypomethylated in nature for this HCC, was not present in the post-surgery plasma samples.

[0267] In one embodiment, the number, percentage or proportion of bins with Z-scores less than -3 can be used to indicate whether cancer is present. For example, as shown in table 2640, 2330 (85.2%) of 2734 bins analyzed showed Z-scores less than -3 in pre-op plasma, while only 171 (6.3%) of 2734 bins analyzed showed Z-scores less than -3 in post-op plasma. The data shows that the amount of tumor DNA in pre-op plasma was significantly greater than the amount of tumor DNA in post-op plasma.

[0268] The cutoff value for the number of bins can be determined using statistical methods. For example, approximately 0.15% of the bins are expected to have a Z value less than -3 based on normal distribution. Thus, the cutoff number of bins can be 0.15% of the total number of bins being analyzed. In other words, if the plasma sample obtained from a non-pregnant individual shows more than 0.15% bins with a Z value less than -3, there is a source of hypomethylated DNA in the plasma, i.e., cancer. For example, 0.15% of the 2734 1 Mb bins analyzed in this example is about 4 bins. Using this value as a cutoff, both the pre-surgery and post-surgery plasma samples contained hypomethylated tumor-derived DNA, but the amount was significantly higher in the pre-surgery plasma samples than in the post-surgery plasma samples. In the four healthy control patients, none of the bins showed significant hypermethylation or hypomethylation. Other cutoff values ​​(e.g., 1.1%) can also be used and may vary depending on the requirements of the assay used. As another example, the cutoff percentage may vary depending on the statistical distribution, as well as the desired sensitivity and acceptable specificity.

[0269] In another embodiment, the cutoff value can be determined by receiver operating characteristic (ROC) curve analysis by analyzing several cancer patients and individuals without cancer.To further confirm the specificity of this approach, plasma samples obtained from patients seeking consultation for non-malignant conditions (C06) were analyzed. 1.1% of the bins had a Z value of less than -3.In one embodiment, different thresholds can be used to classify different levels of disease conditions.A lower threshold rate can be used to distinguish healthy conditions from benign conditions, and a higher threshold rate can be used to distinguish benign conditions from malignant tumors.

[0270] The diagnostic performance of plasma hypomethylation analysis using massively parallel sequencing appears to be superior to that obtained using polymerase chain reaction (PCR)-based amplification of certain classes of repetitive regions, such as long interspersed nucleotide sequences-1 (LINE-1) (P Tangkijvanich et al. 2007 Clin Chim Acta; 379:127-133). One possible explanation for this finding is that hypomethylation is widespread in the tumor genome, but with some heterogeneity from one genomic region to another.

[0271] Indeed, it was observed that the mean plasma methylation density of the reference subjects varied across the genome (Figure 56). Each red dot in Figure 56 represents the mean methylation density of one 1 Mb bin among 32 healthy subjects. The plot shows all 1 Mb bins analyzed across the genome. The numbers in each box represent the chromosome number. It was observed that the mean methylation density varied from bin to bin.

[0272] Simple PCR-based assays would not be able to take such region-to-region heterogeneity into account in their diagnostic algorithms. Such heterogeneity broadens the range of methylation densities observed among healthy individuals. A larger reduction in methylation density is required for a sample to be considered as showing hypomethylation. This results in a decrease in the sensitivity of the test.

[0273] In contrast, in massively parallel sequencing-based approaches, the genome is divided into 1 Mb bins (or bins of other sizes) and the methylation density of such bins is measured individually. This approach reduces the effect of variation in baseline methylation density across different genomic regions, because such regions are compared between test samples and controls. Indeed, within the same bin, the inter-individual variation among the 32 healthy controls was relatively small. 95% of the bins had a coefficient of variation (CV) of 1.8% or less among the 32 healthy controls. It should be noted that to further enhance the sensitivity of detection of cancer-associated hypomethylation, comparisons can be made across multiple genomic regions. Sensitivity would be enhanced by testing multiple genomic regions, since it would be protected from the effects of biological variation in cases where cancer samples happen not to show hypomethylation in a particular region when only one region is tested.

[0274] The approach of comparing the methylation density of corresponding genomic regions between control and test samples (e.g., testing each genomic region separately and then combining such results), and performing this comparison for multiple genomic regions, has a higher signal-to-noise ratio in detecting hypomethylation associated with cancer. This massively parallel sequencing approach is shown as an example. Other methodologies that can determine the methylation density of multiple genomic regions and allow the comparison of the methylation density of corresponding regions between control and test samples are expected to achieve similar effects. For example, hybridization probes or molecular inversion probes that can target plasma DNA molecules derived from specific genomic regions and determine the methylation level of the region can be designed to achieve the desired effect.

[0275] In yet another embodiment, the sum of the Z-scores of all bins can be used to determine whether cancer exists or to monitor the continuous change in the level of plasma DNA methylation.Due to the global hypomethylation nature of tumor DNA, the sum of Z-scores will be lower in plasma collected from individuals with cancer than in healthy controls.The sum of Z-scores of pre- and post-surgery plasma samples of HCC patients was -49843.8 and -3132.13, respectively.

[0276] In other embodiments, other methods can be used to examine the methylation level of plasma DNA. For example, the percentage of methylated cytosine residues relative to the total amount of cytosine residues can be determined by using mass spectrometry (ML Chen et al. 2013 Clin Chem; 59: 824-832) or massively parallel sequencing. However, since most cytosine residues are not present in CpG dinucleotide sequences, the percentage of methylated cytosine in all cytosine residues will be relatively small when compared with the methylation level estimated in the sequence of CpG dinucleotides. The methylation levels of tissue and plasma samples obtained from HCC patients and four plasma samples obtained from healthy control groups were determined. Genome-wide massively parallel sequencing data was used to measure the methylation levels in CpG sequences, all cytosines, CHG sequences, and CHH sequences. H refers to adenine, thymine, or cytosine residues.

[0277] Figure 26D is table 2660, showing the methylation levels of tumor tissue and pre-surgery plasma samples overlapping with some of the control plasma samples when using CHH sequence and CHG sequence.The methylation levels of tumor tissue and pre-surgery plasma samples are consistently lower than those of buffy coat, non-tumorous liver tissue, post-surgery plasma samples and healthy control plasma samples in both CpG and unspecified cytosine.However, the data based on methylated CpG, i.e., methylation density, shows a wider dynamic range than the data based on methylated cytosine.

[0278] In other embodiments, the methylation status of plasma DNA can be determined by methods using antibodies against methylated cytosine (e.g., methylated DNA immunoprecipitation (MeDIP)). However, the accuracy of these methods is expected to be less than sequencing-based methods due to variability in antibody binding. In yet another embodiment, the level of 5-hydroxymethylcytosine in plasma DNA can be determined. In this regard, reduced levels of 5-hydroxymethylcytosine have been found to be an epigenetic signature of certain cancers (e.g., melanoma) (CG Lian, et al. 2012 Cell; 150: 1135-1146).

[0279] In addition to HCC, we also investigated whether this approach could be applied to other cancer types. Plasma samples from two patients with lung adenocarcinoma (CL1 and CL2), two patients with nasopharyngeal carcinoma (NPC1 and NPC2), two patients with colorectal cancer (CRC1 and CRC2), one patient with metastatic neuroendocrine tumor (NE1) and one patient with metastatic leiomyosarcoma (SMS1) were analyzed. Plasma DNA from these subjects was converted by bisulfite and sequenced using the Illumina HiSeq2000 platform for 50 bp at one end. The four healthy control patients mentioned above were used as the reference group for the analysis of these eight patients. 50 bp of sequence reads at one end were used. The whole genome was divided into 1 Mb bins. The data obtained from the reference group were used to calculate the mean and SD of the methylation density for each bin. The results of the eight cancer patients were then expressed as Z scores indicating the number of SDs from the mean of the reference group. A positive value indicates that the methylation density of the test example is lower than the mean value of the reference group, and conversely, a negative value indicates that the methylation density of the test example is higher than the mean value of the reference group. The number of sequence reads and the sequencing depth achieved per sample are shown in Table 2780 of Figure 27I.

[0280] 27A-H are Circos plots of methylation density for eight cancer patients, according to an embodiment of the present invention. Each dot represents the results for a 1 Mb bin. Black dots represent bins with Z-scores between -3 and 3. Red dots represent bins with Z-scores less than -3. Green dots represent bins with Z-scores greater than 3. The space between two consecutive lines represents a difference in Z-score of 20.

[0281] Significant hypomethylation was observed in multiple regions throughout the genome in patients with most cancer types, including lung cancer, nasopharyngeal carcinoma, colorectal cancer, and metastatic neuroendocrine tumors. Interestingly, in addition to hypomethylation, significant hypermethylation was observed in multiple regions throughout the genome in cases with metastatic leiomyosarcoma. The embryonic origin of leiomyosarcoma is mesodermal, whereas the embryonic origin of the other cancer types in the remaining seven patients is ectodermal. Thus, the DNA methylation pattern of sarcomas may differ from that of carcinomas.

[0282] As can be seen from this example, methylation patterns of plasma DNA can be useful to distinguish different types of cancer, in this example, carcinoma and sarcoma. These data also suggest that the approach can be used to detect aberrant hypermethylation associated with malignant tumors. In all eight of these cases, only plasma samples were available, and tumor tissue was not analyzed. This shows that tumor-derived DNA can be easily detected in plasma using the described method, even without prior methylation profile or methylation levels of tumor tissue.

[0283] FIG. 27J is a table 2790 showing the Z-score distribution of 1 Mb bins across the genome in plasma of patients with different malignancies. The percentage of bins with Z-scores below -3, between -3 and 3, and above 3 are shown for each example. In all examples, more than 5% of the bins had Z-scores below -3. Thus, if a cutoff of 5% of bins being significantly hypomethylated is used to classify a sample as cancer positive, all of these examples would be classified as cancer positive. The results indicate that hypomethylation may be a common phenomenon in various cancer types, and plasma methylome analysis would be useful for detecting various cancer types.

[0284] D. Method 28 is a flow chart of a method 2800 of analyzing a biological sample of an organism to determine a classification of a level of cancer according to an embodiment of the present invention. The biological sample contains DNA from normal cells and can potentially contain DNA from cells associated with cancer. At least a portion of the DNA in the biological sample can be cell-free DNA.

[0285] At block 2810, a plurality of DNA molecules from a biological sample are analyzed. Analysis of the DNA molecules may include determining the location of the DNA molecule within the genome of the organism and determining whether the DNA molecule is methylated at one or more sites. The analysis may be performed by obtaining sequence reads from methylation-aware sequencing, and therefore the analysis may be performed only on data previously obtained from the DNA. In other embodiments, the analysis may include actual sequencing or other active steps to obtain data.

[0286] At block 2820, for each of the plurality of sites, the respective number of DNA molecules that are methylated at the site is determined. In one embodiment, the sites are CpG sites, and may be only certain CpG sites selected using one or more criteria described herein. The number of DNA molecules that are methylated is equivalent to a determination of the number that are unmethylated if normalization is performed using the total number of DNA molecules analyzed at a particular site (e.g., the total number of sequence reads). For example, an increase in the CpG methylation density of a region is equivalent to a decrease in the density of unmethylated CpGs in the same region.

[0287] At block 2830, a first methylation level is calculated based on the respective number of DNA molecules that are methylated at the plurality of sites. The first methylation level may correspond to a methylation density determined based on the number of DNA molecules that correspond to the plurality of sites. The sites may correspond to multiple loci or a single locus.

[0288] In block 2840, the first methylation level is compared to a first cutoff value. The first cutoff value may be a reference methylation level or may be related to a reference methylation level (e.g., a particular distance from a normal level). The reference methylation level may be determined from a sample of an individual without cancer or from a genetic locus of the organism that is known not to be associated with cancer of the organism. The first cutoff value may be established from a reference methylation level obtained from a prior biological sample of the organism obtained prior to testing of the biological sample.

[0289] In one embodiment, the first cut-off value is a certain distance (e.g., a certain value of standard deviation) from a reference methylation level established from a biological sample obtained from a healthy organism. The comparison can be performed by determining the difference between the first methylation level and the reference methylation level, and then comparing the difference with a threshold value corresponding to the first cut-off value (e.g., to determine whether the methylation level is statistically different from the reference methylation level).

[0290] At block 2850, a classification of the level of cancer is determined based on the comparison. Examples of levels of cancer include whether the subject has cancer or a precancerous condition, or whether the subject has an increased likelihood of developing cancer. In one embodiment, the first cutoff value may be determined from a sample previously obtained from the subject (e.g., a reference methylation level may be determined from a prior sample).

[0291] In some embodiments, the first methylation level may correspond to the number of regions whose methylation level exceeds a threshold. For example, a plurality of regions of the genome of an organism may be identified. The regions may be identified using criteria described herein (e.g., having a certain length or a certain number of sites). One or more sites (e.g., CpG sites) may be identified within each region. A region methylation level may be calculated for each region. The first methylation level is for the first region. Each region methylation level is compared to a respective region cutoff value, which may be the same or different between regions. The region cutoff value for the first region is the first cutoff value. Each region cutoff value is a certain amount (e.g., 0.5) from a reference methylation level, so that only regions with significant differences from a reference, which may be determined from non-cancer subjects, may be counted.

[0292] A first number of regions whose region methylation levels exceed the respective region cutoff value can be determined and compared to a threshold value to determine a classification. In some implementations, the threshold value is a percentage. Comparing the first number to the threshold value can include dividing the first number of regions by the second number of regions (e.g., all regions) and then comparing to the threshold value, e.g., as part of a normalization process.

[0293] As mentioned above, the concentration fraction of tumor DNA in biological samples can be used to calculate a first cut-off value. The concentration fraction can be simply estimated to be greater than a minimum value, while samples with a concentration fraction lower than the minimum value can be flagged, for example, as not suitable for analysis. The minimum value can be determined based on the expected difference in tumor methylation level compared to the reference methylation level. For example, if the difference is 0.5 (for example, used as cut-off value), the concentration of a certain tumor needs to be high enough to see this difference.

[0294] Certain techniques from method 1300 may be applied to method 2800. In method 1300, copy number variations may be determined in a tumor (e.g., a first chromosome region of a tumor may be examined to see if it has copy number changes compared to a second chromosome region of the tumor). Thus, the presence of a tumor may be inferred by method 1300. In method 2800, a sample may be examined to see if there is any indication of the presence of a tumor, regardless of any copy number characteristics. Some techniques of the two methods may be similar. However, the cutoff values ​​and methylation parameters (e.g., normalized methylation levels) of method 2800 may detect statistically significant differences from a reference methylation level of non-cancerous DNA, as opposed to differences from a reference methylation level of a mixture of cancer DNA and non-cancerous DNA with some regions that may have copy number changes. Thus, the reference value for method 2800 may be determined from a sample that does not contain cancer, for example, from an organism that does not have cancer, or from non-cancerous tissue of the same patient (e.g., from a contemporaneous sample known to be cancer-free, which may be determined from previously collected plasma, or cellular DNA).

[0295] E. Prediction of the Minimum Fractional Concentration of Tumor DNA Detected Using Plasma DNA Methylation Analysis One way to measure the sensitivity of approaches to detect cancer using plasma DNA methylation levels is related to the minimum fractional tumor-derived DNA concentration required to reveal changes in plasma DNA methylation levels compared to control plasma DNA methylation levels. Test sensitivity also depends on the degree of difference in DNA methylation between tumor tissue and baseline plasma DNA methylation levels in healthy controls or blood cell DNA. Blood cells are the main source of DNA in plasma of healthy individuals. The greater the difference, the more easily cancer patients can be distinguished from non-cancerous individuals, which will be reflected as a lower detection limit of tumor-derived material in plasma and a higher clinical sensitivity in detecting cancer patients. In addition, the variation in plasma DNA methylation in healthy subjects or subjects of different ages (G Hannum et al. 2013 Mol Cell; 49: 359-367) will also affect the sensitivity of detecting methylation changes associated with the presence of cancer. The smaller the variation in plasma DNA methylation in healthy subjects, the easier it will be to detect changes caused by the presence of small amounts of cancer-derived DNA.

[0296] Figure 29A is a plot 2900 showing the distribution of methylation density in a reference subject, assuming that the distribution follows a normal distribution, and the analysis is based on each plasma sample giving only one methylation density value (e.g., the methylation density of all autosomes or a specific chromosome). This explains how the specificity of the analysis is affected. In one embodiment, a cutoff of 3 SD below the mean DNA methylation density of the reference subject is used to determine whether the test sample is significantly lower in methylation status than the sample obtained from the reference subject. When this cutoff is used, it is expected that approximately 0.15% of non-cancerous subjects will have a false positive result classified as having cancer, resulting in a specificity of 99.85%.

[0297] FIG. 29B is a plot 2950 showing the distribution of methylation density in the reference subjects and the cancer patients. The cutoff value is 3 SD below the mean value of the methylation density of the reference subjects. If the mean value of the methylation density of the cancer patients is 2 SD below the cutoff value (i.e., 5 SD below the mean value of the reference subjects), 97.5% of the cancer subjects are expected to have a methylation density below the cutoff value. In other words, if one methylation density value is given to each subject, for example, if the global methylation density of the whole genome, all autosomes, or a specific chromosome is analyzed, the expected sensitivity will be 97.5%. The difference between the mean methylation densities of the two populations is influenced by two factors, namely the degree of difference in methylation levels between cancerous and non-cancerous tissues, and the concentration fraction of tumor-derived DNA in the plasma sample. The higher the values ​​of these two parameters, the greater the difference between the methylation density values ​​of the two populations will be. Furthermore, the lower the SD of the methylation density distribution of the two populations, the less the methylation density distributions of the two populations will overlap.

[0298] In the following, we will explain this concept using a hypothetical example. Assume that the methylation density of tumor tissue is approximately 0.45, and that of plasma DNA of healthy subjects is approximately 0.7. These assumed values ​​are similar to those obtained from HCC patients, where the overall methylation density of autosomes was 42.9%, and the average methylation density of autosomes in plasma samples obtained from healthy controls was 71.6%. Assuming a CV of 1% for measuring plasma DNA methylation density across the genome, the cutoff value is 0.7×(100%−3×1%)=0.679. To achieve a sensitivity of 97.5%, the average methylation density of plasma DNA of cancer patients needs to be approximately 0.679−0.7×(2×1%)=0.665. Let f represent the concentration fraction of tumor-derived DNA in the plasma sample. f can be calculated as (0.7−0.45)×f=0.7−0.665. From this, f is approximately 14%. This calculation estimates that the minimum detectable concentration fraction in plasma is 14%, resulting in a diagnostic sensitivity of 97.5% being achieved when global methylation density across the genome is used as the diagnostic parameter.

[0299] This analysis was then performed on data obtained from HCC patients. To account for this, only one methylation density measurement based on values ​​estimated from all autosomes was performed on each sample. The mean methylation density among plasma samples obtained from healthy subjects was 71.6%. The SD of the methylation densities of these four samples was 0.631%. Therefore, the cutoff value of plasma methylation density needs to be 71.6%-3×0.631%=69.7% to achieve a Z score of less than -3 and a specificity of 99.85%. To achieve a sensitivity of 97.5%, the mean plasma methylation density of cancer patients needs to be 2 SD below the cutoff (i.e., 68.4%). The methylation density of tumor tissue was 42.9%, so using the formula: P=BKG×(1-f)+TUM×f, f needs to be at least 11.1%.

[0300] In another embodiment, the methylation density of different genomic regions can be analyzed separately, for example, as shown in FIG. 25A or FIG. 26B. In other words, multiple methylation level measurements are performed for each sample. As shown below, significant hypomethylation can be detected at even lower tumor DNA concentration fractions in plasma, so the diagnostic ability of plasma DNA methylation analysis in cancer detection is enhanced. The number of genomic regions that show significant deviations in methylation density from a reference population can be counted. The number of genomic regions can then be compared to a cutoff value to determine whether there is an overall significant hypomethylation of plasma DNA across the population of examined genomic regions (e.g., 1 Mb bins of the entire genome). The cutoff value can be established by the analysis of a group of reference subjects without cancer, or obtained mathematically, for example, according to a normal distribution function.

[0301] FIG. 30 is a plot 3000 showing the distribution of methylation density of plasma DNA of healthy subjects and cancer patients. The methylation density of each 1 Mb bin is compared to the corresponding value of the reference group. The percentage of bins showing significant hypomethylation (3 SD below the mean value of the reference group) was determined. A cutoff of 10% significant hypomethylation status was used to determine whether tumor-derived DNA was present in the plasma sample. Other cutoff values ​​such as 5%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80% or 90% can also be used depending on the desired sensitivity and specificity of the test.

[0302] For example, 10% of 1Mb bins showing significant hypomethylation (Z value less than -3) can be used as a cutoff to classify a sample as containing tumor-derived DNA. If there are more than 10% of bins that are significantly more hypomethylated than the reference group, the sample is classified as positive in the cancer test. In each 1Mb bin, a cutoff of 3 SD below the mean methylation density of the reference group is used to define the sample as significantly more hypomethylated. For each 1Mb bin, if the mean plasma DNA methylation density of cancer patients is 1.72 SD lower than the mean plasma DNA methylation density of the reference subject, there is a 10% probability that the methylation density value of any particular bin in the cancer patient will be lower than the cutoff (i.e., Z value less than -3) and give a positive result. Then, if all 1Mb bins in the entire genome are examined, it is expected that approximately 10% of the bins will show a positive result with significantly lower methylation density (i.e., Z value less than -3). The global methylation density of plasma DNA in healthy subjects is approximately 0.7, and assuming a coefficient of variation (CV) of 1% for plasma DNA methylation density measurements for each 1 Mb bin, the average methylation density of plasma DNA in cancer patients would need to be 0.7 x (100% - 1.72 x 1%) = 0.68796. This average plasma DNA methylation density is obtained by taking f as the fractional concentration of tumor-derived DNA in plasma. Assuming the methylation density of tumor tissue is 0.45, f can be calculated using the following formula:

[0303]

number

[0304] During the ceremony,

[0305]

number

[0306] represents the average methylation density of plasma DNA in the reference individuals;

[0307]

number

[0308] represents the methylation density of tumor tissues from cancer patients;

[0309]

number

[0310] represents the average methylation density of plasma DNA in cancer patients.

[0311] Using this formula, (0.7-0.45) x f = 0.7-0.68796. Thus, the minimum concentration fraction that can be detected using this approach is estimated to be 4.8%. Sensitivity can be further enhanced by decreasing the cutoff rate of bins that are significantly more hypomethylated, for example from 10% to 5%.

[0312] As shown in the above example, the sensitivity of this method is determined by the degree of difference in methylation levels between cancerous and non-cancerous tissues (e.g., blood cells). In one embodiment, only chromosomal regions that show large differences in methylation density between plasma DNA and tumor tissue of non-cancerous subjects are selected. In one embodiment, only regions with differences in methylation density greater than 0.5 are selected. In other embodiments, differences of 0.4, 0.6, 0.7, 0.8 or 0.9 can be used to select suitable regions. In yet another embodiment, the physical size of the genomic region is not fixed. Instead, the genomic region is defined, for example, based on a fixed read depth or a fixed number of CpG sites. The methylation level in multiple of these genomic regions is evaluated in each sample.

[0313] 31 is a graph 3100 showing the distribution of the difference in methylation density between the mean values ​​of plasma DNA of healthy subjects and tumor tissues of HCC patients. Positive values ​​indicate that the methylation density is higher in plasma DNA of healthy subjects, and negative values ​​indicate that the methylation density is higher in tumor tissues.

[0314] In one embodiment, the bins with the largest difference between the methylation densities of cancer and non-cancerous tissues (e.g., bins with a difference of more than 0.5) can be selected, regardless of whether the tumor is hypomethylated or hypermethylated in these bins. The detection limit of the concentration fraction of tumor-derived DNA in plasma can be lowered by focusing on these bins, where the difference between the distribution of plasma DNA methylation levels between cancer and non-cancer subjects is larger, and the same concentration fraction of tumor-derived DNA in plasma can be given. For example, if only bins with a difference of more than 0.5 are used and a cutoff of 10% of the bins being significantly more hypomethylated is adopted to determine whether a test individual has cancer, the minimum concentration fraction of tumor-derived DNA to be detected (f) can be calculated using the following formula:

[0315]

number

[0316] During the ceremony,

[0317]

number

[0318] represents the average methylation density of plasma DNA in the reference individuals;

[0319]

number

[0320] represents the methylation density of tumor tissue in cancer patients;

[0321]

number

[0322] represents the mean methylation density of plasma DNA in cancer patients.

[0323] Meanwhile, the difference in methylation density between the plasma and tumor tissue of the reference subject is at least 0.5. Then, 0.5×f=0.7-0.68796, and f=2.4%. Therefore, by focusing on bins with larger differences in methylation density between cancer tissues and non-cancerous tissues, the lower limit of tumor-derived DNA fraction can be lowered from 4.8% to 2.4%. Information regarding which bins show a larger degree of methylation difference between cancer tissues and non-cancerous tissues (e.g., blood cells) can be determined from tumor tissues of the same organ or the same tissue type obtained from other individuals.

[0324] In another embodiment, the parameter is obtained from the methylation density of plasma DNA of all bins, and the difference in methylation density between cancerous and non-cancerous tissues can be taken into account. Bins with larger differences can be given heavier weights. In one embodiment, the difference in methylation density between cancerous and non-cancerous tissues of each bin can be directly used as the weight for that particular bin in calculating the final parameter.

[0325] In yet another embodiment, different cancer types may have different patterns of methylation in tumor tissue, and a cancer-specific weight profile may be obtained from the degree of methylation of a particular cancer type.

[0326] In yet another embodiment, the association between bins of methylation density can be determined in subjects with and without cancer. In FIG. 8, it can be observed that in a small number of bins, tumor tissue was more methylated than the plasma DNA of the reference subject. Therefore, the bins with the most extreme difference values ​​(e.g., difference greater than 0.5 and difference less than 0) can be selected. The ratio of the methylation densities of these bins can then be used to indicate whether the test individual has cancer. In other embodiments, the difference and quotient of the methylation densities of different bins can be used as parameters to indicate the association between bins.

[0327] We further evaluated the detection sensitivity of our approach for detecting or assessing tumors using the methylation density of multiple genomic regions, as shown by data obtained from HCC patients. First, reads from pre-surgery plasma were mixed with reads from plasma samples of healthy controls to mimic plasma samples containing tumor DNA concentration fractions ranging from 20% to 0.5%. Then, we counted the percentage of 1 Mb bins (out of 2,734 bins in the whole genome) with methylation density equal to a Z-score of less than -3. When the tumor DNA concentration fraction in plasma was 20%, 80.0% of the bins showed significant hypomethylation. Data corresponding to tumor DNA concentration fractions in plasma of 10%, 5%, 2%, 1% and 0.5% were bins showing hypomethylation of 67.6%, 49.7%, 18.9%, 3.8% and 0.77%, respectively. Because the theoretical limit for the number of bins showing a Z value less than −3 in control samples is 0.15%, our data show that even when the tumor concentration fraction was only 0.5%, there were many more bins (0.77%) that exceeded the theoretical cutoff limit.

[0328] Figure 32A is a table 3200 showing the effect of reducing sequencing depth when plasma samples contained 5% or 2% tumor DNA. A high percentage of bins (>0.15%) showing significant hypomethylation could be detected even when the average sequencing depth was only 0.022 times per haploid genome.

[0329] Figure 32B is a graph 3250 showing the methylation density of repeat and non-repeat regions in plasma of four healthy control patients, buffy coat of HCC patients, normal liver tissue, tumor tissue, pre-surgery plasma sample and post-surgery plasma sample. It can be observed that repeat regions were more methylated (higher methylation density) than non-repeat regions in both cancer and non-cancerous tissues. However, the difference in methylation between repeat and non-repeat regions was larger in non-cancerous tissue and plasma DNA of healthy subjects when compared with tumor tissue.

[0330] As a result, plasma DNA of cancer patients showed a greater decrease in methylation density in repeat regions than in non-repeat regions. The difference in plasma DNA methylation density between the mean value of four healthy controls and HCC patients was 0.163 and 0.088 in repeat and non-repeat regions, respectively. Data on pre-surgery plasma samples and post-surgery plasma samples also showed that the dynamic range of changes in methylation density was greater in repeat regions than in non-repeat regions. In one embodiment, plasma DNA methylation density in repeat regions can be used to determine whether a patient is affected by cancer or to monitor disease progression.

[0331] As mentioned above, the variation of methylation density in the plasma of the reference subject also affects the accuracy of distinguishing cancer patients from non-cancer individuals. The closer the distribution of methylation density (i.e., the smaller the standard deviation), the more accurate the distinction between cancer and non-cancer subjects. In another embodiment, the coefficient of variation (CV) of methylation density of 1 Mb bins can be used as a criterion for selecting bins with low variability of plasma DNA methylation density in the reference group. For example, only bins with CV less than 1% are selected. Other values, such as 0.5%, 0.75%, 1.25% and 1.5%, can also be used as a criterion for selecting bins with low variability of methylation density. In yet another embodiment, the selection criterion can include both the CV of the bins and the difference in methylation density between cancer tissues and non-cancerous tissues.

[0332] Methylation density can also be used to estimate the fractional concentration of tumor-derived DNA in a plasma sample when the methylation density of the tumor tissue is known. This information can be obtained by analysis of the patient's tumor or from a study of tumors from several patients with the same cancer type. As mentioned above, plasma methylation density (P) can be expressed using the following formula:

[0333]

number

[0334] where BKG is the background methylation density obtained from blood cells and other organs, TUM is the methylation density in tumor tissues, and f is the fractional concentration of tumor-derived DNA in the plasma sample. This can be rewritten as follows:

[0335]

number

[0336] The value of BKG can be determined by analyzing a patient's plasma sample at a time when cancer is not present, or from a study of a reference group of cancer-free individuals. Thus, f can be determined after measuring the plasma methylation density.

[0337] F. Combination with other methods The methylation analysis approach described herein can be used in conjunction with other methods based on genetic alterations in tumor-derived DNA in plasma. Examples of such methods include analysis of cancer-associated chromosomal abnormalities (KCA Chan et al. 2013 Clin Chem; 59:211-224; RJ Leary et al. 2012 Sci Transl Med; 4:162ra154) and cancer-associated single nucleotide mutations in plasma (KCA Chan et al. 2013 Clin Chem; 59:211-224). Methylation analysis approaches have advantages over these genetic approaches.

[0338] As shown in Figure 21A, hypomethylation of tumor DNA is a global phenomenon involving regions distributed almost throughout the genome. Thus, DNA fragments from all chromosomal regions are informative regarding the potential contribution of tumor-derived hypomethylated DNA to plasma / serum DNA in patients. In contrast, chromosomal abnormalities (amplification or deletion of chromosomal regions) are only present in some chromosomal regions, and DNA fragments from regions without chromosomal abnormalities in tumor tissue are not informative in the analysis (KCA Chan et al. 2013 Clin Chem; 59: 211-224). Similarly, only a few thousand single nucleotide changes are observed in each cancer genome (KCA Chan et al. 2013 Clin Chem; 59: 211-224). DNA fragments that do not overlap with these single nucleotide changes are not informative in determining whether tumor-derived DNA is present in plasma. Thus, the present methylation analysis approach is potentially more cost-effective than those genetic approaches that detect cancer-associated changes in circulating blood.

[0339] In one embodiment, the cost-effectiveness of plasma DNA methylation analysis can be further enhanced by enriching DNA fragments from the most informative regions, e.g., regions with the highest differential methylation differences between cancerous and non-cancerous tissues. Exemplary methods for enriching these regions include the use of hybridization probes (e.g., Nimblegen SeqCap system and Agilent SureSelect Target Enrichment system), PCR amplification, and solid-phase hybridization.

[0340] G. Tissue-specific analysis / donor Tumor-derived cells invade and metastasize to adjacent or distant organs. Invaded tissues or metastatic foci donate DNA into the plasma as a result of cell death. By analyzing the methylation profile of DNA in the plasma of cancer patients and detecting the presence of tissue-specific methylation signatures, the tissue types involved in the disease process can be detected. This approach provides a non-invasive anatomical scan of the tissues involved in the cancer process, helping to identify the organs involved as primary and metastatic sites. Also, by monitoring the relative concentration of the methylation signatures of diseased organs in the plasma, it is possible to assess the tumor burden of those organs and determine whether the cancer process in that organ is worsening or improving or has been cured. For example, gene X is specifically methylated in the liver. Metastatic involvement of that liver by a cancer (e.g., colorectal cancer) would then be expected to increase the concentration of the methylated sequence from gene X in the plasma. There would also be another sequence or groups of sequences with similar methylation characteristics to gene X. In that case, the results obtained from such sequences could be combined. Similar considerations are applicable to other tissues, such as brain, bone, lung and kidney.

[0341] On the other hand, it is known that DNA from different organs shows tissue-specific methylation signatures (BW Futscher et al. 2002 Nat Genet; 31:175-179; SSC Chim et al. 2008 Clin Chem; 54: 500-511). Therefore, methylation profiling in plasma can be used to elucidate the contribution of tissues from different organs to plasma. Since plasma DNA is thought to be released when cells die, elucidation of such contributions can be used to evaluate organ injury. For example, liver pathology such as hepatitis (e.g. due to a virus, autoimmune process, etc.) or hepatoxicity (e.g. due to a drug overdose (e.g. due to paracetamol) or drug-induced toxins (e.g. alcohol) is associated with hepatocellular injury and would be expected to be associated with increased levels of liver-derived DNA in plasma. For example, suppose gene X is specifically methylated in the liver. Then liver pathology would be expected to increase the concentration of gene X-derived methylated sequences in plasma. Conversely, suppose gene Y is specifically hypomethylated in the liver. Then liver pathology would be expected to decrease the concentration of gene Y-derived methylated sequences in plasma. In yet other embodiments, gene X or gene Y may be replaced by any genomic sequence, not necessarily a gene, that shows differential methylation in different tissues within the body.

[0342] The techniques described herein can also be applied to assess donor-derived DNA in the plasma of organ transplant recipients (YMD Lo et al. 1998 Lancet; 351:1329-1330). Polymorphic differences between donors and recipients were used to distinguish donor-derived DNA from recipient-derived DNA in plasma (YW Zheng et al. 2012 Clin Chem; 58: 549-558). We report that tissue-specific methylation signatures of transplanted organs can also be used as a method to detect donor DNA in recipient plasma.

[0343] By monitoring the concentration of donor DNA, the condition of transplanted organs can be evaluated non-invasively. For example, transplant rejection is associated with higher cell death rates and thus higher concentrations of donor DNA in recipient plasma (or serum), which is reflected by the methylation signature of transplanted organs, and may be increased when compared to when the patient is in stable condition, or when compared to other stable transplant recipients or non-transplanted healthy controls. Similar to what has been reported for cancer, donor-derived DNA can be identified in transplant recipient plasma by detecting all or some of the following characteristics, including polymorphic differences, shorter DNA sizes of transplanted solid organs (YW Zheng et al. 2012 Clin Chem; 58: 549-558), and tissue-specific methylation characteristics.

[0344] H. Size-based methylation normalization As described above and in Lun et al (FMF Lun et al. Clin. Chem. 2013; doi:10.1373 / clinchem.2013.212274), methylation density (e.g., plasma DNA methylation density) correlates with DNA fragment size. The distribution of methylation density of shorter plasma DNA fragments was significantly lower than that of longer fragments. We propose that some non-cancerous conditions with abnormal fragmentation patterns of plasma DNA (e.g., systemic lupus erythematosus (SLE)) may show apparent plasma DNA hypomethylation due to the presence of a larger amount of short plasma DNA fragments that are in a hypomethylated state. In other words, the size distribution of plasma DNA may be a confounding factor of plasma DNA methylation density.

[0345] Figure 34A shows the size distribution of plasma DNA in SLE patient SLE04. The size distribution of nine healthy control patients is shown as a gray dotted line, and the size distribution of SLE04 is shown as a black solid line. Short plasma DNA fragments were more abundant in SLE04 than in nine healthy control patients. This size distribution pattern may confound methylation analysis of plasma DNA, leading to further apparent hypomethylation, since shorter DNA fragments are generally more hypomethylated.

[0346] In some embodiments, the measured methylation level can be normalized to reduce the confounding effect of size distribution on plasma DNA methylation analysis. For example, the size of DNA molecules at multiple sites can be measured. In various implementations, the measurement can obtain a specific size (e.g., length) of DNA molecules, or simply determine that the size is within a certain range, which can also correspond to the size. The normalized methylation level can then be compared to a cutoff value. There are several ways to perform normalization to reduce the confounding effect of size distribution on plasma DNA methylation analysis.

[0347] In one embodiment, size fractionation of DNA (e.g., plasma DNA) may be performed. Size fractionation may ensure that DNA fragments having similar sizes are used to determine methylation levels in a manner consistent with the cutoff value. As part of the size fractionation, DNA fragments having a first size (e.g., a first length range) may be selected, where the first cutoff value corresponds to the first size. Normalization may be achieved by calculating methylation levels using only those selected DNA fragments.

[0348] Size fractionation can be achieved in various ways, for example, by physical separation of DNA molecules of different sizes (e.g., by electrophoresis or microfluidics-based or centrifugation-based technologies), or by in silico analysis. In in silico analysis, in one embodiment, paired-end massively parallel sequencing of plasma DNA molecules can be performed. The size of the sequenced molecules can then be estimated by comparing the positions of each of the two ends of the plasma DNA molecules with a reference human genome. Subsequent analysis can then be performed by selecting sequenced DNA molecules that fit one or more size selection criteria (e.g., size criteria that are within a certain range). Thus, in one embodiment, the methylation density of fragments with smaller sizes (e.g., within a certain range) can be analyzed. The cutoff value (e.g., in block 2840 of method 2800) can be determined based on fragments within the same size range. For example, methylation levels can be determined from samples known to have or not have cancer, and the cutoff value can be determined from these methylation levels.

[0349] In another embodiment, a functional relationship between methylation density and size of circulating DNA can be determined. The functional relationship can be defined by the data points or coefficients of the function. The functional relationship can provide a corresponding scaling value for each size (e.g., shorter sizes can have a corresponding increase in methylation). In various implementations, the scaling value can be between 0 and 1 or greater than 1.

[0350] Normalization can be performed based on average size.For example, the average size corresponding to the DNA molecule used to calculate the first methylation level can be calculated by computer, and the first methylation level can be multiplied by the corresponding scaling value (i.e., corresponding to the average size).As another example, the methylation density of each DNA molecule can be normalized according to the size of the DNA molecule and the relationship between DNA size and methylation.

[0351] In another implementation, normalization can be performed on a molecule-by-molecule basis. For example, the size of each of the DNA molecules at a particular site can be obtained (e.g., as described above), and a scaling value corresponding to each size can be determined from the functional relationship. In a non-normalized calculation, each molecule is equally counted in determining the methylation index at a site. In a normalized calculation, the contribution of a molecule to the methylation index can be weighted by a scaling factor corresponding to the size of the molecule.

[0352] Figures 34B and 34C show methylation analysis of plasma DNA obtained from SLE patient SLE04 (Figure 34B) and HCC patient TBR36 (Figure 34C). The outer ring shows the Z of plasma DNA without in silico size fractionation. meth The inner ring represents the Z region of plasma DNA of 130 bp or more. meth Results are shown. In SLE patient SLE04, 84% of bins showed hypomethylation without in silico size fractionation. The percentage of bins showing hypomethylation was reduced to 15% when only fragments ≥ 130 bp were analyzed. In HCC patient TBR36, 98.5% and 98.6% of bins showed plasma DNA hypomethylation with and without in silico size fractionation, respectively. These results suggest that in silico size fractionation can effectively reduce false-positive hypomethylation results associated with increased plasma DNA fragmentation (e.g., in patients with SLE or other inflammatory diseases).

[0353] In one embodiment, the results of the analysis with and without size fractionation can be compared to indicate whether there is a confounding effect of size on methylation results. Thus, in addition to or instead of normalization, calculation of methylation levels at specific sizes can be used to determine whether there is a possibility of a false positive when the proportion of bins above a cutoff value differs with and without size fractionation, or whether only a specific methylation level differs. For example, the presence of a significant difference between the results of samples with and without size fractionation can be used to indicate the possibility of a false positive result due to an abnormal fragmentation pattern. The threshold for determining whether a difference is significant can be established by analysis of a cohort of cancer patients and a cohort of non-cancer control patients.

[0354] I. Genome-wide analysis of CpG island hypermethylation in plasma In addition to global hypomethylation, CpG island hypermethylation is commonly observed in cancer (SSB Baylin et al. 2011 Nat Rev Cancer; 11: 726-734; PA Jones et al. 2007, Cell; 128: 683-692; M Esteller et al. 2007 Nat Rev Genet 2007; 8: 286-298; M Ehrlich et al. 2002 Oncogene 2002; 21: 5400-5413). In this section, we describe the use of genome-wide analysis of CpG island hypermethylation to detect and monitor cancer.

[0355] 35 is a flow chart of a method 3500 for determining a classification of a cancer level based on hypermethylation of CpG islands according to an embodiment of the present invention. The multiple sites of method 2800 may include CpG sites, where the CpG sites are organized into multiple CpG islands, each CpG island containing one or more CpG sites. The methylation level of each CpG island may be used to determine a classification of a cancer level.

[0356] In block 3510, the CG islands to be analyzed are identified. In this analysis, as an example, a set of CG islands to be analyzed is defined, characterized by a relatively low methylation density in the plasma of healthy reference subjects. In one aspect, the change in methylation density in the reference group can be relatively small to more easily allow detection of cancer-associated hypermethylation. In one embodiment, the CG islands have a mean methylation density less than a first percentage in the reference group, and the coefficient of variation of the methylation density in the reference group is less than a second percentage.

[0357] As an example, and for illustrative purposes, the following criteria may be used to identify useful CG islands: i. The mean methylation density of C-chain-red islands in a reference group (e.g., healthy subjects) is less than 5% ii. The coefficient of variation of the analysis of methylation density in plasma of a reference group (e.g., healthy subjects) is less than 30%. These parameters can be adjusted for the specific application. From our dataset, 454 CG islands in the genome met these criteria.

[0358] The methylation density of each CG island is calculated at block 3520. The methylation density may be calculated as described herein.

[0359] At block 3530, it is determined whether each CpG island is hypermethylated. For example, in an analysis of CpG island hypermethylation of a test case, the methylation density of each CpG island is compared to the corresponding data of a reference group. The methylation density (an example of a methylation level) may be compared to one or more cutoff values ​​to determine whether a particular island is hypermethylated.

[0360] In one embodiment, the first cutoff value may correspond to the mean value of the methylation density of the reference group plus a particular percentage. Another cutoff value may correspond to the mean value of the methylation density of the reference group plus a particular standard deviation value. In some implementations, the Z score (Z meth) was calculated and compared to the cutoff value. As an example, a CpG island in a test subject (e.g., a subject being screened for cancer) was considered to be significantly hypermethylated if it met the following criteria: i. its methylation density is more than 2% higher than the mean of the reference group; ii. Z meth is greater than 3. Additionally, these parameters may be adjusted for a particular application.

[0361] In block 3540, the methylation density (e.g., Z-score) of the hypermethylated CG islands is used to determine a cumulative score. For example, after identifying all significantly hypermethylated CG islands, a score can be calculated that includes the sum of the Z-scores or a function of the Z-scores of all hypermethylated CG islands. One example of a score is the cumulative probability (CP) score, which is described in another section. A cumulative probability score is a Z-score that is calculated according to a probability distribution (e.g., a Student's t probability distribution with 3 degrees of freedom) to determine the probability of having such an observation by chance. meth is used.

[0362] At block 3550, the cumulative score is compared to a cumulative threshold to determine a classification of the level of cancer. For example, if the total hypermethylation in the identified CG islands is sufficiently great, the organism may be identified as having cancer. In one embodiment, the cumulative threshold corresponds to the highest cumulative score obtained from a reference group.

[0363] IX. Methylation and CNAs As mentioned above, the methylation analysis approach described herein can be used in conjunction with other methods based on genetic alterations in tumor-derived DNA in plasma. Examples of such methods include analysis of cancer-associated chromosomal abnormalities (KCA Chan et al. 2013 Clin Chem; 59: 211-224; RJ Leary et al. 2012 Sci Transl Med; 4: 162ra154). Aspects of copy number abnormalities (CNAs) are described in US Patent Application Publication No. 13 / 308,473.

[0364] A. CNA Copy number abnormality can be detected by counting the DNA fragments that align to a specific part of genome, normalizing the counts, and comparing the counts to a cutoff value.In various embodiments, normalization can be performed by counting the DNA fragments that align to different haplotypes in the same part of genome (relative haplotype dosage (RHDO)) or by counting the DNA fragments that align to different parts of genome.

[0365] The RHDO method relies on the use of heterozygous loci. The embodiments described in this section can also be used for homozygous loci by comparing two regions and not two haplotypes of the same region, and are therefore non-haplotype specific. In the relative chromosomal region dosage method, the number of fragments from one chromosomal region (e.g., determined by counting sequence reads aligned to the region) is compared to an expected value (which can be obtained from a reference chromosomal region or from the same region of another sample known to be healthy). In this way, fragments of a chromosomal region are counted regardless of which haplotype the sequenced tag is from. Thus, sequence reads containing non-heterozygous loci can also be used. To make the comparison, in an embodiment, the tag counts can be normalized before the comparison. Each region is defined by at least two loci (separated from each other), and the fragments at these loci can be used to obtain a collective value for the region.

[0366] The normalized number of sequenced reads (tags) of a specific region can be calculated by dividing the number of sequenced reads that are aligned to that region by the total number of sequenced reads that can be aligned to the whole genome. This normalized tag number allows the results obtained from one sample to be compared with the results of another sample. For example, the normalized number can be the ratio (e.g., percentage or fraction) of sequenced reads that are expected to come from a specific region, as described above. In other embodiments, other methods for normalization are possible. For example, the number of counts of one region can be normalized by dividing the number of counts of a reference region (in the above example, the reference region is simply the whole genome). This normalized tag number can then be compared to a threshold value that can be determined from one or more reference samples that do not show cancer.

[0367] The normalized tag count of the test case is then compared to the normalized tag count of one or more reference subjects (e.g., one or more reference subjects without cancer). In one embodiment, the comparison is performed by calculating the Z-score of the case in a particular chromosomal region. The Z-score can be calculated using the following formula: Z-score=(normalized tag count of case-mean) / SD, where "mean" is the mean of the normalized tag counts aligned to a particular chromosomal region of the reference samples; SD is the standard deviation of the normalized tag counts aligned to a particular region of the reference samples. Thus, the Z-score is the number of standard deviations that the normalized tag count of a chromosomal region of the test case deviates from the normalized average tag count of the same chromosomal region of one or more reference subjects.

[0368] In the situation where the test organism has cancer, the chromosomal regions that are amplified in the tumor tissue are overrepresented in the plasma DNA. This results in a positive Z value. On the other hand, the chromosomal regions that are deleted in the tumor tissue are underrepresented in the plasma DNA. This results in a negative Z value. The magnitude of the Z value is determined by several factors.

[0369] One factor is the concentration fraction of tumor-derived DNA in biological samples (e.g., plasma). The higher the concentration fraction of tumor-derived DNA in a sample (e.g., plasma), the greater the difference between the normalized tag numbers of test cases and reference cases. Thus, a larger magnitude of Z-score is obtained.

[0370] Another factor is the variation of normalized tag number in one or more reference cases.If the over-representation of chromosomal regions in the biological samples (e.g., plasma) of test cases is the same, the smaller the variation (i.e., the smaller the standard deviation) of normalized tag number in the reference group, the larger the Z value.Similarly, if the under-representation of chromosomal regions in the biological samples (e.g., plasma) of test cases is the same, the smaller the standard deviation of normalized tag number in the reference group, the more negative the Z value.

[0371] Another factor is the magnitude of chromosomal abnormality in tumor tissue. The magnitude of chromosomal abnormality refers to the copy number change (increase or decrease) of a particular chromosomal region. The higher the copy number change in tumor tissue, the higher the degree of over- or under-representation of a particular chromosomal region in plasma DNA. For example, the loss of both copies of a chromosome leads to a greater under-representation of a chromosomal region in plasma DNA than the loss of one of the two copies of a chromosome, thus resulting in a more negative Z-score. Typically, there are multiple chromosomal abnormalities in cancer. The chromosomal abnormality in each cancer can further vary by its nature (i.e., amplification or deletion), its degree (increase or decrease of single or multiple copies) and its scope (the size of the abnormality in relation to the length of the chromosome).

[0372] The accuracy of measuring normalized tag number is affected by the number of molecules analyzed.To detect chromosomal abnormalities with one copy change (gain or loss), if the concentration fraction is approximately 12.5%, 6.3% and 3.2%, respectively, it is expected that 15,000, 60,000 and 240,000 molecules need to be analyzed.Further details of tag counting in detecting cancer in different chromosomal regions are described in U.S. Patent Application Publication No. 2009 / 0029377, entitled "Diagnosing Fetal Chromosomal Aneuploidy Using Massively Parallel Genomic Sequencing" by Lo et al., the entire contents of which are incorporated herein by reference for all purposes.

[0373] In embodiments, size analysis may be used instead of tag counting. Size analysis may also be used instead of normalized tag count. Size analysis may use various parameters described herein and in US Patent Application Publication No. 12 / 940,992. For example, the Q value or F value from above may be used. Such size values ​​do not require normalization by counting from other regions because these values ​​do not correspond to the number of reads. The techniques of haplotype-specific methods, such as the RHDO method described above and in more detail in US Patent Application Publication No. 13 / 308,473, may also be used for non-specific methods. For example, techniques including depth and refinement of regions may be used. In some embodiments, the GC tendency of a particular region may be taken into account when comparing two regions. Since the RHDO method uses the same region, such correction is not necessary.

[0374] Although certain cancers may typically have abnormalities in certain chromosomal regions, such cancers do not always have abnormalities exclusively in such regions. For example, additional chromosomal regions may show abnormalities, and the location of such additional regions may be unknown. Furthermore, when screening patients to identify early stages of cancer, it may be desirable to identify widespread cancers that may show abnormalities throughout the genome. To address these situations, in embodiments, multiple regions may be systematically analyzed to determine which regions show abnormalities. The number of abnormalities and their locations (e.g., whether they are close together) may be used, for example, to confirm abnormalities, determine the stage of cancer, diagnose cancer (e.g., if the number is greater than a threshold value), and make a prognosis based on the number and location of various regions showing abnormalities.

[0375] Thus, in embodiments, an organism may be identified as having cancer based on the number of regions that show abnormalities. Thus, multiple regions (e.g., 3,000) may be examined to identify the number of regions that show abnormalities. The regions may span the entire genome or may span only parts of the genome (e.g., non-repetitive regions).

[0376] 36 is a flow chart of a method 3600 of analyzing a biological sample of an organism using multiple chromosomal regions according to an embodiment of the present invention. The biological sample includes nucleic acid molecules (also referred to as fragments).

[0377] In block 3610, a plurality of regions (e.g., non-overlapping regions) of the genome of an organism are identified. Each chromosomal region includes a plurality of gene loci. The regions can be 1 Mb in size, or some other equivalent size. If the regions are 1 Mb in size, the entire genome can include about 3,000 regions, each having a predetermined size and location. Such predetermined regions can vary to accommodate the length of a particular chromosome or the specific number of regions used, and any other criteria described herein. If the regions have different lengths, such lengths can be used, for example, to normalize the results, as described herein. The regions can be specifically selected based on certain criteria of a particular organism and / or based on findings of the cancer being tested. The regions can also be selected arbitrarily.

[0378] At block 3620, a location of the nucleic acid molecule within the reference genome of the organism is identified for each of the plurality of nucleic acid molecules. The location may be determined by any of the methods described herein, for example, by sequencing the fragments to obtain sequenced tags and aligning the sequenced tags to the reference genome. The particular haplotype of the molecule may also be determined for haplotype-specific methods.

[0379] For each chromosomal region, blocks 3630-3650 are performed. In block 3630, each population of nucleic acid molecules is identified as originating from the chromosomal region based on the identified locations. Each population can include at least one nucleic acid molecule located at each of a plurality of loci of the chromosomal region. In one embodiment, the population can be fragments that align to a particular haplotype of the chromosomal region (e.g., to RHDO as described above). In another embodiment, the population can be any fragments that align to the chromosomal region.

[0380] At block 3640, the computer system calculates each value of each nucleic acid molecule group. Each value defines the identity of the nucleic acid molecule of each group. Each value can be any of the values ​​described herein. For example, the value can be the number of fragments in a group or a statistical value of the size distribution of fragments in a group. Each value can also be a normalized number, such as the number of tags in a region divided by the total number of tags in a sample or the number of tags in a reference region. Each value can be a difference or ratio from another value (e.g., in a RHDO), thereby obtaining a characteristic of the difference in the region.

[0381] In block 3650, each value is compared with a reference value to determine whether the first chromosomal region shows a deletion or amplification. The reference value may be any threshold or reference value described herein. For example, the reference value may be a threshold value determined in normal samples. In RHDO, each value may be a difference or ratio of tag numbers in two haplotypes, and the reference value may be a threshold value for determining that there is a statistically significant deviation. As another example, the reference value may be a tag number or size value of another haplotype or region, and the comparison may include obtaining a difference or ratio (or a function of such ratio) and then determining whether the difference or ratio is greater than a threshold value.

[0382] The reference value may vary based on the results of other regions. For example, if adjacent regions also show deviations (albeit small compared to a certain threshold (e.g., a Z value of 3)), a smaller threshold may be used. For example, if all three contiguous regions are above a first threshold, cancer is likely. This first threshold may therefore be lower than another threshold required to identify cancer from non-contiguous regions. Having three regions (or more than three) with small deviations may sufficiently reduce the chance of chance effects that sensitivity and specificity may be preserved.

[0383] At block 3660, the amount of genomic regions classified as exhibiting deletion or amplification is determined. There may be limitations on the chromosomal regions counted. For example, only regions that are adjacent to at least one other region may be counted (or adjacent regions may be required to be a certain size (e.g., 4 or more regions)). In embodiments where the regions are unequal, the number may also account for the length of each (e.g., the number may be the total length of the abnormal region).

[0384] At block 3670, the amount is compared to an amount threshold to determine a classification of the sample. By way of example, the classification may be whether the organism has cancer, the stage of the cancer, and the prognosis of the cancer. In one embodiment, all abnormal regions are counted and a single threshold is used regardless of where they appear. In another embodiment, the threshold may vary based on the location and size of the counted regions. For example, the amount of a region on a particular chromosome or chromosome arm may be compared to a threshold for that particular chromosome (or arm). Multiple thresholds may be used. For example, the amount of an abnormal region on a particular chromosome (or arm) must be greater than a first threshold, and the total amount of abnormal regions in the genome must be greater than a second threshold. The threshold may be a percentage of the region determined to exhibit deletion or amplification.

[0385] This threshold for the amount of regions may also depend on how strong the imbalance is in the counted regions. For example, the amount of regions used as a threshold for determining the classification of cancer may depend on the specificity and sensitivity (abnormal threshold) used to detect abnormalities in each region. For example, if the abnormal threshold is low (e.g., Z value of 2), the amount threshold may be selected to be high (e.g., 150). However, if the abnormal threshold is high (e.g., Z value of 3), the amount threshold may be low (e.g., 50). Also, the amount of regions showing abnormalities may be a weighted value, for example, one region showing a high degree of imbalance may be weighted higher than a region showing only a small imbalance (i.e., there are more classifications for abnormalities than just positive and negative). As an example, the sum of Z values ​​may be used, whereby a weighted value is used.

[0386] Therefore, the amount (which may include number and / or size) of chromosomal regions that show significant over- or under-representation of normalized tag counts (or other values ​​in the group characteristic) can be used to reflect the severity of disease. The amount of chromosomal regions with abnormal normalized tag counts can be determined by two factors: the number (or size) of chromosomal abnormalities in tumor tissue and the concentration fraction of tumor-derived DNA in biological samples (e.g., plasma). More advanced cancers tend to show more (and larger) chromosomal abnormalities. Thus, more cancer-associated chromosomal abnormalities may be detectable in samples (e.g., plasma). In patients with more advanced cancers, higher tumor burden will result in a higher concentration fraction of tumor-derived DNA in plasma. As a result, tumor-associated chromosomal abnormalities will be more easily detected in plasma samples.

[0387] One possible approach to improve sensitivity without sacrificing specificity is to consider the results of adjacent chromosomal segments. In one embodiment, the cutoff of Z-score remains greater than 2 and less than -2. However, a chromosomal region is classified as potentially abnormal only if two consecutive segments show the same type of abnormality (e.g., both segments have Z-score greater than 2). In other embodiments, the Z-scores of adjacent segments can be summed, using a higher cutoff value. For example, the Z-scores of three consecutive segments can be summed, and a cutoff value of 5 can be used. This concept can be extended to more than three consecutive segments.

[0388] The combination of quantity and abnormal threshold may also depend on the purpose of the analysis and any prior knowledge (or lack thereof) of the organism.For example, when screening a healthy population for cancer, typically high specificity is used, both in the quantity of regions (i.e., high threshold for the number of regions) and possibly the abnormal threshold when a region is identified as having abnormality.However, for patients with higher risk (e.g., patients who complain of lump or family history, smokers, chronic human papillomavirus (HPV) carriers, hepatitis virus carriers, or other virus carriers), the threshold may be lowered to have higher sensitivity (fewer false negatives).

[0389] In one embodiment, if a lower detection limit of 1 Mb resolution and 6.3% tumor-derived DNA is used to detect chromosomal abnormalities, the number of molecules in each 1 Mb segment would need to be 60,000. This translates to approximately 180 million alignable reads in the entire genome (60,000 reads / Mb x 3,000 Mb).

[0390] Smaller segment sizes provide higher resolution for detecting smaller chromosomal abnormalities. However, this increases the requirement for the number of molecules analyzed in total. Larger segment sizes reduce the number of molecules required for analysis at the expense of resolution. Thus, only larger abnormalities can be detected. In some implementations, larger regions can be used, segments showing abnormalities can be subdivided, and these smaller regions can be analyzed to obtain better resolution (e.g., as described above). If one has an estimate for the size of the deletion or amplification to be detected (or the minimum concentration for detection), the number of molecules analyzed can be reduced.

[0391] B. CNA based on sequencing of bisulfite-treated plasma DNA Genome-wide hypomethylation and CNAs can be frequently observed in tumor tissue. Here, we show that information of CNAs and cancer-associated methylation changes can be obtained simultaneously from bisulfite sequencing of plasma DNA. Since the two types of analysis can be performed on the same data set, there is virtually no additional cost for CNA analysis. In other embodiments, different procedures can be used to obtain methylation information and genetic information. In other embodiments, a similar analysis for cancer-associated hypermethylation can be performed in combination with CNA analysis.

[0392] FIG. 37A shows the CNA analysis (inside to outside) of tumor tissue, non-bisulfite (BS) treated plasma DNA and bisulfite treated plasma DNA of patient TBR36. FIG. 37A shows the CNA analysis (inside to outside) of tumor tissue, non-bisulfite (BS) treated plasma DNA and bisulfite treated plasma DNA of patient TBR36. The outermost ring shows the chromosome diagram. Each dot represents the result of a 1 Mb region. The green, red and grey dots represent regions with copy number gain, copy number loss and no copy number change, respectively. In the plasma analysis, the Z-score is shown. There is a difference of 5 between the two concentric lines. In the tumor tissue analysis, the copy number is shown. There is a copy difference of 1 between the two concentric lines. FIG. 38A shows the CNA analysis of tumor tissue, non-bisulfite (BS) treated plasma DNA and bisulfite treated plasma DNA of patient TBR34. The patterns of CNAs detected in bisulfite-treated and non-bisulfite-treated plasma samples were consistent.

[0393] The patterns of CNA detected in tumor tissue, non-bisulfite treated plasma and bisulfite treated plasma were consistent. To further evaluate the agreement between the results of bisulfite treated plasma and non-bisulfite treated plasma, scatter plots are generated. Figure 37B is a scatter plot showing the correlation between Z-scores in the detection of CNA using bisulfite treated plasma and non-bisulfite treated plasma of 1 Mb bins of patient TBR36. A positive correlation was observed between the Z-scores of the two analyses (r=0.89, p<0.001, Pearson correlation). Figure 38B is a scatter plot showing the correlation between Z-scores in the detection of CNA using bisulfite treated plasma and non-bisulfite treated plasma of 1 Mb bins of patient TBR34. A positive correlation was observed between the Z-scores of the two analyses (r=0.81, p<0.001, Pearson correlation).

[0394] C. Synergistic Analysis of Cancer-Related CNAs and Methylation Alterations As mentioned above, the analysis of CNA may include counting the number of sequence reads in each 1 Mb region, while the analysis of methylation density may include detecting the proportion of cytosine residues in CpG dinucleotides that are methylated. The combination of these two analyses may provide synergistic information in cancer detection. For example, the methylation classification and the CNA classification may be used to determine a third classification of cancer levels.

[0395] In one embodiment, the presence of either cancer-related CNA or methylation change can be used to indicate the potential presence of cancer. In such an embodiment, the sensitivity of detecting cancer can be increased when CNA or methylation change is present in the plasma of the test subject. In another embodiment, the presence of both changes can be used to indicate the presence of cancer. In such an embodiment, the specificity of the test can be improved because either of these two types of changes can potentially be detected in some non-cancer subjects. Thus, the third classification can be positive for cancer only if both the first and second classifications indicate cancer.

[0396] 26 HCC patients and 22 healthy subjects were recruited. Blood samples were collected from each subject, and plasma DNA was sequenced after bisulfite treatment. In HCC patients, blood samples were collected at the time of diagnosis. The presence of significant amounts of CNA was defined as having more than 5% of bins with a Z-score of, for example, less than -3 or more than 3. The presence of significant amounts of cancer-associated hypomethylation was defined as having more than 3% of bins with a Z-score of less than -3. As an example, the amount of a region (bin) can be expressed as the raw number of bins, the percentage, and the length of the bin.

[0397] Table 3 shows the detection of significant amounts of CNAs and methylation changes in the plasma of 26 HCC patients using massively parallel sequencing on bisulfite-treated plasma DNA.

[0398] [Table 3]

[0399] The detection rates for cancer-associated methylation changes and CNAs were 69% and 50%, respectively. When the presence of either criterion was used to indicate the potential presence of cancer, the detection rate (i.e., diagnostic sensitivity) improved to 73%.

[0400] Results are shown for two patients showing the presence of CNA (Figure 39A) or methylation changes (Figure 39B). Figure 39A is a Circos plot showing CNA (inner ring) and methylation analysis (outer ring) of bisulfite-treated plasma of HCC patient TBR240. In the CNA analysis, green, red and grey dots represent regions with chromosomal gains, chromosomal losses and no copy number changes, respectively. In the methylation analysis, green, red and grey dots represent regions with hypermethylation, hypomethylation and normal methylation, respectively. In this patient, cancer-associated CNA was detected in the plasma, but the methylation analysis did not show a significant amount of cancer-associated hypomethylation. Figure 39B is a Circos plot showing CNA (inner ring) and methylation analysis (outer ring) of bisulfite-treated plasma of HCC patient TBR164. In this patient, cancer-associated hypomethylation was detected in the plasma. However, no significant amount of CNA could be observed. The results for two patients showing the presence of both CNA and methylation changes are shown in Figure 48A (TBR36) and Figure 49A (TBR34).

[0401] Table 4 shows the detection of significant amounts of CNA and methylation changes in the plasma of 22 control patients using massively parallel sequencing on bisulfite-treated plasma DNA. A bootstrapping (i.e., leave-one-out) method was used to evaluate each control patient. Thus, when a particular subject was evaluated, the other 21 subjects were used to calculate the mean and SD of the control group.

[0402] [Table 4]

[0403] The specificity of detecting significant amounts of methylation changes and CNAs was 86% and 91%, respectively, and increased to 95% when the presence of both criteria was required to indicate the potential presence of cancer.

[0404] In one embodiment, samples positive for CNA and / or hypomethylation are considered positive for cancer, and samples are considered negative if both are not detectable. Using "or" logic provides greater sensitivity. In another embodiment, only samples positive for both CNA and hypomethylation are considered positive for cancer, thereby providing greater specificity. In yet another embodiment, a three-tier classification may be used. Subjects are classified as i. both normal; ii. one abnormal; iii. both abnormal.

[0405] Different follow-up strategies can be used for these three classifications.For example, (iii) subjects can be subjected to the most intensive follow-up protocol (e.g., including whole body imaging); (ii) subjects can be subjected to a less intensive follow-up protocol (e.g., repeated plasma DNA sequencing after a relatively short time interval of several weeks); (i) subjects can be subjected to the least intensive follow-up protocol (e.g., re-examination after several years).In other embodiments, methylation and CNA measurements can be combined with other clinical parameters (e.g., imaging results or serum biochemistry) to further subdivide classification.

[0406] D. Prognostic Value of Plasma DNA Analysis Following Curative Treatment The presence of cancer-related CNA and / or methylation changes in plasma indicates the presence of tumor-derived DNA in the circulating blood of cancer patients. Reduction or elimination of these cancer-related changes is expected after treatment (e.g., surgery). Meanwhile, persistence of these changes in plasma after treatment may indicate incomplete removal of all tumor cells from the body, and may be a useful prognostic factor for disease recurrence.

[0407] Blood samples were collected from two HCC patients, TBR34 and TBR36, one week after curative surgical resection of the tumor. CNA and methylation analysis was performed on bisulfite-treated post-treatment plasma samples.

[0408] Figure 40A shows CNA analysis on bisulfite-treated plasma DNA taken before (inner ring) and after (outer ring) surgical resection of the tumor of HCC patient TBR36. Each dot represents the result of a 1 Mb region. Green, red, and gray dots represent regions with copy number gain, copy number loss, and no change in copy number, respectively. The majority of CNAs observed before treatment disappeared after tumor resection. The percentage of bins showing Z-scores below -3 or above 3 decreased from 25% to 6.6%.

[0409] Figure 40B shows methylation analysis on bisulfite-treated plasma DNA taken before (inner ring) and after (outer ring) surgical resection of the tumor of HCC patient TBR36. Green, red, and grey dots represent regions with hypermethylation, hypomethylation, and normal methylation, respectively. There was a significant decrease in the percentage of bins showing significant hypomethylation from 90% to 7.9%, as well as a significant decrease in the degree of hypomethylation. The patient achieved complete clinical remission 22 months after tumor resection.

[0410] Figure 41A shows CNA analysis on bisulfite-treated plasma DNA collected before (inner ring) and after (outer ring) surgical resection of the tumor of HCC patient TBR34. Although the number of bins exhibiting CNA and the magnitude of CNA in affected bins both decrease after surgical resection of the tumor, residual CNA can be observed in the post-surgery plasma sample. The red ring highlights the area where residual CNA was most evident. The percentage of bins exhibiting Z-scores below -3 or above 3 decreased from 57% to 12%.

[0411] Figure 41B shows methylation analysis on bisulfite-treated plasma DNA collected before (inner ring) and after (outer ring) surgical resection of the tumor of HCC patient TBR34. The magnitude of hypomethylation decreased after tumor resection, with the mean Z-score of hypomethylated bins decreasing from -7.9 to -4.0. However, the percentage of bins with Z-scores below -3 showed the opposite change, increasing from 41% to 85%. This observation potentially indicates the presence of residual cancer cells after treatment. Clinically, multiple foci of tumor nodules were detected in the remaining non-resected liver 3 months after tumor resection. Pulmonary metastases were observed starting 4 months after surgery. The patient died 8 months after surgery from local recurrence and metastatic disease.

[0412] The observations in these two patients (TBR34 and TBR36) suggest that the presence of residual cancer-associated alterations in CNA and hypomethylation can be used to monitor and prognose cancer patients after therapeutic treatment. The data also showed that the degree of change in the amount of detected plasma CNA can be used synergistically with the assessment of the degree of change in the extent of plasma DNA hypomethylation for prognosticating and monitoring the efficacy of treatment.

[0413] Thus, in some embodiments, one biological sample is obtained before treatment and a second biological sample is obtained after treatment (e.g., surgery). A first value is obtained in the first sample, such as the Z-score of the region (e.g., region methylation level and normalized number of CNA) and the number of regions showing hypomethylation and CNA (e.g., amplification or deletion). A second value can be obtained in the second sample. In another embodiment, a third or additional sample can be obtained after treatment. The number of regions showing hypomethylation and CNA (e.g., amplification or deletion) can be obtained from the third or additional sample.

[0414] As described above for Figures 40A and 41A, a first number of regions exhibiting hypomethylation in a first sample can be compared to a second amount of regions exhibiting hypomethylation in a second sample. As described above for Figures 40B and 41B, a first amount of regions exhibiting hypomethylation in a first sample can be compared to a second amount of regions exhibiting hypomethylation in a second sample. A comparison of the first amount to the second amount and the first number to the second number can be used to determine a prognosis of the treatment. In various embodiments, only one of the comparisons can be determinative of the prognosis, or both comparisons can be used. In embodiments in which a third or additional sample is obtained, one or more of these samples can be used to determine a prognosis of the treatment by themselves or in combination with the second sample.

[0415] In one implementation, the prognosis is predicted to be worse when the first difference between the first amount and the second amount is below a first difference threshold. In another implementation, the prognosis is predicted to be worse when the second difference between the first number and the second number is below a second difference threshold. The thresholds may be the same or different. In one embodiment, the first difference threshold and the second difference threshold are zero. Thus, in the above example, the difference between the methylation values ​​indicates a worse prognosis for patient TBR34.

[0416] Prognosis may be better when the first difference and / or the second difference are above the same or each threshold. Prognosis classification may depend on how much difference is above or below the threshold. Multiple thresholds may be used to give different classifications. A larger difference may predict a better prognosis, and a smaller difference (and even a negative value) may predict a worse prognosis.

[0417] In some embodiments, the time points at which the various samples are taken are also recorded. Such time parameters can be used to determine the kinetics or rate of change of the amount. In one embodiment, a rapid decrease in tumor-associated hypomethylation in plasma and / or a rapid decrease in tumor-associated CNA in plasma is predictive of a good prognosis. Conversely, a static or rapid increase in tumor-associated hypomethylation in plasma and / or a static or rapid increase in tumor-associated CNA is predictive of a poor prognosis. Measurement of methylation and CNA can be used in combination with other clinical parameters (e.g., imaging results or serum biochemistry or protein markers) to predict clinical outcome.

[0418] In embodiments, other samples besides plasma may be used. For example, tumor-associated methylation abnormalities (e.g., hypomethylation) and / or tumor-associated CNAs may be measured from tumor cells circulating in the blood of cancer patients, from cell-free DNA or tumor cells in urine, stool, saliva, sputum, bile fluid, pancreatic juice, cervical swabs, secretions from the reproductive system (e.g., vagina), peritoneal fluid, pleural fluid, semen, sweat, and tears.

[0419] In various embodiments, tumor-associated methylation abnormalities (e.g., hypomethylation) and / or tumor-associated CNAs can be detected in the blood or plasma of patients with breast cancer, lung cancer, colorectal cancer, pancreatic cancer, ovarian cancer, nasopharyngeal cancer, cervical cancer, melanoma, brain tumors, and the like. Indeed, the approach can be used for all cancer types, since genetic alterations such as methylation and CNAs are universal phenomena in cancer. Measurements of methylation and CNAs can be used in combination with other clinical parameters (e.g., imaging results) for predicting clinical outcomes. Embodiments can also be used for screening and monitoring patients with precancerous lesions (e.g., adenomas).

[0420] Thus, in one embodiment, biological samples are collected before treatment, and CNA and methylation measurements are repeated after treatment. From said measurements, a first amount after the region determined to show deletion or amplification can be obtained, and a second amount after the region determined to have a region methylation level above each region cutoff value can be obtained. The first amount is compared to the first amount after, and the second amount is compared to the second amount after, to determine the prognosis of the organism.

[0421] The comparison to determine a prognosis of the organism may include determining a first difference between the first amount and the later first amount, where the first difference may be compared to one or more first difference thresholds to determine the prognosis. The comparison to determine a prognosis of the organism may include determining a second difference between the second amount and the later second amount, where the second difference may be compared to one or more second difference thresholds. The threshold may be zero or another number.

[0422] The prognosis may be predicted to be worse when the first difference is below the first difference threshold than when the first difference is above the first difference threshold. The prognosis may be predicted to be worse when the second difference is below the second difference threshold than when the second difference is above the second difference threshold. Examples of treatments include immunotherapy, surgery, radiation therapy, chemotherapy, antibody-based therapy, gene therapy, epigenetic therapy, or targeted therapy.

[0423] E. Performance The diagnostic performance of different numbers of sequence reads and different numbers of bin sizes in CNA and methylation analysis is described below.

[0424] 1. Number of sequence reads According to one embodiment, plasma DNA was analyzed from 32 healthy control patients, 26 patients with hepatocellular carcinoma, and 20 patients with other cancer types (e.g., nasopharyngeal carcinoma, breast cancer, lung cancer, neuroendocrine carcinoma, and leiomyosarcoma). 22 of the 32 healthy subjects were randomly selected as the reference group. The mean and standard deviation (SD) of these 22 reference individuals were used to determine the normal range of methylation density and genomic representation. DNA extracted from each individual's plasma sample was used for sequencing library construction using an Illumina paired-end sequencing kit. The sequencing library was then subjected to bisulfite treatment to convert unmethylated cytosine residues to uracil. The bisulfite-treated sequencing library of each plasma sample was sequenced using one lane of an Illumina HiSeq2000 sequencer.

[0425] After base calling, adapter sequences at the ends of fragments and low quality bases (i.e., quality scores less than 5) were removed. The trimmed reads in FASTQ format were then processed by a methylation data analysis pipeline called Methy-Pipe (P Jiang et al. 2010, IEEE International Conference on Bioinformatics and Biomedicine, doi:10.1109 / BIBMW.2010.5703866). To align the bisulfite-converted sequencing reads, first, all cytosine residues were in silico converted to thymine on the Watson and Crick strands separately using the reference human genome (NCBI build 36 / hg19). Then, in silico conversion of each cytosine to thymine was performed in all processed reads, preserving the positional information of each converted residue. SOAP2 was used to align the converted reads to two pre-converted reference human genomes (R Li et al. 2009 Bioinformatics 25:1966-1967), and up to two mismatches were allowed in each of the aligned reads. Only reads mappable to unique genomic locations were used for downstream analysis. Ambiguous reads and duplicate (clonal) reads mapping to both Watson and Crick strands were eliminated. Cytosine residues within CpG dinucleotide sequences were used for downstream methylation analysis. After alignment, cytosines originally present on the sequenced reads were restored based on the position information retained during in silico conversion. The restored cytosines in CpG dinucleotides were scored as methylated. Thymines in CpG dinucleotides were scored as unmethylated.

[0426] In the methylation analysis, the genome was divided into bins of equal size. The bin sizes examined included 50kb, 100kb, 200kb and 1Mb. The methylation density of each bin was calculated as the number of methylated cytosines in the sequence of CpG dinucleotides divided by the total number of cytosines at CpG positions. In other embodiments, the bin sizes may be unequal across the genome. In one embodiment, each bin in such unequal sized bins is compared across multiple subjects.

[0427] To determine whether the plasma methylation densities of the test cases were normal, the methylation densities were compared with those of the reference group. Twenty-two of 32 healthy subjects had a methylation Z score (Z meth ) were randomly selected as the reference group.

[0428]

number

[0429] During the ceremony,

[0430]

number

[0431] is the methylation density of a particular 1 Mb bin of the test case;

[0432]

number

[0433]

number

[0434] is the SD of the methylation density of the corresponding bin in the reference group.

[0435] In the CNA analysis, the number of sequenced reads mapping to each 1 Mb bin was determined (KCA Chan el al. 2013 Clin Chem 59:211-24). After correction for GC bias using Locally Weighted Scatter Plot Smoothing regression as previously reported (EZ Chen et al. 2011 PLoS One 6: e21791), the sequenced read density in each bin was determined. In the plasma analysis, the sequenced read density of the test cases was compared to the reference group to obtain the Z-score of CNA,

[0436]

number

[0437] was calculated.

[0438]

number

[0439] During the ceremony,

[0440]

number

[0441] is the sequenced read density of a particular 1 Mb bin of the test case;

[0442]

number

[0443] is the average sequenced read density of the corresponding bin in the reference group;

[0444]

number

[0445] was the SD of the sequenced read density of the corresponding bin in the reference group. CNA A CNA score < −3 or > 3 was defined as indicating CNA.

[0446] A mean of 93 million aligned reads (range: 39–142 million) were obtained per case. To assess the impact of a reduced number of sequenced reads on diagnostic performance, 10 million aligned reads were randomly selected from each case. The same reference population was used to establish reference ranges for each 1 Mb bin in the dataset with reduced sequenced reads. Significant hypomethylation (i.e., Z meth The proportion of bins showing CNA (i.e., Z < -3) CNA The proportion of bins with a β score <-3 or >3 was determined for each case. Receiver operating characteristic (ROC) curves were used to illustrate the diagnostic performance of genome-wide hypomethylation and CNA analysis for a dataset with all sequenced reads from one lane and 10 million reads per case. In the ROC analysis, all 32 healthy subjects were included in the analysis.

[0447] Figure 42 shows the diagnostic performance of genome-wide hypomethylation analysis using different numbers of sequenced reads. In hypomethylation analysis, the area under the curve of the ROC curve was not significantly different between the two datasets that analyzed all sequenced reads from one lane and 10 million reads per case (P=0.761). In CNA analysis, diagnostic performance deteriorated with a significant decrease in the area under the curve when the number of sequenced reads was reduced from using one lane of data to 10 million (P<0.001).

[0448] 2. The impact of using different bin sizes Besides dividing the genome into 1 Mb bins, it was also investigated whether smaller bin sizes could be used. In theory, the use of smaller bins could reduce the variability in methylation density within a bin, since methylation densities can vary widely between different genomic regions. If the bins are larger, the probability of containing regions with different methylation densities increases, leading to an overall increase in the variability in the methylation density of the bins.

[0449] Using smaller bin sizes can reduce the variability in methylation density associated with inter-region differences, but this, in turn, reduces the number of sequenced reads that map to a particular bin. The reduction in the number of reads that map to individual bins increases variability due to sampling variation. The optimal bin size that can produce the smallest overall variability in methylation density can be experimentally determined depending on the requirements of a particular diagnostic application (e.g., the total number of sequenced reads per sample and the type of DNA sequencing device used).

[0450] Figure 43 shows the ROC curves for cancer detection based on genome-wide hypomethylation analysis using different bin sizes (50kb, 100kb, 200kb and 1Mb). The P-values ​​shown are the P-values ​​for area under the curve comparison using a bin size of 1Mb. A trend of improvement can be seen when the bin size is reduced from 1Mb to 200kb.

[0451] F. Cumulative Probability Score The amount of methylation and CNA regions can be various values. The above example describes the number of regions exceeding a cutoff value or the percentage of regions showing significant hypomethylation or CNA as parameters for classifying whether a sample is associated with cancer. Such an approach does not take into account the magnitude of abnormalities in individual bins. For example, a Z of -3.5 meth and a bin with Z of -30 meth Bins with Z are identical because they are both classified as having significant hypomethylation. However, the degree of hypomethylation change in plasma (i.e., Z methThe magnitude of the value is influenced by the amount of cancer-associated DNA in the sample and can therefore complement the information of the percentage of bins showing abnormalities to reflect the tumor burden. A higher concentration fraction of tumor DNA in the plasma sample leads to a lower methylation density, which in turn leads to a lower Z meth is converted to a value.

[0452] 1. Cumulative Probability Score as a Diagnostic Parameter To utilize the information gained from the magnitude of anomalies, an approach called the Cumulative Probability (CP) score is developed. Based on the normal distribution probability function, meth Values ​​were converted to the probability of obtaining such an observation by chance.

[0453] CP scores are Z less than -3. meth For bin(i) having

[0454]

number

[0455] where Prob i is the Z value for bin (i) that follows a Student's t distribution with 3 degrees of freedom. meth where log is the natural logarithm function. In another embodiment, the logarithm down to 10 (or other number) may be used. In other embodiments, other distributions (such as, but not limited to, normal and gamma distributions) may be applied to convert the Z-score to CP.

[0456] A larger CP score indicates a lower probability of having such aberrant methylation density by chance in a normal population, and thus indicates a higher likelihood of having aberrant hypomethylated DNA in the sample (e.g., the presence of cancer-associated DNA).

[0457] Compared with the percentage of bins showing abnormality, CP score measurement has a higher dynamic range. Although tumor burden between different patients can vary widely, a larger range of CP values ​​is useful to reflect the tumor burden of patients with relatively high and relatively low tumor burden. In addition, the use of CP score can potentially be more sensitive to detect changes in the concentration of tumor-associated DNA in plasma. This is advantageous for monitoring treatment response and prognosis. Thus, a decrease in CP score during treatment indicates a good treatment response. A lack of decrease or even an increase in CP score during treatment indicates a poor response or lack of response. In prognosis, a high CP score indicates a high tumor burden, suggesting a poor prognosis (e.g., a higher probability of death or tumor progression).

[0458] Figure 44A shows the diagnostic performance of cumulative probability (CP) and proportion of bins with abnormalities. There was no significant difference between the areas under the curve of the two types of diagnostic algorithms (P=0.791).

[0459] Figure 44B shows the diagnostic performance of plasma analysis of global hypomethylation, CpG island hypermethylation and CNA. With one lane of sequencing per sample (200 kb bin size for hypomethylation analysis, 1 Mb bin size for CNA, and CpG island defined according to a database hosted by the University of California, Santa Cruz (UCSC)), the area under the curve for all three analyses was greater than 0.90.

[0460] In subsequent analyses, the highest CP score in control patients was used as the cutoff for each of the three analyses. The selection of these cutoffs gave a diagnostic specificity of 100%. The diagnostic sensitivity of global hypomethylation, CpG island hypermethylation and CNA analyses was 78%, 89% and 52%, respectively. At least one of these three abnormalities was detected in 43 out of 46 cancer patients, which gave a sensitivity of 93.4% and a specificity of 100%. These results indicate that these three analyses can be used synergistically to detect cancer.

[0461] Figure 45 shows a table including the results of global hypomethylation, CpG island hypermethylation and CNA in patients with hepatocellular carcinoma. The cutoff values ​​of CP score in these three analyses were 960, 2.9 and 211, respectively. The positive CP score results were bolded and underlined.

[0462] Figure 46 shows a table including the results of global hypomethylation, CpG island hypermethylation and CNA in patients with cancer other than hepatocellular carcinoma. The cutoff values ​​of CP score in these three analyses were 960, 2.9 and 211, respectively. Positive CP score results are bolded and underlined.

[0463] 2. Application of the CP score to cancer monitoring Serial samples were taken from HCC patient TBR34 before and after treatment and were analyzed for global hypomethylation.

[0464] Figure 47 shows the serial analysis of plasma methylation in case TBR34. The innermost ring shows the methylation density of the buffy coat (black) and tumor tissue (purple). In these plasma samples, the Z meth The difference between the two lines is a Z of 5. methThe red and grey dots represent bins with hypomethylation and no change in methylation density compared to the reference group. From the second innermost ring outwards, plasma samples were taken before treatment, 3 days after tumor resection and 2 months after tumor resection, respectively. Before treatment, a high degree of hypomethylation can be observed in the plasma, with more than 18.5% of the bins having a Z below -10. meth Three days after tumor resection, it can be observed that the degree of hypomethylation in plasma was decreased, with all bins showing a Z value below -10. meth did not have.

[0465] [Table 5]

[0466] Table 5 shows that the magnitude of hypomethylation changes decreased 3 days after surgical tumor resection, whereas the proportion of abnormal bins increased. On the other hand, the CP score may more accurately indicate the decrease in the degree of hypomethylation in plasma and better reflect the change in tumor burden.

[0467] Even 2 months after surgery (OT), there was a significant proportion of bins showing hypomethylation changes. The CP score also remained static at approximately 15,000. The patient was later diagnosed with multifocal tumor deposits (previously unknown at the time of surgery) in the remaining non-resected liver at 3 months, and was reported to have multiple lung metastases 4 months after surgery. The patient died of metastatic disease 8 months after surgery. These results suggested that the CP score may be a more powerful reflection of tumor burden than the proportion of bins with abnormalities.

[0468] Overall, CP can be useful in applications that require the measurement of the amount of tumor DNA in plasma. Examples of such applications include prognosis and monitoring of cancer patients (e.g., to observe response to treatment or to observe tumor progression).

[0469] The cumulative Z score is a direct sum of the Z scores (i.e., there is no conversion to probability). In this example, the cumulative Z score behaves the same as the CP score. In other examples, CP may be more sensitive than the cumulative Z score in monitoring residual disease due to the larger dynamic range of the CP score.

[0470] X. Effect of CNAs on Methylation The use of CNA and methylation to determine each classification of cancer levels, where these classifications are combined to give a third classification, has been described above. In addition to such combinations, CNA can be used to change the cutoff value in the methylation analysis and to identify false positives by comparing the methylation levels of regions with different CNA features. For example, excessive methylation levels (e.g., Z > 3) can be used to identify false positives. CNA ) is the normal abundance of methylation levels (e.g., −3 <Z CNA <3). First, the effect of CNA on methylation levels is described.

[0471] A. Changes in methylation density in regions with chromosomal gains and losses Because tumor tissues generally exhibit global hypomethylation, the presence of tumor-derived DNA in the plasma of cancer patients results in a decreased methylation density when compared to non-cancer controls. The degree of hypomethylation in the plasma of cancer patients is theoretically proportional to the fractional concentration of tumor-derived DNA in the plasma sample.

[0472] In the region of chromosomal gain in tumor tissue, an additional amount of tumor DNA is released from the amplified DNA segment into plasma.This increased contribution of tumor DNA to plasma theoretically leads to a higher degree of hypomethylation in the plasma DNA of the affected region.Another factor is that the genomic region showing amplification is expected to confer growth advantage to tumor cells, and therefore is expected to be expressed.Such regions are generally hypomethylated.

[0473] In contrast, in regions that show chromosomal loss in tumor tissue, the reduced contribution of tumor DNA to plasma leads to lower hypomethylation compared to regions that have no copy number change.Another factor is that the genomic regions that are deleted in tumor cells may contain tumor suppressor genes, and it may be advantageous for tumor cells to have such silenced regions.Therefore, it is expected that such regions are more likely to be hypermethylated.

[0474] Here, the results of two HCC patients (TBR34 and TBR36) are used to illustrate this effect. Figure 48A (TBR36) and Figure 49A (TBR34) have rings highlighting the regions with chromosomal gains or losses, and the corresponding methylation analysis. Figure 48B and Figure 49B show plots of methylation z-scores for losses, normals, and gains in TBR36 and TBR34 patients, respectively.

[0475] FIG. 48A shows a Circos plot showing CNAs (inner ring) and methylation changes (outer ring) in bisulfite-treated plasma DNA from HCC patient TBR36. The red ring highlights regions with chromosomal gains or losses. Regions with chromosomal gains were more hypomethylated than regions with no copy number changes. Regions with chromosomal losses were less hypomethylated than regions with no copy number changes. FIG. 48B shows a plot of methylation Z-scores for regions with chromosomal gains and losses, as well as regions with no copy number changes, from HCC patient TBR36. Compared to regions with no copy changes, regions with chromosomal gains had more negative Z-scores (more hypomethylated) and regions with chromosomal losses had less negative Z-scores (less hypomethylated).

[0476] Figure 49A shows a Circos plot showing CNAs (inner ring) and methylation changes (outer ring) in bisulfite-treated plasma DNA from HCC patient TBR34. Figure 49B shows a plot of methylation Z-scores for regions with chromosomal gains and losses and regions with no copy number changes from HCC patient TBR34. The difference in methylation density between regions with chromosomal gains and losses was greater in patient TBR36 than in patient TBR34, due to the higher fractional concentration of tumor-derived DNA in patient TBR36.

[0477] In this example, the region used to determine CNA is the same as the region used to determine methylation. In one embodiment, each region cutoff value depends on whether each region shows deletion or amplification. In one implementation, each region cutoff value (e.g., Z score cutoff used to determine hypomethylation) has a larger magnitude (e.g., the magnitude can be greater than 3 and a cutoff less than -3 can be used) when each region shows amplification than when amplification is not shown. Thus, in testing for hypomethylation, each region cutoff value can have a more negative value when each region shows amplification than when amplification is not shown. Such implementation is expected to improve the specificity of the test for detecting cancer.

[0478] In another implementation, each region cutoff value has a smaller magnitude (e.g., less than 3) when each region shows deletion than when deletion is not shown. Thus, in a test for hypomethylation, each region cutoff value may have a less negative value when each region shows deletion than when deletion is not shown. Such an implementation is expected to improve the sensitivity of the test for detecting cancer. The adjustment of the cutoff value in the above implementation may vary depending on the desired sensitivity and specificity of a particular diagnostic scenario. In other embodiments, the measurement of methylation and CNA may be used to predict cancer in combination with other clinical parameters (e.g., imaging results or serum biochemistry).

[0479] B. Use of CNA to Select Regions As mentioned above, it has been shown that plasma methylation density changes in regions with copy number abnormalities in tumor tissue. In regions with copy number gains in tumor tissue, the increased contribution of hypomethylated tumor DNA to plasma leads to a greater degree of hypomethylation of plasma DNA compared to regions without copy number abnormalities. Conversely, in regions with copy number losses in tumor tissue, the decreased contribution of hypomethylated cancer-derived DNA to plasma leads to a lower degree of hypomethylation of plasma DNA. This association between plasma DNA methylation density and relative presentation can potentially be used to distinguish the results of hypomethylation associated with the presence of cancer-related DNA from other non-cancerous causes of hypomethylation in plasma DNA (e.g., SLE).

[0480] To illustrate this approach, plasma samples from two hepatocellular carcinoma (HCC) patients and two patients without cancer but with SLE were analyzed. These two SLE patients (SLE04 and SLE10) showed hypomethylation and the apparent presence of CNA in plasma. In patient SLE04, 84% of the bins showed hypomethylation and 11.2% of the bins showed CNA. In patient SLE10, 10.3% of the bins showed hypomethylation and 5.7% of the bins showed CNA.

[0481] Figures 50A and 50B show the results of plasma hypomethylation and CNA analysis for SLE patients SLE04 and SLE10. The outer ring shows the methylation Z-scores (Z meth ) indicates methylation of less than -3 Z meth Bins with Z > -3 are in red. meth Bins with α were grey. The inner ring represents the Z-score of the CNA (Z CNA ) The green, red, and gray dots indicate Z values ​​greater than 3, less than 3, and between -3 and 3, respectively. CNA In these two SLE patients, hypomethylation and CNA alterations were observed in the plasma.

[0482] To determine whether changes in methylation and CNAs were consistent with the presence of cancer-derived DNA in plasma, we performed cytotoxicity assays with Z >3, <3, and -3 to 3. CNA Z of the area having meth The methylation changes and CNAs contributed by cancer-derived DNA in plasma were compared. CNA Regions with meth In contrast, Z > 3 CNA Regions with meth For illustration, a one-tailed rank sum test was applied to determine the regions with CNAs (i.e., Z < -3 or > 3). CNA (area having Z meth However, the region without CNA (i.e., Z between -3 and 3) CNA (area having Z meth In other embodiments, other statistical tests may be used, such as, but not limited to, Student's t-test, analysis of variance (ANOVA) test, and Kruskal-Wallis test.

[0483] Figures 51A and 51B show the Z-reactivity profile for regions with and without CNA in the plasma of two HCC patients (TBR34 and TBR36).meth Analysis shows Z less than -3 CNA and areas with Z > 3 CNA In both TBR34 and TBR36, regions that are under-represented in plasma (i.e., regions with a Z less than -3) are shown to be over-represented and under-represented in plasma. CNA The region with normal expression in plasma (i.e., Z between -3 and 3) is CNA (regions with significantly higher Z meth (P value < 10 -5 , one-tailed rank sum test). Normal representation is consistent with that expected for a euploid genome. Regions with over-representation in plasma (i.e., Z > 3) CNA In regions with normal expression in plasma, they had significantly lower Z meth (P value < 10 -5 , one-tailed rank-sum test). All these changes were consistent with the presence of hypomethylated tumor DNA in the plasma samples.

[0484] Figures 51C and 51D show the Z-regulation of regio...

Claims

1. A method for analyzing a biological sample of an organism, wherein the biological sample includes cell-free DNA derived from normal cells and cell-free DNA that may be derived from cancer-related cells, and the method includes (1) a step of analyzing a plurality of cell-free DNA molecules from the biological sample, wherein analyzing the plurality of cell-free DNA molecules includes determining the positions of the cell-free DNA molecules in the genome of the organism, determining whether the cell-free DNA molecules are methylated at one or more sites, wherein each of the plurality of cell-free DNA molecules provides a plurality of sites, and determining whether the cell-free DNA molecules are methylated at the one or more sites includes performing methylation recognition sequence determination, and the methylation recognition sequence determination includes (i) contacting the plurality of cell-free DNA molecules with hybridization probes to enrich cell-free DNA molecules derived from a specific genomic region, wherein the specific genomic region includes genomic regions having methylation differences between cancer and non-cancer for a tissue, (ii) determining the sequences of the cell-free DNA molecules, including; (2) for each site of the plurality of sites, a step of individually calculating the number of each cell-free DNA molecule at the highly methylated sites; and (3) a step of calculating a first methylation level using the number of each cell-free DNA molecule highly methylated at the plurality of sites including, the method.

2. The method according to claim 1, wherein performing methylation recognition sequence determination includes determining the sequences of at least 60,000 cell-free DNA molecules.

3. The method according to claim 1, wherein performing methylation recognition sequence determination further includes amplifying the cell-free DNA molecules.

4. Performing methylation recognition sequence determination includes a step of treating the cell-free DNA molecules with sodium bisulfite; and; a step of determining the sequences of the treated cell-free DNA molecules including, the method according to claim 1.

5. The method according to claim 4, wherein treating the cell-free DNA molecules with sodium bisulfite is part of Tet-assisted bisulfite sequencing or oxidative bisulfite sequencing for the detection of 5-hydroxymethylcytosine.

6. The method according to claim 1, further including determining a first classification of the cancer level based on the first methylation level.

7. Determining a first classification of the level of cancer based on the first methylation level, comprising: comparing the first methylation level with a first cut-off value; and determining a first classification of the level of the cancer based on the comparison The method according to claim 6.

8. Wherein the first classification indicates that cancer is present in the organism, and the method further comprises identifying the type of cancer associated with the organism by comparing the first methylation level with a corresponding value determined from other organisms, wherein at least two of the other organisms are identified as having different types of cancer. The method according to claim 7.

9. The method according to claim 7, wherein the first cut-off value is a specific distance from a reference methylation level established from a biological sample obtained from a healthy organism.

10. The method according to claim 9, wherein the specific distance is a number of specific standard deviations from the reference methylation level.

11. The method according to claim 7, wherein the cut-off value is established from a reference methylation level determined from a previous biological sample of the organism obtained before the biological sample is examined.

12. Comparing the first methylation level with the first cut-off value, comprising: determining the difference between the first methylation level and the reference methylation level; and comparing the difference with a threshold value corresponding to the first cut-off value The method according to claim 7.

13. Determining the concentration fraction of tumor DNA in a biological sample; calculating a first cut-off value based on the concentration fraction of tumor DNA in the biological sample The method according to claim 7, further comprising.

14. Measuring the sizes of cell-free DNA molecules at a plurality of sites and obtaining the sizes measured thereby; normalizing the first methylation level using cell-free DNA molecules having a first size before comparing the first methylation level with the first cut-off value The method according to claim 7, further comprising.

15. The method according to claim 14, wherein the first size is a length within a certain range.

16. The method according to claim 14, wherein cell-free DNA molecules having a first size are selected based on size-dependent physical separation. Step of performing paired-end sequencing of the plurality of cell-free DNA molecules to obtain a pair of sequences for each of the cell-free DNA molecules; Step of determining the size of the cell-free DNA molecules by comparing the pair of sequences with a reference genome; and Step of selecting cell-free DNA molecules having a first size The method according to claim 14, further comprising selecting cell-free DNA molecules having a first size by the above steps.

18. The step of normalizing a first methylation level using the cell-free DNA molecules having a first size includes Step of obtaining a functional relationship between size and methylation level; and Step of normalizing the first methylation level using the functional relationship The method according to claim 14, wherein the functional relationship gives a scaling value corresponding to each size.

19. Step of calculating, by a computer, an average size corresponding to the used cell-free DNA molecules to calculate a first methylation level; and Step of multiplying the first methylation level by a corresponding scaling value The method according to claim 18, further comprising the above steps.

20. For each of the plurality of sites, For each of the cell-free DNA molecules located at the site, Step of obtaining the size of each of the cell-free DNA molecules at the site; and Step of normalizing the contribution of each of the cell-free DNA molecules to the number of cell-free DNA molecules highly methylated at the site using a scaling value corresponding to each size The method according to claim 18, further comprising the above steps.

21. The method according to claim 7, wherein the plurality of sites include CpG sites, the CpG sites are integrated into a plurality of CpG islands, each CpG island includes two or more CpG sites, and the first methylation level corresponds to a first CpG island.

22. For each of the plurality of CpG islands, Step of determining whether the CpG island is highly methylated relative to a reference group of samples of other organisms by comparing the methylation level of the CpG island with respective cutoff values, thereby determining highly methylated CpG islands; Step of determining the methylation density of each of the highly methylated CpG islands; Step of calculating a cumulative score from each of the methylation densities; and Step of comparing the cumulative score with a cumulative cutoff value to determine a first classification The method according to claim 21, further comprising the above steps. A step of determining whether the concentration fraction of tumor DNA in the biological sample is greater than a minimum value; and a step of flagging the biological sample if the concentration fraction of the tumor DNA is not greater than the minimum value The method according to claim 1, further comprising. The method according to claim 23, wherein the minimum value is determined based on an expected difference in the methylation level of the tumor relative to a reference methylation level. The method according to claim 1, wherein the plurality of sites are present on a plurality of chromosomes. The method according to claim 1, wherein the plurality of sites are from regions separated from each other. The method according to claim 1, wherein the methylation recognition sequence determination further comprises contacting cell-free DNA molecules with a protein that binds to methylated DNA. A method for analyzing a biological sample of an organism, comprising: the biological sample includes nucleic acid molecules derived from normal cells and nucleic acid molecules that may be derived from cancer-related cells, and the method includes: (1) A step of analyzing a plurality of cell-free DNA molecules from the biological sample, wherein analyzing each of the plurality of cell-free DNA molecules includes: determining the position of the cell-free DNA molecule in the genome of the organism; determining whether the cell-free DNA molecule is methylated at one or more sites, wherein the one or more sites of the plurality of cell-free DNA molecules provide a plurality of sites, and determining whether the cell-free DNA molecule is methylated at one or more sites includes performing methylation recognition sequence determination, and the methylation recognition sequence determination includes: (i) treating the plurality of cell-free DNA molecules with sodium bisulfite; (ii) contacting the plurality of cell-free DNA molecules with a hybridization probe to enrich cell-free DNA molecules derived from a specific genomic region, wherein the specific genomic region includes a genomic region having a methylation difference between cancer and non-cancer for a tissue; and (iii) determining the sequence of the cell-free DNA molecule including, (2) For each of the plurality of sites, a step of individually calculating the number of each cell-free DNA molecule at a highly methylated site; and (3) A step of calculating a first methylation level using the number of each cell-free DNA molecule highly methylated at the plurality of sites including, method. **Claim 29** A method for analyzing a biological sample of an organism, comprising: the biological sample includes a nucleic acid molecule derived from a normal cell and a nucleic acid molecule that may be derived from a cancer-related cell, and the method comprises: (1) a step of analyzing a plurality of cell-free DNA molecules from the biological sample, wherein analyzing each of the plurality of cell-free DNA molecules comprises: determining the position of the cell-free DNA molecule in the genome of the organism; determining whether the cell-free DNA molecule is methylated at one or more sites, wherein one or more sites of each of the plurality of cell-free DNA molecules provide a plurality of sites, and determining whether the cell-free DNA molecule is methylated at one or more sites comprises performing methylation recognition sequence determination, and the methylation recognition sequence determination comprises: (i) contacting the plurality of cell-free DNA molecules with a protein that binds to methylated DNA; (ii) contacting the plurality of cell-free DNA molecules with a hybridization probe to enrich cell-free DNA molecules derived from a specific genomic region, wherein the specific genomic region includes a genomic region having a methylation difference between cancer and non-cancer for a tissue; and (iii) determining the sequence of the cell-free DNA molecule comprising; (2) for each of the plurality of sites, a step of individually calculating the number of each of the cell-free DNA molecules at the highly methylated sites; and (3) a step of calculating a first methylation level using the number of each of the cell-free DNA molecules highly methylated at the plurality of sites A method comprising. **Claim 30** A computer-readable medium storing a plurality of instructions for controlling a computer device to execute the method according to any one of claims 1 to 29. **Claim 31** A computer device comprising means for executing the method according to any one of claims 1 to 29, and further comprising at least one processor and a memory.