Method of clinical assessment of disease using genomic methylation patterns
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MEDICOVER GENETICS LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-08-06
Smart Images

Figure 00000052_0000 
Figure 00000053_0000 
Figure 00000054_0000
Abstract
Description
[0001] Medicover CH Kilger Anwaltspartnerschaft mbB Cyprus FasanenstraRe 29 Our Ref.: B281-0058W01 10719 Berlin
[0002] METHOD OF CLINICAL ASSESSMENT OF DISEASE USING GENOMIC METHYLATION PATTERNS
[0003] FIELD OF THE INVENTION
[0004] The invention is in the field of biology, medicine, and chemistry, in particular the field of molecular biology and more in particular in the field of molecular diagnostics. The invention is in particular in the field of in-vitro diagnostics using cell-free nucleic acids and DNA and methylation patterns therein. The invention is also in the field of diagnostic kits.
[0005] BACKGROUND OF THE INVENTION
[0006] Cancer has been marked as one of the leading causes of death with metastatic cancer being responsible for over 90% of cancer deaths. Characterized by an abnormal growth of cells, cancer can invade and destroy other types of tissues in all parts of the body. The discovery of circulating-tumor DNA (ctDNA) is considered an emerging biomarker for many cancers, making it a crucial biomarker in the diagnosis of cancer, as ctDNA analysis offers an attractive non-invasive means for cancer detection and monitoring.
[0007] Cell-free DNA in plasma consists of a mixture of fragmented DNA molecules released from various tissues within the body. Fragmentation of cell-free DNA (cfDNA) is a non-random process. Each cfDNA fragment bears molecular signatures of its cell or tissue of origin, signatures which once deconvoluted are found to primarily comprise of DNA methylation patterns, nucleosome footprints, topological differences and sequence motifs.
[0008] Every single cell of a multicellular organism carries the same genetic material encrypted in their DNA sequence. The diversity of their morphology and functions caused by differential gene expression is mainly controlled by epigenetic modifications. DNA methylation is the chemical modification of cytosines covalently linked to a methyl group, preferentially where a cytosine is followed by a guanine (CpG site) in mammalian DNA. The human genome contains about 0.6 billion cytosines and 32 million CpG sites. This epigenetic modification can have an impact on a range of biological processes involving gene expression such as homeostasis, embryonic development, cancer and aging. Being an important regulator of gene transcription, DNA methylation has been generally associated with transcriptionalsilencing and chromatin organization leading to various diseases. Abnormal DNA methylation has been identified as one of the processes leading to tumor onset, development, progression and relapse, since tumor cells have been characterized by a different methylome from that of normal cells. Therefore, methylated DNA has been studied as a potential biomarker in the tissues of most tumor types. Aberrant methylation markers provide increased specificity and sensitivity and could be more informative compared to DNA mutations since pathophysiological methylation variations are much more frequent than single nucleotide variations (SNVs). In spite of that, the use of DNA methylation detection methodologies is hindered by the required high DNA starting quantities, thus making its utilization challenging.
[0009] Further studies of high frequency CpG site regions, also known as CpG islands, have brought to light important findings when applied to either animal models or human cell lines. Different parts of the same CpG island may have different levels of methylation while methylation levels may be distributed bimodally between highly methylated and unmethylated domains. This phenomenon supports the concept of binary switch-like pattern of DNA methyltransferase enzymes involved in the DNA methylation process.
[0010] Hypomethylation, defined by a lower methylation level in CpGs, has been considered in physiopathology as a marker of genomic instability and often has the responsibility of activation of oncogenes in cancer. Normally, pericentromeric heterochromatin is highly methylated which ensures any satellite or repetitive genomic sequences to remain silenced therefore maintaining the integrity and stability of the genome. The loss of DNA methylation levels in these normally deactivated regions can result in the activation of transposable elements that can integrate at random genomic sites and lead to mutagenesis and genomic instability. Abnormal hypomethylation can also reactivate suppressed viral sequences that are integrated in the genome which can initiate tumor progression. Hypomethylation in gene bodies can also lead to the overexpression of certain genes, including those involved in cell proliferation or survival, contributing to cancer development. DNA hypomethylation has also been linked with anti-cancer drug resistance since promoters of specific resistance genes have been found hypomethylated, thus transcriptionally active.
[0011] Hypermethylation on the other hand, defined as the increased methylation of CpG sites, can alter the normal activity of genes. Such alterations can negatively impact the mechanisms of transcription, double-strand break repair and recombination as well as DNA mismatch repair promoting tumorigenesis, invasion, metastasis, tumor progression and cancer-cell survival.
[0012] Hypermethylation refers to an increase in DNA methylation levels, usually in the promoter regions of genes. Promoter hypermethylation can lead to the silencing of tumor suppressor genes, which areresponsible for controlling cell growth and preventing the development of cancer. When tumor suppressor genes are silenced by hypermethylation, it can contribute to the uncontrolled growth of cancer cells.
[0013] It is important to note that the specific DNA methylation changes observed in cancer can vary depending on the type of cancer and individual genetic or environmental factors. Additionally, other molecular alterations, such as mutations or chromosomal abnormalities, also contribute to cancer development.
[0014] Aberrant DNA methylation patterns are commonly observed in cancer cells. Methylation-specific PCR (MSP) and other methylation-specific detection methods can be used to identify specific DNA methylation changes associated with particular types of cancer. These methylation markers can serve as diagnostic tools for early detection and as indicators of disease progression.
[0015] DNA methylation profiles can distinguish between different subtypes of cancer. For example, in some types of leukemia, DNA methylation analysis can help identify specific subtypes with different clinical behaviors and treatment responses. This information can guide personalized treatment strategies and improve patient outcomes.
[0016] DNA methylation patterns can be detected in various bodily fluids, such as blood, urine, sweat, tears, sputum, cerebrospinal fluid, stool sample, ascites fluid, pleural effusion fluid, saliva, bronchoalveolar lavage fluid and / or solid tissues. This characteristic allows for the development of non-invasive diagnostic tests, such as liquid biopsies. By analyzing DNA methylation markers present in these fluids, it is possible to detect and monitor certain diseases including cancer, without the need for invasive procedures.
[0017] DNA methylation status can provide valuable prognostic and predictive information. By examining specific methylation patterns, clinicians can assess the likelihood of disease detection, recurrence, disease progression, or treatment response. This information helps guide treatment decisions and improve patient management.
[0018] One of the most common approaches for the detection of cytosine methylation, more specifically 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC), as well as further assessment of the methylation status of a DNA sample, has been bisulfite sequencing. Bisulfite sequencing maps the methylome by chemically modifying the DNA through converting unmethylated cytosines to uracil while the methylated forms of 5-methylcytosine or 5-hydroxymethylcytosine remain intact, which can be thereafter identified through sequencing. However, despite its wide use, the limitations of bisulfite sequencing such as its lack of specificity, the generation of damage in the DNA resulting in DNA loss orfragmentation and therefore biased sequencing data, have led to the need for a new approach which would overcome such constraints.
[0019] To distinguish between 5mC and 5hmC, oxidative bisulfite sequencing (oxBS-seq) has been developed. OxBS-seq discriminates between 5mC and 5hmC via a highly selective chemical oxidation of 5hmC to 5-formylcytosine (5fC). Specific oxidation of 5hmC to 5fC is achieved with potassium perruthenate (KRuO4). The approach also exploits the difference in reactivity between 5hmC and 5fC towards bisulfite. Bisulfite treatment causes 5fC to be deformylated and deaminated to form uracil (which is read as thymine at the sequencing stage) in contrast to 5hmC, which is not deaminated. Thus, the only base that is not deaminated and therefore read as a cytosine after oxBS-seq is 5mC. Consequently, this approach gives an accurate readout of all the 5mC present enabling quantitative, single nucleotide resolution sequencing on widely available platforms. However, one of the main limitations of the method is the significant loss of DNA due to the severe oxidation conditions. Simultaneously, the low levels of 5hmC in the genome require high sequencing coverage and 5mC subtraction to obtain the accurate levels of 5hmC, increasing the noise level and enhancing the need to increase sequencing coverage and number of replicates, hence cost.
[0020] Similarly, Tet-assisted bisulfite sequencing (TAB-seq) employs a clever use of the ten-eleven translocation (TET) enzyme to distinguish between 5mC and 5hmC using bisulfite. The Tet enzymes are responsible for stepwise oxidative demethylation of 5mC to 5-carboxylcytosine (5caC) without affecting 5hmC. In TAB-seq, 5hmC is protected by glycosylation. The DNA is then treated with bisulfite, converting 5caC to uracil, while leaving the glycosylated 5hmC untouched. Any cytosines read in the resulting sequence are thus interpreted as 5hmC. The biggest limitation of the method, besides the DNA damage resulting from the bisulfite sequencing step, lies in that the sensitive and specific detection of 5hmC by TAB-seq depends on the sequencing depth, which means that samples with less abundant 5hmC modifications will require more sequencing depth.
[0021] DNA damage resulting from bisulfite treatment led to the expansion of bisulfite-free methods to detect methylated cytosines at single-base resolution, such as TET-assisted pyridine borane sequencing (TAPS). TAPS combines TET-mediated oxidation of 5mC and 5hmC to 5-carboxylcytosine (5caC) with pyridine borane reduction of 5caC to dihydrouracil (DHU). Subsequent PCR converts DHU to thymine, enabling a cytosine-to-thymine transition of 5mC and 5hmC. TAPS detects modifications directly with high sensitivity and specificity, without affecting unmodified cytosines. This method is nondestructive and detects both 5mC and 5hmC. TAPS uses mild chemistry to detect DNA methylation directly and shows improved sequence quality, mapping rate, and coverage compared to bisulfite sequencing, while reducing sequencing cost by half. The combination of direct methylation detection and the nondestructive nature of TAPS makes it ideal for simultaneous genetic analysis in cfDNA, which couldenhance noninvasive cancer detection by liquid biopsies. TAPS comes in two other varieties which surpass the standard version's limitation of detection of both methylated and hydroxymethylated cytosines: TAPSP: p-glucosyltransferase labels 5hmC with glucose, protects 5hmC from the oxidation and reduction reactions and allows for specific detection of 5mC; Chemical-assisted pyridine borane sequencing (CAPS): Potassium perruthenate acts as a chemical replacement for Tetl and specifically oxidizes 5hmC, thus allowing for direct detection. However, CAPS involves multiple steps and treatments, possibly leading to loss of material.
[0022] Another bisulfite-free method for profiling 5mC at single-base resolution using nanogram quantities of DNA has been dubbed Direct Methylation sequencing (DM-Seq). DM-Seq employs two key DNA-modifying enzymes: a neomorphic DNA methyltransferase and a DNA deaminase capable of precise discrimination between cytosine modification states. Coupling these activities with deaminase-resistant adapters enables accurate detection of only 5mC via a cytosine-to-thymine transition in sequencing. The use of a new CpG-specific carboxymethyltransferase -called M.MpelN374K- with an E. coli metabolite -called carboxy-S-adenosyl-L-methionine (CxSAM)- to protect unmodified cytosines on newly generated hemi-methylated strands, showed higher yield than BS-Seq. This was achieved by adding custom, deaminase-resistant adapters with 5-propynylcytosine (5pyC) instead of the traditional 5mC version then copying DNA with these adapters resulting in a methylated copy strand and hence hemi-methylation without confounding 5mC and 5hmC at clinically important sites. Nevertheless, the DM-Seq method employs a significant experimental complexity while the need for hemi-methylated strands requires strand duplication which can further lead to increased error rates.
[0023] The most recent technological development in the field has been a single base-resolution sequencing methodology that sequences complete genetics and the two most common cytosine modifications in a single workflow. DNA is copied and bases are enzymatically converted. Paired decoding of bases across the original and copy strand provides a phased digital readout. This 5-letter sequencing approach is accurate, requires low DNA input and has a relatively simple workflow and analysis pipeline. Simultaneous phased reading of genetic and epigenetic bases provides a more complete picture of the information stored in genomes. Hairpins are ligated to double-stranded DNA and the strands are separated. An additional copy strand is synthesized using Klenow exo- polymerase and short sequencing adapters are ligated. Modified cytosines are protected through oxidation by TET2 and glycosylation by beta-glucosyltransferase (BGT). Treatment by APOBEC3A and UvrD helicase is used to simultaneously open and deaminate the hairpin. Unprotected cytosines are deaminated from cytosines to uracils (read as thymines) and by using a set of resolution rules, the pairs of bases across the two strands are resolved into one of five states: adenine, guanine, thymine, cytosine and modified cytosine (5mC & 5hmC).The evolution of this approach allows disambiguation of 5mC from 5hmC without compromising genetic base calling within the same sample fragment. The first three steps of the workflow are identical to five-letter seq to generate the adapter ligated sample fragment with the synthetic copy strand. To enable six-letter seq, methylation at 5mC is enzymatically copied across the CpG unit to the cytosine on the copy strand using DNA methyltransferase 5 (DNMT5), whereas 5hmC is glycosylated via beta-glucosyltransferase to prevent such a copy. DNMT5 has been selected to perform the copy step because of its specificity for copy methylation over de novo methylation. The un-modified cytosines are then deaminated to uracil, which is subsequently read as thymine via the joint action of TET2, APOBEC3A and UvrD helicase. The DNA is subjected to PCR amplification and sequencing. Each of unmodified cytosines, 5mC and 5hmC can be resolved as the three CpG units have distinct sequencing readouts of the two-base code. The method is highly accurate, performs with very limited sample amounts, and is compatible with commercially available sequencing platforms. Nonetheless, the high cost and the experimental complexity, due to the multiple steps involved in the process, are two of the main limitations for both methods, 5-letter and 6-letter sequencing.
[0024] More recently, APOBEC3A has been shown to efficiently deaminate methylated, but not TET-oxidized, cytosine bases in DNA allowing for a non-destructive, enzymatic deamination of cytosine modifications. APOBEC-Coupled Epigenetic Sequencing (ACE-Seq) permits cytosine, 5mC and 5hmC to be parsed and reveals genomic features with a need for about 1000 times less DNA than TAB-seq.
[0025] Along those lines, Enzymatic Methyl-sequencing (EM-seq) is an enzymatic-based approach that has been found to outperform bisulfite sequencing in aspects of coverage, duplication, sensitivity, nucleotide composition as well as GC distribution, improved correlation across input amounts and increased representation of genomic features. Specifically, EM-seq effectively converts unmethylated cytosines into uracil while maintaining and protecting DNA integrity by using only the enzymatic reactions of oxidation, glycosylation and deamination to map 5-methylcytosine and 5-hydroxymethylcytosine (Vaisvila R. et al).
[0026] A person skilled in the art could use any of the above-mentioned methods for the detection, at a baselevel, of modified cytosines in combination with the invention disclosed further below. Numerous other episignatures, i.e. functionally relevant changes to the genome that do not involve a change in the nucleotide sequence, can be present in DNA or RNA which can be specifically targeted for insolution hybridization using the method disclosed herein. A person skilled in the art would appreciate that different types of conversion could be applied to specifically target different episignatures originating either from DNA or RNA molecules.Next Generation Sequencing (NGS) technologies have been applied in the development of non-invasive molecular testing by using next generation massively parallel shotgun sequencing (MPSS), a whole genome-based approach. Furthermore, targeted-based NGS methods have been developed to sequence only particular sequences of interest. Such targeted approaches substantially reduce the cost since sequencing is only performed on the target sequences of interest. However, the enrichment performance resulting from the use of conventional probes to capture chemically- or enzymatically converted libraries is usually suboptimal because converted DNA is not homologous anymore to the same non-converted regions. This problematic part can be circumvented by converting the library after enrichment by hybridization. However, PCR-free hybridization is needed for this approach, which greatly minimizes the sample's diversity when using low starting amounts of DNA such as the ones found for cfDNA in plasma. The method disclosed further below circumvents both above mentioned limitations by using a mixture of specifically designed converted and non-converted TAC oligonucleotides, allowing an optimal enrichment of regions of interest for methylation status assessment leading to improved sensitivity and specificity detection, particularly for applications using minute amounts of starting DNA requiring to be amplified before enrichment.
[0027] The disclosed method of the invention delivers improved sensitivity and specificity achieving efficient enrichment of converted and amplified DNA originating from low amounts of material while preserving accurate methylation levels.
[0028] This invention provides a new method allowing for diagnosis of cancer and disease, cancer and disease risk prediction based on methylation risk scores, cancer and disease prognosis, accurate cancer and disease sub-classification, minimal residual disease (MRD) detection, drug response prediction, treatment planning, therapy selection and treatment efficacy monitoring in a subject.
[0029] SUMMARY OF THE INVENTION
[0030] The invention relates to a method of disease diagnosis, disease risk prediction based on methylation risk scores, disease prognosis, accurate disease sub-classification, cancer status determination, MRD detection, drug response prediction, treatment planning, therapy selection, treatment efficacy monitoring and patient stratification for clinical trial enrolment by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:
[0031] a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;b. converting the tested sample's DNA molecules;
[0032] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0033] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0034] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0035] f. sequencing the enriched library;
[0036] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0037] h. determining cancer and disease risk prediction, cancer and disease diagnosis, cancer and disease prognosis, disease sub-classification, drug response prediction, MRD status, treatment planning, therapy selection and / or treatment efficacy of the tested sample by utilizing the information from step (g).
[0038] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0039] In various embodiments, the vertebrate sample is a human DNA sample and is selected from a group comprising of a peripheral blood sample, a plasma sample, a urine sample, a sputum sample, a cerebrospinal fluid sample, a stool sample, an ascites sample, a tear sample, a sweat sample, a saliva sample, a pleural effusion sample, a bronchoalveolar lavage sample, a buffy coat sample and / or a solid tissue sample.
[0040] In accordance with the present disclosure, cell-free DNA fragments from a test subject from which the sample stems from are treated to convert unmethylated cytosines to uracils subsequently replaced by thymine during the amplification, sequenced then compared to a reference genome to identify the status of methylation at one or more sites of the fragments. The identification of anomalously methylated cfDNA fragments can be a challenging task. The classification of one or more cfDNA fragments as anomalously methylated can only prevail when compared to control subjects with physiological methylation levels. Furthermore, the state of methylation can vary among a group of control subjects which makes the identification of a subject as anomalously methylated more difficult due to inter-individual variability (Gross et al).In some aspects, the present invention provides a method of analyzing a biological sample, including a mixture of cell-free nucleic acid molecules, to determine the level of methylation in a subject from which the biological sample is obtained, the mixture potentially including methylated nucleic acid molecules, unmethylated nucleic acid molecules and / or a partially methylated nucleotide acid molecules.
[0041] In some embodiments, the current method comprises performing targeted methylation sequencing on the cell-free nucleic acid molecules to generate information on the methylation status of cytosines of interest. In some embodiments, the sample's nucleic acids, hence the cytosines, are present in very small amounts.
[0042] The methylation status is considered as the methylated fraction or percentage at a particular site of interest e.g., at a single nucleotide, at a specific locus or at a longer sequence of interest, in proportion to the total amount of DNA fragments in the sample comprising that specific site (Ahlquist D.A. et al).
[0043] In some embodiments, a determination of increased methylation in one or more DNA fragments or cytosines comprises a detection of higher methylation level within a region selected from the group consisting of a CpG island and a CpG island shore, the long region which lies on both sides of a CpG island, or CpGs in gene bodies (Ahlquist D.A. et al).
[0044] In some aspects, the disclosed method of the invention leverages the benefit of optimal TAC oligonucleotides designed to efficiently capture the converted and amplified sequencing library from samples having very low amount of DNA while preserving accurate episignatures.
[0045] In some embodiments, the optimal design of TAC oligonucleotides having specific features and their production combined with their use as a mixture of converted and non-converted with various proportions during in-solution hybridization, led to improved sensitivity and specificity performance, especially with samples having very little amount of DNA. Said type of samples are particularly challenging to analyze and therefore suffer from increased bias when state of the art methods are employed. The disclosed method also allows for flexibility and cost-effectiveness to analyze challenging samples without compromising sensitivity and specificity.
[0046] DETAILED DESCRIPTION OF THE INVENTION
[0047] Cancer & Disease DiagnosisThe invention relates to a method of diagnosing cancer and diseases, by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:
[0048] a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;
[0049] b. converting the tested sample's DNA molecules;
[0050] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0051] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0052] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0053] f. sequencing the enriched library;
[0054] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0055] h. determining cancer and disease diagnostic status of the tested sample by utilizing the information from step (g).
[0056] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0057] In one embodiment, the CpG methylation status pattern on each sequenced fragment from a plurality of sequenced fragments is computed in order to determine the tissue of origin. Said sequenced fragments are defined as pairs of sequenced reads, with said CpG methylation pattern being defined by the binary status (0: unmethylated; 1: methylated) of each CpG on said fragment, i.e. a fragment spanning four methylated CpGs in a row and two unmethylated CpGs in a row would have a CpG methylation pattern defined by 111100.
[0058] A wide variety of DNA methylation biomarkers have been identified and employed in cancer diagnosis, being important informants of the course and probable outcome of the disease. Examples of methylation biomarkers used in cancer diagnosis comprise:NDRG4: The N-Myc Downstream-Regulated Gene 4, also known as SMAP-8 and BDM1, is one of the four members of the NDRG gene family, a group of genes involved in cell proliferation, differentiation, development and stress. NDRG4 is specifically expressed by enteric neurons while promoter methylation of N-myc Downstream-Regulated Gene 4 (NDRG4) in fecal DNA is an established early detection marker for colorectal cancer (CRC). Also, NDRG4 is one of the molecular markers of an FDA-approved, multi-target stool DNA test which detects significantly more cancers than the leading immunochemical test.
[0059] BMP3: Bone morphogenetic protein 3, also known as osteogenin, is a member of the transforming growth factor beta superfamily and has the ability to inhibit the ability of other BMPs to induce bone and cartilage development and negatively regulate bone density. High frequencies of methylated BMP3 biomarkers have been found in several cases of colorectal cancer patients. Simultaneously, hypermethylated BP3 in stool samples established the protein as a prominent early diagnosis biomarker in lesions with neoplastic CRC potential.
[0060] SHOX2: Short-stature homeobox 2, also known as homeobox protein Ogl2X or paired-related homeobox protein SHOT, is a member of the homeobox family of genes that encode proteins containing DNA-binding domain. Hypermethylation of the SHOX2 gene has been closely associated with lung carcinogenesis, therefore determining the gene as a sensitive and specific biomarker for identifying subjects with Non-Small Cell Lung Carcinoma (NSCLC) and is currently used to aid in the diagnosis of lung cancer in patients at increased risk for the disease.
[0061] PTGER4: Prostaglandin E Receptor 4 is one of the four identified EP receptors that bind with and mediate cellular responses to PGE2 and certain other prostanoids, but with less affinity and responsiveness. Said receptor can activate T-cell factor signaling and has been shown to mediate PGE2 induced expression of early growth response-1 (EGR1), regulate the level and stability of cyclogenase-2 mRNA, and lead to the phosphorylation of glycogen synthase kinase-3. In its hypermethylated form, PTGER4 has been used along with other receptors to distinguish patients with lung cancer from patients with benign lesions and patients at high risk irrelevant of tumor volume, smoking status and age.
[0062] Cancer & Disease Risk Prediction
[0063] The invention relates to a method of cancer and disease risk prediction, by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;
[0064] b. converting the tested sample's DNA molecules;
[0065] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0066] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0067] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0068] f. sequencing the enriched library;
[0069] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0070] h. determining the cancer and disease risk status of the tested sample by utilizing the information from step (g).
[0071] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0072] CpG islands are regions of DNA that contain a high density of CpG dinucleotides. Aberrant hypermethylation of CpG islands in promoter regions can lead to gene silencing and is frequently observed in cancer. Hypermethylation of tumor suppressor genes, such as CDKN2A, APC, BRCA1, and MLH1, has been implicated in various cancer types. Promoter hypermethylation of the 06-methylguanine-DNA methyltransferase (MGMT) gene, a DNA repair enzyme that removes alkyl groups from DNA, is associated with decreased expression of the enzyme making tumor cells more susceptible to DNA damage. MGMT promoter methylation has been studied extensively in glioblastoma and is used as a predictive biomarker for response to certain chemotherapy agents. Promoter hypermethylation of Glutathione S-transferase Pl (GSTP1), an enzyme involved in detoxification processes, has been observed in prostate cancer and is associated with the loss of its enzymatic activity. GSTP1 promoter methylation is being investigated as a potential diagnostic and prognostic biomarker for prostate cancer. Promoter hypermethylation of Septin 9 (SEPT9), a gene involved in cell division and cytoskeleton organization, has been found in colorectal cancer and is being used as a non-invasive biomarker for the early detection of colorectal cancer through blood-based testing. Rasassociation domain family member 1A (RASSF1A) is a tumor suppressor gene involved in cell cycle control and apoptosis and its hypermethylated promoter has been observed in various cancers, including lung, breast, and bladder cancer. These and other markers may be analyzed with the present method and used for cancer risk prediction or cancer onset.
[0073] Cancer & Disease Prognosis
[0074] The invention relates to a method of cancer and disease prognosis, by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:
[0075] a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;
[0076] b. converting the tested sample's DNA molecules;
[0077] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0078] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0079] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0080] f. sequencing the enriched library;
[0081] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0082] h. determining a cancer and a disease prognostic status of the tested sample by utilizing the information from step (g).
[0083] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.Several DNA methylation biomarkers have been identified and utilized for cancer prognosis, providing valuable information about the likely course and outcome of the disease. Examples of methylation biomarkers commonly used in cancer prognosis comprise:
[0084] BRCA1 promoter methylation: Methylation of the BRCA1 gene promoter has been observed in breast and ovarian cancers. BRCA1 is a tumor suppressor gene involved in DNA repair mechanisms. Methylation-induced silencing of BRCA1 can affect response to therapy and prognosis, especially in patients with breast cancer.
[0085] APC promoter methylation: Adenomatous polyposis coli (APC) is a tumor suppressor gene commonly associated with colorectal cancer. Methylation of the APC promoter region has been found in colorectal cancer and can serve as a prognostic indicator for disease progression and survival.
[0086] CDKN2A(pl6INK4a) promoter methylation: Methylation of the CDKN2A gene, which encodes the pl6INK4aprotein, is associated with cell cycle regulation. Promoter methylation of CDKN2A has been identified in various cancers, including melanoma, lung, and pancreatic cancer. It has been linked to worse prognosis and shorter survival in some cases.
[0087] RASSF1A promoter methylation: Methylation of the RASSFlA gene promoter has been investigated as a prognostic biomarker in several cancer types, including lung, breast, bladder, and gastric cancer. RASSF1A is involved in cell cycle regulation and apoptosis. Its methylation status has been associated with tumor aggressiveness and poor patient outcomes.
[0088] Drug Response Prediction
[0089] The invention relates to a method of drug response prediction, by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:
[0090] a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;
[0091] b. converting the tested sample's DNA molecules;
[0092] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0093] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAColigonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0094] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0095] f. sequencing the enriched library;
[0096] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0097] h. determining a drug response prediction of the tested sample by utilizing the information from step (g).
[0098] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0099] The following examples are specified herein:
[0100] MGMT promoter methylation: O6-methylguanine-DNA methyltransferase (MGMT) is a DNA repair enzyme. Methylation of the MGMT promoter is associated with decreased expression of MGMT and has been extensively studied in glioblastoma. It serves as a predictive biomarker in determining the response to alkylating chemotherapy agents like temozolomide.
[0101] BRCA1 promoter methylation: Methylation of the BRCA1 gene promoter has been observed in breast and ovarian cancers. BRCA1 is a tumor suppressor gene involved in DNA repair mechanisms. BRCA1 promoter methylation status in breast cancer can predict clinical outcome following therapy with certain chemotherapy agents and targeted therapies such platinum-derived therapy and PARP inhibitors.
[0102] BDNF methylation: brain-derived neurotrophic factor is a member of the neurotrophin family of growth factors and helps to support survival of existing neurons, encourages growth and differentiation of new neurons and synapses by activating certain neurons TrkB tyrosine kinase receptors. BDNF DNA hypomethylation has been shown to impair antidepressant treatment response.
[0103] CREBBP methylation: cAMP response element-binding (CREB) binding protein is a coactivator that has intrinsic acetyltransferase functions. CREB is able to add acetyl groups to both transcription factors and histone lysines altering chromatin structure thus making genes more accessible for transcription. DNA methylation changes in the CREBBP gene have been found significantly correlated with improved clinical outcomes in treatment-resistant schizophrenia when medicated by clozapine.Pharmacodynamics / Pharmacokinetics
[0104] The invention relates to a method of therapy selection, treatment planning and / or treatment efficacy monitoring, by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:
[0105] a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;
[0106] b. converting the tested sample's DNA molecules;
[0107] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0108] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0109] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0110] f. sequencing the enriched library;
[0111] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0112] h. determining therapy selection, treatment planning and treatment efficacy monitoring of the tested sample by utilizing the information from step (g).
[0113] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0114] Methylation biomarkers can be used for the determination of the most effective treatment plan for a patient. Each patient has a different genetic makeup, metabolizing the chemical structure of drugs differently. Treatment planning, particularly dosing decisions are beneficial to many patients since, in some cases, patients with low metabolism in certain drugs present multiple adverse side effects due to drug accumulation in their body. The tight monitoring of methylation biomarkers affected by administration of drugs that give indications on biochemical and physiological consequences, e.g., therapeutic window, duration of action and undesirable effects; is a powerful tool for clinicians.Clozapine is an FDA-approved atypical antipsychotic medication for treatment-resistant schizophrenia. Clozapine is not the first-line drug of choice due to its range of adverse effects, making compliance an issue for many patients. This is due to its high interindividual differences in plasma concentration at a given dose and the risk of serious adverse drug reactions (ADR) at high concentrations. Therefore, epigenome-wide association studies on clozapine plasma concentrations reporting that few CpG sites could link the methylation status to observed interindividual concentrations may benefit doctors in determining their patients' most effective personalized treatment plan.
[0115] MRD detection and recurrence
[0116] The invention relates to a method of minimal residual disease (MRD) and recurrence detection, by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:
[0117] a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;
[0118] b. converting the tested sample's DNA molecules;
[0119] c. amplifying the tested sample's converted DNA to create a sequencing library;
[0120] d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;
[0121] e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;
[0122] f. sequencing the enriched library;
[0123] g. performing statistical analysis on the enriched library sequences thereby determining the methylation pattern of a specific region or a plurality of genomic regions of interest; and
[0124] h. determining MRD detection of the tested sample by utilizing the information from step (g).
[0125] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.Minimal Residual Disease detection typically refers to the identification of persistent cancer cells in the patient's body indicative of micrometastasis, or small portion of primary tumor remaining in the body after treatment. As a significant contributor to recurrence or metastasis, MRD is relevant for patients with tumors who can be treated surgically or undergo treatment with the intent to cure. MRD detection can predict the molecular relapse of the disease before the manifestation of clinical symptoms or the presence of tumor in radiological imaging allowing for early intervention. The invention allows for accurate and sensitive detection of cancer relapse, and its use as a prognostic tool for patient outcome. Monitoring patient during treatment and surveillance (after the end of treatment) allows for individualization of cancer treatment.
[0126] MRD testing is mainly used in blood cancers (leukemia, lymphoma and myeloma) but is being studied in other cancers.
[0127] Bladder cancer is the 5th most common cancer in the western world, and 70%-80% of the patients are diagnosed with non-muscle invasive disease. Nucleix's Bladder EpiCheck® received FDA clearance for monitoring of Non-Muscle Invasive Bladder Cancer (NMIBC) Recurrence. The Multiplex DNA methylation-based PCR assay analyzes subtle disease-specific changes in DNA methylation markers, allowing for the detection of high-risk (non Ta-LG) cancers.
[0128] The timely identification of residual disease in colorectal cancer is critical for patient survival since it could lead to treatment failure if it is not identified early. Determination of BCAT1 and IKZF1 methylation levels to quantify circulating tumor DNA (ctDNA) can be used to assess cancer recurrence, such as in Colvera™ test developed by Clinical Genomics Pathology Inc.
[0129] In certain embodiments the invention pertains to a method that involves hybridization-based enrichment of selected converted target regions across the vertebrate genome followed by sequencing, coupled with a novel bioinformatics and mathematical analysis pipeline for increased sensitivity and specificity of methylation assessment of numerous cytosines of interest that are present in minute amounts in a DNA sample, wherein the DNA sample used in the method comprises a mixture of methylated and un-methylated cell-free DNA fragments, using a fraction of the cost used in the state-of-art assay.
[0130] The key component allowing for this invention features the use of optimally designed converted TAC oligonucleotides and / or non-converted TAC oligonucleotides and / or amplified converted TAC oligonucleotides and / or amplified non-converted TAC oligonucleotides generated from vertebrate DNA.Herein, the mixture of nucleic acid fragments is preferably isolated from a sample taken from a eukaryotic organism, preferably a vertebrate, preferably a primate, more preferably a human.
[0131] In the context of the present invention, the term "subject" refers to animals, preferably mammals, and, more preferably, humans. Generally, the sample is obtained from the "subject".
[0132] Sensitivity as used in the present disclosure can refer to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify a proportion of the population that truly has a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects within a population having cancer.
[0133] Specificity as used in the present disclosure, can refer to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify a proportion of the population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population not having cancer.
[0134] In the context of the present invention, the term "probe" refers to synthetic TArget Capture (TAC) oligonucleotides. It refers to an oligonucleotide which occurs either naturally as in a purified restriction digest or produced synthetically, recombinantly or by specific PCR amplification and is able to hybridize to another oligonucleotide of interest. Probes or TAC oligonucleotides are useful in the detection, isolation, enrichment and detection of genomic sequences of interest.
[0135] In various embodiments, a pool of TAC oligonucleotides refers to a collection of TAC oligonucleotides comprising at least more than two distinct sequences.
[0136] In one embodiment, the use of families of TAC oligonucleotides within the TAC oligonucleotide pool significantly increases enrichment for the target sequences of interest, as evidenced by a greater than 50% average increase in read-depth for the family of TAC oligonucleotides versus a single TAC oligonucleotide.
[0137] Each TAC oligonucleotide family comprises a plurality of members that bind to the same genomic sequence of interest but have different start and / or stop positions with respect to a reference coordinate system for the genomic sequence of interest. Typically, the reference coordinate system that is used for analyzing human genomic DNA is the human reference genome build hgl9, which is publicly available, although other versions (e.g., hg38) may be used. Alternatively, the reference coordinate system can be an artificially created genome based on build hgl9 that contains only the genomic sequences of interest.Each TAC oligonucleotide family comprises at least 2 members that bind to the same genomic sequence of interest. In various embodiments, each TAC oligonucleotide family comprises at least 2 member sequences, or at least 3 member sequences, or at least 4 member sequences, or at least 5 member sequences, or at least 6 member sequences, or at least 7 member sequences, or at least 8 member sequences, or at least 9 member sequences, or at least 10 member sequences. In various embodiments, the plurality of TAC oligonucleotide families comprises different families having different numbers of member sequences. For example, a pool of TAC oligonucleotides can comprise one TAC oligonucleotide family that comprises 3 member sequences, another TAC oligonucleotide family that comprises 4 member sequences, and yet another TAC oligonucleotide family that comprises 5 member sequences, and the like. The person skilled in the art will appreciate that the TAC oligonucleotide family can comprise more than 5 member sequences. In one embodiment, a TAC oligonucleotide family comprises 3-5 member sequences. Each member of the TAC oligonucleotide family binds to the same genomic sequence of interest but with different start and / or stop positions relative to a reference coordinate system of the genomic sequence of interest.
[0138] In one embodiment, the pool of TAC oligonucleotides comprises a plurality of TAC oligonucleotide families. Thus, a pool of TAC oligonucleotides comprises at least 2 TAC oligonucleotide families. In various embodiments, a pool of TAC oligonucleotides comprises at least 3 different TAC oligonucleotide families, or at least 5 different TAC oligonucleotide families, or at least 10 different TAC oligonucleotide families, or at least 50 different TAC oligonucleotide families, or at least 100 different TAC oligonucleotide families, or at least 500 different TAC oligonucleotide families, or at least 1000 different TAC oligonucleotide families, or at least 2000 TAC oligonucleotide families, or at least 4000 TAC oligonucleotide families, or at least 5000 TAC oligonucleotide families.
[0139] Each member within a family of TAC oligonucleotides binds to the same genomic region of interest but with different start and / or stop positions, with respect to a reference coordinate system for the genomic sequence of interest, such that the binding pattern of the members of the TAC oligonucleotide family is staggered. In various embodiments, the start and / or stop positions are staggered by at least 3 base pairs, or at least 4 base pairs, or at least 5 base pairs, or at least 6 base pairs, or at least 7 base pairs, or at least 8 base pairs, or at least 9 base pairs, or at least 10 base pairs, or at least 15 base pairs, or at least 20 base pairs, or at least 25 base pairs. In a preferred embodiment, the start and / or stop positions are staggered by 5-10 base pairs. In one embodiment, the start and / or stop positions are staggered by 5 base pairs. In another embodiment, the start and / or stop positions are staggered by 10 base pairs.
[0140] In the context of the present invention, the genomic regions of interest include, but are not limited to, converted regions, non-converted regions and partially converted regions.As used herein, "nucleic acid" is used in general to refer to any deoxyribonucleic acid, either modified or unmodified DNA. "Nucleic acids" include, without limitation, single- and double-stranded nucleic acids. The term "nucleic acid" also refers to a DNA molecule that contains one or more bases modified chemically, enzymatically, or metabolically. The term "nucleic acid" also refers to an RNA molecule that contains one or more bases modified chemically, enzymatically, or metabolically.
[0141] In the context of the present invention, the term "nucleic acid fragments" and "fragmented nucleic acids" can be used interchangeably. In various embodiments, the term "fragment" indicates a fragment of a nucleic acid molecule for example, a fragment can refer to a cfDNA or cfRNA molecule in a plasma sample or a cfDNA or cfRNA molecule that has been extracted from a plasma sample, and / or an amplification product of a cfDNA or cfRNA molecule. In another embodiment, a fragment can also refer to a sequence read, or a group of sequence reads required for subsequent analysis.
[0142] As used in the current invention, the term "methylation pattern" refers to the positioning of methylated and unmethylated cytosine sites across a DNA fragment. When compared, two nucleic acids can have a similar or even the same methylation frequency or methylation percentage but a different methylation pattern when the number of methylated and unmethylated nucleotides is the same or at least similar throughout a region but the said methylated and unmethylated cytosines differ in location.
[0143] In various embodiments, the terms "normal", "physiological" or "healthy" are defined as characteristic of or appropriate to an organism's healthy or normal functioning, free of disease or any known pathology. Conversely, the terms "abnormal", "pathological", "pathophysiological", "disease", "sick", "disorder" refer to a condition of the living body or of one of its parts that impairs normal functioning and is typically manifested by distinguishing signs and symptoms.
[0144] In the current invention, the term "abnormal methylation pattern" refers to the methylation pattern of a nucleic acid molecule that is expected to be found in a sample less or more frequently than a threshold value. Methods disclosed herein characterize DNA fragments as abnormal when the said DNA fragment contains unusual methylation patterns in the tested samples.
[0145] The terms "conversion", "converted", "converted DNA", "converted cfDNA", "converted library", "converted regions", "converted strand", "converted probe", "converted TAC oligonucleotide", "converted cytosine" or "converted genome" as used in this invention, refer to DNA molecules that have been modified with the intention of differentiating an epigenetic marker on a nucleic acid from an unmarked nucleic acid. In one embodiment, this is achieved through the chemical, physical or enzymatic modification of the epigenetic marker that can be detected thereafter. In one embodiment, this is achieved through the chemical, physical or enzymatic modification of the unmarked nucleic acidloci that can be detected thereafter. In one embodiment, this is achieved through the chemical or enzymatic modification of deamination of an unmethylated cytosine base into uracil.
[0146] As used in various embodiments, the term "deamination" is defined as the chemical or enzymatic modification of an amino group from a molecule.
[0147] In various embodiments, the term "sequence reads" refers to nucleotide sequence information of a nucleic acid molecule of a sample obtained by a subject. Such sequence reads can be obtained through various methods known in the art including, but not limited to, hybridization array, sequencing and amplification techniques.
[0148] The term "primer" refers to a short oligonucleotide (typically between 18 and 24 bases) with the ability to act as a starting point of synthesis given that is placed under conditions where specific generation of a primer extension product complementary to the primer synthesis can take place. In various embodiments, the primer is single stranded for maximum efficiency in amplification.
[0149] In the context of the present invention, "amplification" refers to the production of multiple copies of a polynucleotide, or a portion of a polynucleotide.
[0150] In one embodiment, the DNA sample is selected from the group consisting of a plasma sample, a urine sample, a sputum sample, a cerebrospinal fluid sample, a sweat sample, a tears sample, an ascites sample, a saliva sample, a pleural effusion sample, a stool sample, a bronchoalveolar lavage sample, a buffy coat and / or a solid tissue sample.
[0151] The method of the invention can be used with a range of biological samples. Essentially any biological sample containing genetic material, e.g., DNA, and specifically cell-free DNA (cfDNA), can be used as a sample in the methods allowing for genetic analysis of the DNA therein. For example, in various embodiments, the DNA sample is a plasma sample containing cfDNA.
[0152] In a preferred embodiment the blood sample is a peripheral blood sample, and the sample can be acquired by standard methods. Alternatively, the blood sample can be a fractionated portion of peripheral blood, such as a plasma sample. As little as 0.5ml to 8ml of plasma is adequate to provide suitable DNA material for analysis according to the disclosed method. Total cell-free DNA can then be extracted from the sample using standard techniques, non-limiting examples of which include a QIAsymphony protocol (QIAGEN) suitable for cell-free DNA isolation or any other manual or automated extraction method suitable for cell-free DNA isolation.
[0153] Provided herein is an invention related to a method of evaluation of the methylation status of a sample (e.g., peripheral blood sample, urine sample, sputum sample, stool sample, saliva sample,cerebrospinal fluid sample, sweat sample, tears sample, bronchoalveolar lavage sample) obtained from a subject (e.g., a human subject), the method comprising comparing the genetic transition levels of cytosine to thymine and guanine to adenosine for each cytosine or guanine in the converted cfDNA fragments of interest to a reference genome and identifying the fragment as hypomethylated or hypermethylated when the methylation proportion of the fragment is less methylated or more methylated, respectively, compared to healthy (non-disease affected) subjects.
[0154] In one embodiment, the method can also be applied for the determination of the combined methylation and hydroxymethylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample.
[0155] In another embodiment, the method can also be applied for the determination of epigenetic status of a single or a plurality of loci in genomic DNA molecules of a vertebrate sample, such as but not limited to 5-Methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC), N6-Methyladenine (6mA) or N4-Methylcytosine (4mC).
[0156] In one embodiment, the method can also be applied for the determination of epigenetic status of a single or a plurality of loci in genomic RNA molecules of a vertebrate sample, such as but not limited to N6-Methyladenine (6mA), pseudouridines, 5-methylcytosine (5mC), Nl-methyladenosine (1mA), inosine or 2'-O-methylation.
[0157] In the context of this invention, the term episignature is defined as a nucleic acid modification such as, but not limited to, 5-Methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC), N6-Methyladenine (6mA) or N4-Methylcytosine (4mC), N6-Methyladenine (6mA), pseudouridines, Nl-methyladenosine (1mA), inosine or 2'-O-methylation.
[0158] In another embodiment, the method can also be applied for the determination of any epigenetic status of a single or a plurality of loci in genomic DNA or RNA molecules of a vertebrate sample that can be detected by modifying the sequence, structure and / or hybridization dynamics.
[0159] In some embodiments, assessing the methylation status of the cfDNA fragment of the sample comprises the methylation status of one base. In some embodiments, assessment of the methylation status of the cfDNA fragment in the sample comprises the degree of methylation at a plurality of bases. Additionally, in some embodiments the methylation status of the cfDNA fragment indicates an increased methylation level of the fragment compared to a normal methylation status of the fragment. In some embodiments, the methylation status of the cfDNA fragment shows a decreased methylation of the fragment compared to a normal methylation status of the fragment. In some embodiments, themethylation status of the cfDNA fragment comprises a different pattern of methylation of the fragment in comparison to a normal methylation pattern of the fragment.
[0160] As used in the current disclosure, "a methylation status" of a cfDNA fragment refers to the presence or absence of one or a plurality of methylated nucleotide bases in the DNA molecule. As an example, a cfDNA fragment containing at least a methylated cytosine could be considered methylated. A cfDNA fragment which does not contain any methylated cytosines is considered unmethylated. The methylation status of a portion of the fragment can provide indication of the methylation status of a part or all other cytosines within the fragment. Simultaneously, a methylation status of a nucleic acid sequence can provide indications of the methylation density within said sequence, often but not always nor limited to, providing the location of methylation sites in said sequence.
[0161] Accordingly, the term "methylation status" can also refer to the multiplicity of methylated or unmethylated cytosines throughout any region of a nucleic acid. If the cytosine residues within a nucleic acid sequence of a subject are methylated, they can be referred to as "hypermethylated" or "having increased methylation" while if the cytosine residues within a nucleic acid sequence are not methylated, they can be referred to as "hypomethylated" or "having decreased methylation". Simultaneously, if the cytosine residues of a nucleic acid sequence are more methylated related to another nucleic acid sequence, that sequence is considered hypermethylated or having increased methylation. Alternatively, if the cytosine residues of a nucleic acid sequence are less methylated related to another nucleic acid sequence, that sequence is considered hypomethylated or having decreased methylation.
[0162] In some embodiments, the "hypermethylated" or "hypomethylated" status is relative to the methylation status of the same loci from normal samples.
[0163] In some embodiments, a sample comprises a nucleic acid comprising a differentially methylated region (DMR) while in other embodiments, it comprises a plurality of DMRs. Preferably, each nucleic acid of interest includes at least one DMR.
[0164] The term "differentially methylated regions" (DMRs) refers to regions of the genome with different methylation levels from abnormal compared to normal samples, regarded as possible functional regions involved in physiological regulation.
[0165] Simultaneously, the term "differentially methylated cytosines" (DMCs) refers to cytosines with different methylation levels which could be regarded as possible functional loci involved in biological regulation. The joint effect of a group of adjacent DMCs forming a DMR could regulate gene expression.In the method according to the invention, the TArget Capture oligonucleotides in step (d), are long probes and each of the TAC oligonucleotides (i) is between 150-350 nucleotides in length, (ii) has a 5' end and a 3' end, (iii) preferably binds to the sequence of interest at least 50 base pairs away, on both the 5' end and the 3' end, from regions harboring segmental duplications or repetitive DNA elements, and (iv) has a GC content between 25% and 85%.
[0166] In one embodiment, the TAC oligonucleotides are between 150-350 nucleotides in length. In another embodiment, the TAC oligonucleotides are between 150-250 nucleotides, 250-350 nucleotides or 350-450 nucleotides in length.
[0167] In one embodiment, the TAC oligonucleotides have GC contents between 25% to 45%, 45% to 65% or 65% to 85%.
[0168] In one embodiment, the TAC oligonucleotides can bind to regions harboring segmental duplications or repetitive DNA elements.
[0169] Hybridization as used herein, refers to the annealing of one or more TAC oligonucleotides to target nucleotide sequences. Hybridization conditions typically include a temperature that is below the melting temperature of the TAC oligonucleotides while limiting non-specific hybridization of the TAC oligonucleotides.
[0170] To achieve isolation of the desired enriched sequences, usually the TAC oligonucleotide sequences are adapted in such a way that sequences that hybridize to the TAC oligonucleotides can be physically separated from sequences that do not bind to the TAC oligonucleotides. Typically, this is accomplished by fixing the TAC oligonucleotides to a support. This allows for physical separation of those sequences that bind the TAC oligonucleotides from those sequences that do not bind the TAC oligonucleotides. For example, each sequence within the pool of TAC oligonucleotides can be labelled with biotin and the pool can then be bound to beads coated with a biotin-binding substance, for instance streptavidin or avidin. In a preferred embodiment, the TAC oligonucleotides are labeled with biotin and bound to streptavidin-coated paramagnetic beads, thereby allowing separation by utilizing the magnetic property of the beads. In one embodiment, the biotin can be chemically linked to the primer used to generate the TAC oligonucleotide. In a second embodiment, the biotinylation step can be performed to each TAC oligonucleotide or to the pool of sequences that can hybridize the target region. The ordinarily skilled artisan will appreciate, nevertheless, that other affinity binding systems are known in the art and can be used rather than biotin-streptavidin / avidin. This includes, but is not limited to, an antibody-based method in which the TAC oligonucleotides are labeled with an antigen and then bound to antibody-coated beads. Furthermore, the TAC oligonucleotides can integrate on one end a sequence tag and can be bound to a support via a complementary sequence on the support that hybridizes tothe sequence tag. In addition to magnetic beads, other types of support can be used, such as polymer beads and so forth.
[0171] In one embodiment, the TAC oligonucleotides are provided in a form that allows them to be bound to a support, such as biotinylated TAC oligonucleotides. In another embodiment, the TAC oligonucleotides are provided together with a support, such as biotinylated TAC oligonucleotides provided together with streptavidin-coated magnetic beads. In another embodiment, the TAC oligonucleotides are found free in solution and are not bound to any solid support until after the hybrid capture.
[0172] Following enrichment of the sequence(s) of interest using the TAC oligonucleotides, thereby forming an enriched library of DNAs, the members of the enriched library are eluted, amplified and sequenced using standard methods known in the art.
[0173] Next Generation Sequencing is often used, even though other accurate counting methods such as digital PCR, single molecule sequencing, and microarrays can be used.
[0174] In some embodiments, the members of the converted sequencing library that bind to the pool of TAC oligonucleotides are fully complementary to the TAC oligonucleotides. In other embodiments, the members of the converted sequencing library that bind to the pool of TAC oligonucleotides are partially complementary to the TAC oligonucleotides. For instance, in certain circumstances it may be useful to exploit and analyze data that are from DNA fragments that have been captured by the enrichment process but do not necessarily belong to the genomic regions of interest (i.e., such DNA fragments could bind to a converted TAC oligonucleotide because of complete or partial homologies) and would produce coverage throughout the genome when sequenced.
[0175] In the context of this invention, the term "partial homology" indicates a genetic region of 5, 10, 20, 30 or more than 40 bases of the nucleic acid fragment sufficient for hybridization capture. Subsequently, the term "complete homology", as used herein, indicates a region which encompasses all of the nucleic acid fragments. In one embodiment, the TAC oligonucleotides hybridize to at least one location within the nucleic acid fragment. Nonetheless, more than one location within the genome can also be targeted by the same TAC oligonucleotide region, depending on its conversion status.
[0176] In one aspect, the method provides kits for carrying out the method of the invention. In one embodiment, the kit contains a container comprising of the pool of TAC oligonucleotides and hybridization reagents, the converted sequencing library preparation reagents, software and instructions for performing the method. In one embodiment, the TAC oligonucleotides are provided in a form that allows them to be bound to a solid support, such as biotinylated TAC oligonucleotides. In another embodiment, the TAC oligonucleotides are provided together with a solid support, such asbiotinylated TAC oligonucleotides provided together with streptavidin-coated magnetic beads. In various other embodiments, the kit can comprise additional components for carrying out other aspects of the method. For example, further to the pool of TAC oligonucleotides, the kit can consist of one or more of the following: (i) one or more components for isolating DNA from a sample, (ii) one or more components for preparing the converted sequencing library (e.g., primers, adapters, buffers, linkers, ligation reagents, polymerase reagents, DNA conversion reagents, DNA glycosylation reagents, DNA oxidation reagents and enhancers, DNA restriction reagents, and so forth), (iii) one or more components for enriching the converted sequencing library (e.g., probes, hybridization reagents, washing buffers and so forth), and / or (iv) software and instructions for performing the statistical analysis for the determination of the methylation pattern of a specific region or a plurality of genomic regions of interest in a sample and assessment of clinical disease in the said sample. In yet another embodiment, the kit comprises a container comprising the pool of TAC oligonucleotides, library reagents and instructions for performing the method wherein (i) each TAC oligonucleotide is between 150-350 nucleotides in length, each member sequence having a 5' end and a 3' end; (ii) each TAC oligonucleotide binds to the sequence of interest at least 50 base pairs away, on both the 5' end and the 3' end, from regions harboring segmental duplications or repetitive DNA elements; (iii) the GC content of the TAC oligonucleotides is between 25% and 85%; (iv) the pool of TAC oligonucleotides comprises of 30 or more distinct sequences; and (v) the pool of TAC oligonucleotides comprises of converted TAC oligonucleotides alone or in a mixture including partially converted and / or nonconverted TAC oligonucleotides.
[0177] In one embodiment, the TAC oligonucleotides are between 150-350 nucleotides in length. In another embodiment, the TAC oligonucleotides are between 150-250 nucleotides, 250-350 nucleotides or 350-450 nucleotides in length.
[0178] In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0179] In some embodiments, the technology is related to assessing the amount and methylation state of the cfDNA fragments obtained from a biological sample, cfDNA fragments which comprise one or more DMCs or DMRs.
[0180] In the present disclosure, the determination of methylation levels of the plurality of DMRs in a sample establishes said sample as normal, abnormal, hypomethylated or hypermethylated. After sequencing and obtaining a methylation status for each cfDNA fragment, said fragment is further identified as being normal, abnormal, hypomethylated or hypermethylated compared to healthy controls. As used herein, "hypomethylated" cfDNA fragments can be identified as fragments having at least oneunmethylated cytosine and / or CpG site. In the same way, "hypermethylated" cfDNA fragments can be identified as fragments having at least one methylated cytosine and / or CpG sites (Gross et al.).
[0181] In some embodiments, the method disclosed further comprises the detection and identification of the sample fragment as hypermethylated when the said fragment encompasses at least a threshold number of cytosines and / or CpG sites with more than a threshold percentage of the cytosines and / or CpG sites being methylated. In some embodiments, the threshold number of cytosines and / or CpG sites is 1 or more cytosines and / or CpG sites and the threshold percentage of said methylated cytosines and methylated CpG sites is 0.1% or greater. In some embodiments, the disclosure further comprises the detection and identification of the sample fragment as hypomethylated when the said fragment encompasses at least a threshold number of cytosines and / or CpG sites with more than a threshold percentage of the cytosines and / or CpG sites being unmethylated. In some embodiments, the threshold number of cytosines and / or CpG sites is 1 or more cytosines and / or CpG sites and the threshold percentage of unmethylated cytosines and unmethylated CpG sites is 0.1% or greater.
[0182] Cell free DNA in blood plasma consists of a mixture of fragmented DNA molecules of apoptotic cells from different tissues of the body. Cells of each tissue are differentiated to perform specific functions, therefore different patterns of genetic activity and epigenetic signatures are expected to be seen by each tissue type fragments.
[0183] In one embodiment, the origin of a nucleic acid can be detected and / or determined from a specific tissue or cell type based on its DNA methylation signature.
[0184] As used herein, "mapping quality" is defined as the probability that a read is aligned in the wrong place (phred-scaled posterior probability that the mapping position of this read is incorrect).
[0185] In one embodiment, preparing the DNA sequencing library comprises the step of including unique molecular identifiers (UMIs) to uniquely tag each molecule and to create UMI families.
[0186] In the context of the present invention, "frequency" may be used interchangeably with abundance or occurrence. In one embodiment of the invention, a variant allele frequency at a given locus describes the ratio of aligned cfDNA molecules supporting a variant allele at the given locus over the total number of aligned cfDNA molecules that span the locus.
[0187] In the context of the present invention, the term minimal residual disease (MRD) may refer to the very small number of cancer cells that remain in the body during or after treatment.Thus, in one embodiment the method of the invention relates to and solves the problem of detecting relapse and monitoring MRD, wherein the detection of abnormal methylation levels in a patient following treatment can indicate insufficient treatment response or disease recurrence.
[0188] Herein, next-generation sequencing (NGS) may be used for nucleic acid sequence analysis, although other sequencing technologies can also be employed, which provide very accurate counting in addition to sequence information. Accordingly, other accurate counting methods, such as but not limited to digital PCR, single molecule sequencing, nanopore sequencing, DNA nanoball sequencing, sequencing by ligation, Pyrosequencing, Ion semiconductor sequencing, semiconductor sequencing, sequencing by synthesis, and microarrays can also be used.
[0189] As used herein, the terms "Target Capture oligonucleotides", "TAC oligonucleotides" "TAC" or "probe" are used interchangeably and refer to DNA oligonucleotides that are complementary to the region(s) of interest on a genomic sequence(s) of interest and which are used as bait to capture and enrich the region of interest from a large library of sequences, such as a whole genomic converted sequencing library prepared from a biological sample. A pool of TAC oligonucleotides is used for enrichment wherein the oligonucleotides within the pool have been optimized with regard to: (i) the length of the oligonucleotides, (ii) the distribution of the TAC oligonucleotides across the region(s) of interest, (iii) the GC content of the TAC oligonucleotides, (iv) the capture of converted DNA, and (v) the specific capture of converted or non-converted DNA of interest. The number of oligonucleotides within the TAC oligonucleotide pool (pool size) has also been optimized.
[0190] Preferably, and in accordance with the invention, a TAC oligonucleotide that binds the target's plus or minus strand can be modified:
[0191] a) by converting all its cytosines into uracils, or
[0192] b) by converting only a portion of its cytosines into uracils.
[0193] Ideally, the modification of the TAC oligonucleotides involves converting one or more of the unmethylated cytosines.
[0194] In a preferred embodiment of the invention,
[0195] a. each TAC oligonucleotide is between 150-350 nucleotides in length;
[0196] b. each TAC oligonucleotide has a 5' end and a 3' end;
[0197] c. each TAC oligonucleotide has a GC content between 25% and 85%;d. each TAC oligonucleotide binds at least 50 base pairs away from regions harbouring segmental duplications or repetitive DNA elements; and / or
[0198] e. there are 30 or more different TAC oligonucleotides present in the method.
[0199] f. each TAC oligonucleotide is present in certain proportions of converted and un-converted form.
[0200] In another embodiment, the TAC oligonucleotides are between 150-250 nucleotides, 150-350 nucleotides, 250-350 nucleotides or 350-450 nucleotides in length.
[0201] In one embodiment, the TAC oligonucleotides have GC contents between 25% to 45%, 45% to 65% or 65% to 85%.
[0202] In one embodiment, the TAC oligonucleotides can bind to regions harboring segmental duplications or repetitive DNA elements.
[0203] The inventors have achieved superior results if the deamination of cytosine to uracil of the genomic DNA and / or the TAC oligonucleotides was performed enzymatically. Nonetheless, the person skilled in the art would appreciate the fact that other methods can be used to achieve similar results in the detection of 5mC. This includes but is not limited to chemical deamination by sodium bisulfite.
[0204] In one embodiment of the invention, one or more of the TAC oligonucleotides were modified only partially and not all cytosines were converted to uracil. This means that for example if a TAC oligonucleotide is made-up of 20 potential sites for conversion, then in the method of the present invention it may be decided to convert all of the 20 sites or maybe only a portion of these. This is an important aspect of the invention as this principle can be used to take into account varying degrees of methylation of the genomic DNA. In one embodiment, the decision to partially convert a TAC oligonucleotide is based on the desire to enrich a fragment of interest with a specific epigenetic level, e.g., hypermethylated.
[0205] Therefore, in accordance with the present invention, varying amounts of conversion may be performed for example and in a preferred embodiment 100%, 80%, 70%, 60%, 50%, or less than 50% of the cytosines are converted to uracil. This is a method of adjusting the hybridization specificity to the genomic DNA. Due to the fact that the TAC oligonucleotides may be converted specifically on a sample-by-sample basis, the method can be adjusted to varying methylation degrees.
[0206] In the context of the present invention, conversion to detect epigenetic markers can be performed either enzymatically (EM-seq) or chemically (e.g. bisulfite sequencing, TABseq, TAPS). The person skilled in the art will appreciate that amplified converted TAC oligonucleotides could generateadditional TAC oligonucleotide sequences which could optimize the hybrid capture efficiency, besides improving the ease of TAC oligonucleotide production. A person skilled in the art would appreciate that different types of conversion could be applied to specifically target different episignatures originating either from DNA or RNA molecules.
[0207] In a preferred embodiment, the fragments which have been enriched are amplified before they are sequenced. Standard amplification methods include but are not limited to, e.g., the polymerase chain amplification, or isothermal amplification.
[0208] The method of the current invention encompasses additional steps such as aligning the sequenced enriched fragments to a standardized vertebrate genome sequence, preferably human, comparing the genetic transition levels of cytosine to thymine and guanine to adenine. Hence the invention comprises:
[0209] aligning all sequenced fragments to the vertebrate genome and to the converted vertebrate genome;
[0210] assigning the methylation status in the cytosines of the aligned fragment by detecting the induced conversions from step (b), where an induced conversion is exhibited as a genetic transition when compared to the vertebrate genomes; and / or
[0211] performing statistical analysis to assess the level of methylation of the cytosines of interest.
[0212] As discussed above, the sample can have many origins. Preferably, the nucleic acid sample is a vertebrate sample, preferably a human sample, preferably a blood sample, and more preferably a plasma sample.
[0213] Most preferably, the nucleic acid sample comprises cell-free DNA (cfDNA) or cell-free RNA (cfRNA) and the cfDNA or cfRNA is the nucleic acid preferentially analysed.
[0214] For quality control reasons and quantification, it is preferred if the DNA sample is combined with spikein control DNAs having fully methylated or unmethylated CpGs. Preferably, the control DNAs are from non-vertebrates and are fragmented in size, wherein the TAC pool is also enriched by spike-in control DNAs.
[0215] In a preferred embodiment, adapters are added to the genomic DNA sample prior to a conversion of cytosine to uracil and more preferably an adapter where its cytosines are protected from deamination.Also preferably, the adapters are added to the DNA fragments sample prior to a conversion of cytosine to uracil. Such adapters include, but are not limited to, NEBNext© Multiplex Oligos, NEXTFLEX© Bisulfite-Seq Barcodes from PerkinElmer, IDT xGen™ Methylation Sequencing methylated stubby adapter, or self-designed protected from conversion adapters produced by any oligonucleotide synthesis company that can produce oligos with methylated cytosines. However, a person of ordinary skills in the art will appreciate that other enzymatic or chemical reactions may be used to modify methylated cytosines to achieve the same effect. One alternative example provided herein is the use of Pyrollo-dc for protecting cytosines from being converted to uracil by cytidine deaminase.
[0216] The inventors have had outstanding results when protection of methylated cytosines from conversion was performed by a methylcytosine dioxygenase enzyme, including but not limited to Ten-Eleven Translocation dioxygenase 2 (TET2), such as recombinant TET2 (1129-2002) protein from Active Motif and / or DNA beta-glucosyltransferase (BGT), for the simultaneous protection of hydroxymethylated cytosines from conversion, including but not limited to recombinant T4 DNA beta-glucosyltransferase protein from MyBiosource.
[0217] Conversion can be performed by a deaminase enzyme, including but not limited to APOBEC3A, recombinant APOBEC3A (A3A) from Active Motif, or Cytidine deaminase human from Sigma Aldrich. Nonetheless, the ordinary skilled artisan will appreciate that deamination can also be achieved by other methods. This includes, but is not limited to, chemical conversion through deamination by sodium bisulfite. A person skilled in the art would appreciate that different types of conversion could be applied to specifically target different episignatures originating either from DNA or RNA molecules.
[0218] The following combination is preferred wherein the deamination involves the use of a methylcytosine dioxygenase enzyme, preferably Ten-Eleven Translocation dioxygenase 2 (TET2), a glucosyltransferase enzyme, preferably T4-phage beta-glucosyltransferase (T4-betaGT), and a DNA cytidine deaminase, preferably apolipoprotein B mRNA editing catalytic polypeptide-like 3A enzyme (APOBEC3A). However, the person skilled in the art acknowledges that deamination can also be done chemically.
[0219] Preferably, the TAC oligonucleotides comprise an adapter and the adapter is a biotin adapter. Once the TAC oligonucleotides bind the target, the adapter is used to immobilize the oligo. This is used to isolate the heteroduplex from any background buffer or other nucleic acids.
[0220] It is preferred that the TAC oligonucleotides comprise an adapter and the adapter is preferably a biotin adapter. This is done to bind the TAC oligonucleotides to a solid support, such as streptavidin-coated magnetic beads.In one embodiment, the biotinylated adapter can be used to amplify the whole pool of TAC oligonucleotides to prepare the latter in an efficient and cost-effective manner. When amplification occurs in the presence of converted TAC oligonucleotides, it generates additional sequences that further improve the enrichment efficiency due to the possibility of specifically capturing supplemental converted sequencing library fragments by said new TAC oligonucleotide.
[0221] While the bioinformatic analysis can be done in many ways, determining the methylation pattern preferably comprises steps selected from the group of: i) a demultiplexing procedure whereby sequenced reads are assigned to their samples, as identified via an index sequence, producing FASTQ. files, ii) running an aligner that maps the sequenced reads on the genomic coordinates that have the highest mapping likelihood, said likelihood is calculated from mapping in both the converted and nonconverted versions of the human genome, iii) processing the aligned sequences using a software suite that determines cytosine methylation status from converted sequence reads, iv) extracting the number of methylated and non-methylated cytosines for each target region, v) applying an exact binomial test to calculate the confidence interval for the methylation level of each CpG site, vi) computing reference CpG methylation levels for various vertebrate tissues, vii) quantifying the methylation level strand bias and / or the proportion of methylated cytosines (not CpG sites) and / or the coefficient of variation of each CpG, and viii) computing the depth of coverage (sequencing read depth) for each targeted CpG. The information obtained from steps g) to h) can lead to the determination of cancer and disease risk prediction, cancer and disease diagnosis, cancer and disease prognosis, cancer and disease subclassification, drug response prediction, MRD status, risk of recurrence, treatment planning, therapy selection and / or treatment efficacy of the tested sample and patient stratification for clinical trial enrollment.
[0222] In an embodiment, strand bias for a CpG position in a methylation assay is defined as the presence of statistically significant difference between the measured methylation level of that position on the forward strand and the measured methylation level of that position on the reverse strand.
[0223] The process of enzymatic detection of methylated cytosines to the fifth carbon (5mC) involves three enzymes and two sets of reactions. 5-methylated cytosines (5mC) are protected from deamination firstly by an oxidative step then a glycosylation step involving methylcytosine dioxygenase enzyme and glycosylase enzymes (e.g., TET2, T4PGT). Initially, 5-methyl cytosines (5mC) are oxidated to 5-hydromethylcytosines, 5-formylcytosines and 5-carboxycytosines. The second step involves the glycosylation of 5-hydroxymethylcytosines (pre-existing in the genome or formed by the oxidation step) to 5-p-glucosyloxymethylcytosines. Following the protection steps, only non-modified cytosines are converted to uracil, thus making the detection of methylated cytosines specific. The person skilledin the art would appreciate that detection of methylated cytosines can also be achieved with the use of other methods including, but not limited to OxBS-seq, DM-seq and TAB-seq.
[0224] In a preferred embodiment, the adapter incorporated during the sequencing library must be specifically designed to be protected from conversion of cytosines to uracils, e.g., by using methylated cytosine. In another embodiment, the adapter for conversion of the sequencing library must be specifically designed to take into account the conversion of cytosines to uracils.
[0225] Ideally, the amplified libraries are mixed with blocking oligos (e.g., 200pM), Cot-1 DNA (e.g., 5pg from Thermo Fisher Scientific at 1 mg / ml), and / or Salmon Sperm DNA (e.g., 50pg from Invitrogen at lOmg / ml).MEDICOVER CH Kilger Anwaltspartnerschaft mbB Germany FasanenstraRe 29 Our Ref.: B281-0058W01 10719 Berlin FIGURES
[0226] Figure 1: A. A converted sequencing library containing both converted and non-converted cytosines, hybridized with a pool of non-converted TAC oligonucleotides presenting selective capture of the DNA methylation level.
[0227] B. Illustration of the basic principles of the invention. The use of a mixture of regular and converted TAC oligonucleotides to optimally enrich, through in-solution hybridization method, any fragments of interest; whether they were converted or not (i.e., methylated or not). The method allows for the sensitive and specific detection of methylated cytosine in a cfDNA.
[0228] Figure 2: Schematic of the method. The converted sequencing library, containing both converted (unmethylated) and non-converted (methylated) cytosines, is hybridized with the pool of TAC oligonucleotides containing both converted and non-converted biotinylated TAC oligonucleotides for the optimal assessment of the DNA methylation level. Subsequently, sequencing of the enriched library and bioinformatic analysis are performed for the determination of the methylation status of cytosines of interest present in the sample and the detection of possible biomarkers or disease status assessment.
[0229] Figure 3: Boxplot demonstrating the enrichment efficiency (average read depth per Gb of output) with two different sizes of TAC oligonucleotides. Two different pools of 157 mixed converted and nonconverted TAC oligonucleotides targeting the same regions were used to enrich two sets of 6 identical pool plasma libraries each. The first pool comprises TAC oligonucleotides of median length 150 nucleotides while the second pool has a median length of 250 nucleotides. The larger TAC oligonucleotides result in more than four times the enrichment efficiency compared to the smaller TAC oligonucleotides.
[0230] Figure 4: A. Comparison of methylation levels at selected CpGs using the method of the current invention with a mixture of converted and non-converted TAC oligonucleotides versus a standard state of the art whole methylome sequencing (WMS) method. A strong correlation is observed (r=0.975).B. Increased depth of coverage (read depth) at the same selected CpGs using the method of the current invention with a mixture of converted and non-converted TAC oligonucleotides versus a standard state of the art whole genome methylation sequencing (WMS) method. A significantly improved coverage is observed in almost all regions (t-test p-value < 0.001) along with strongly correlated methylation levels.
[0231] Figure 5: Comparison of methylation levels at selected hyper- (methylation level greater than 70%) and hypo- (methylation level lower than 10%) methylated CpGs using the method of the current invention with a mixture of converted and non-converted TAC oligonucleotides versus the method of the current invention with a pool of non-converted TAC oligonucleotides only. The variance of the former approach is significantly better (lower) than the latter (F-test p-value < 0.001) for the hypomethylated CpGs. An over-estimation of methylation levels is observed using solely nonconverted TAC oligonucleotides compared to using a mixture of converted and non-converted TAC oligonucleotides.
[0232] Figure 6: Expected methylation levels at CpG sites of methylated and non-methylated controls using the method of the current invention with a mixture of converted and non-converted TAC oligonucleotides are estimated at 96% and 0.6%, respectively.
[0233] Figure 7: A. Methylation levels (y-axis) at 23 consecutive CpGs for 3 normal samples in a region with an expected methylation level less than 1% (from WMS). CpG methylation levels of each sample are represented by a dotted-solid line.
[0234] B. Methylation levels (y-axis) at 16 consecutive CpGs for 3 normal samples in a region with an expected methylation level greater than 90%. CpG methylation levels of each sample are represented by a dotted-solid line.
[0235] Figure 8: Example of two identified differentially methylated regions. A total of 17 normal samples and 7 abnormal samples were compared to identify differentially methylated regions. Said normal samples comprise 13 plasma and 4 buffy coat samples taken from healthy donors not previously diagnosed with cancer. Said abnormal samples include 7 tissue samples taken from patients diagnosed with NSCLC cancer.Figure 9: Receiver operating characteristic curve illustrating the classification performance of one embodiment of the current invention using 18 plasma samples taken from healthy donors not previously diagnosed with cancer and 11 abnormal plasma samples taken from patients diagnosed with colorectal cancer. A sensitivity greater than 0.8 can be achieved with a high specificity at 0.95.Medicover CH Kilger Anwaltspartnerschaft mbB Cyprus FasanenstraRe 29 Our Ref.: B281-0058W01 10719 Berlin
[0236] EXAMPLES
[0237] TArget Capture oligonucleotide Design
[0238] As used herein, the term "TArget Capture oligonucleotides" or "TAC oligonucleotides" or "probes" refers to DNA sequences that are complementary to the region(s) of interest (e.g., DMRs) present in a converted sample which are used as "baits" to capture and enrich the region of interest from a large library of sequences. A pool of TAC oligonucleotides is used for enrichment wherein the sequences within the pool have been optimized relating to: i) the length of the sequences; ii) the distribution of the TAC oligonucleotides across the region(s) of interest; iii) the GC content of the TAC oligonucleotides; iv) methylation level of the region of interest; and v) strand bias. Furthermore, the number of sequences within the TAC oligonucleotide pool (size of the pool), tiling, overlapping and proportions of converted and non-converted baits have been optimized.
[0239] The TAC oligonucleotides are designed with specific GC content characteristics so as to reduce data GC bias and to allow a custom and innovative data analysis pipeline. It has been established that TAC oligonucleotides with a GC content of 25-85% achieve optimal enrichment and perform best with cell-free DNA. Within the pool of TAC oligonucleotides, different sequences can have different % of GC content. Even though to be selected for inclusion with the pool, the % GC content of each sequence is chosen as between 25-85%, as determined by calculating the GC content of each member within the pool of TAC oligonucleotides. Specifically, every TAC oligonucleotide used has a % GC content within the given percentage range (e.g., between 25-85% GC content). Nonetheless, the ordinarily skilled artisan would consider the change of % GC content following conversion.
[0240] In some instances, the pool of TAC oligonucleotides (i.e., each member within each family of TAC oligonucleotides) may be chosen to have a different % GC content range, deemed to be more suitable for the enrichment and assessment of the proportion of methylated cytosines. Non-limiting examples of various % GC content ranges can be between 25% and 85%, or between 25% and 80%, or between 25% and 75%, or between 25% and 70%, or between 25% and 65%, or between 25% and 60%, or between 25% and 55%, or between 25% and 50%. Nonetheless, the ordinarily skilled artisan would appreciate that there are more variations of % GC content ranges that can be used instead of the ones mentioned above.
[0241] To design a pool of TAC oligonucleotides having the optimal characteristics set forth above with respect to size, position in the human genome, converted DNA sequence, proportions of converted and non-converted baits and % GC content, both manual and computerized analysis methods known in the art can be applied to the analysis of the human reference genome. In one embodiment, a semi-automatic method is applied where regions are firstly manually designed based on the human reference genome built 19 (hgl9) ensuring that the aforementioned repetitive regions are avoided and subsequently curated for GC-content using software that calculates the % GC-content of each region based on its coordinates on the human reference genome built 19 (hgl9). In another embodiment, custom-built software is used to analyze the human reference genome in order to identify suitable TAC oligonucleotide regions which comply with certain criteria, such as but not limited to, % GC content, converted DNA sequence, proximity to repetitive regions and / or proximity to other TAC oligonucleotides.
[0242] In the current invention, the number of TAC oligonucleotides in the pool has been carefully examined and adjusted to achieve the best balance between result robustness and assay cost / throughput. The pool typically contains at least 30 or more TAC oligonucleotides, but can include as many as 500 or more TAC oligonucleotides, 1000 or more TAC oligonucleotides, 2000 or more TAC oligonucleotides, 4000 or more TAC oligonucleotides, 7000 or more TAC oligonucleotides, 13000 or more TAC oligonucleotides, 25000 or more TAC oligonucleotides, 50000 or more TAC oligonucleotides, 100000 or more TAC oligonucleotides, 200000 or more TAC oligonucleotides, 400000 or more TAC oligonucleotides.
[0243] In view of the foregoing, in another aspect, the invention provides a method for preparing a pool of TAC oligonucleotides for use in the method of the invention for evaluating the level of methylation of sequences of interest, wherein the method for preparing the pool of TAC oligonucleotides comprises: selecting regions in one or more converted DNA sequences of interest having the criteria set forth above (e.g., at least 50 base pairs away on either end from the aforesaid repetitive sequences and GC content of between 25% and 85%, as determined by calculating the GC content of each member within the pool of TAC oligonucleotides), using primers that amplify sequences that hybridize to the selected regions, and amplifying the sequences, wherein each sequence is 150-350 nucleotides in length. In one embodiment, the biotin can be chemically linked to the primer used to generate the TAC oligonucleotide. In a second embodiment, the latter can be generated by biotinylating the pool of sequences that can hybridize the target region. Similarly, the TAC oligonucleotides can be biotinylated individually or after being pooled altogether.
[0244] In one embodiment, the TAC oligonucleotides are between 150-350 nucleotides in length. In another embodiment, the TAC oligonucleotides are between 150-250 nucleotides, 250-350 nucleotides or 350-450 nucleotides in length.TAC oligonucleotide conversion
[0245] In the current invention, TAC oligonucleotides are used in mixtures of converted, partially converted and non-converted forms for the more sensitive and specific detection of methylated cytosines from the tested sample. The person skilled in the art would follow any one of the above-mentioned methods, described in the background of the invention section, to convert TAC oligonucleotides such as, including but not limited to, chemical and / or enzymatic treatment. In one embodiment, the method uses a combination of converted and non-converted TAC oligonucleotides. In another embodiment, the method uses a combination of converted and partially converted TAC oligonucleotides. In yet another embodiment, the method uses a combination of partially converted and non-converted TAC oligonucleotides. Non-limiting examples of various ratios of the aforementioned combinations can be 1:1:0, 1:0:1, 2:2:1, 0:1:1, 3:2:1, 2:1:2, 1:3:2. Nonetheless, the ordinary skilled artisan would appreciate that there are more variations of ratios that can be used instead of the ones mentioned above.
[0246] Sample collection and preparation
[0247] The methods of the invention are performed on any vertebrate sample containing vertebrate DNA, preferably human sample containing human DNA. Typically, the sample is a human plasma sample, although other tissue sources that contain human DNA can be used. Blood plasma can be obtained from a peripheral whole blood sample from a human individual and the plasma can be obtained by standard methods. As little as 0.5 - 8ml of plasma is sufficient to provide suitable DNA material for analysis according to the method of the invention. Total cell free DNA can then be extracted from the sample using standard techniques, non-limiting examples of which include a Mag-bind protocol (Omega Bio-Tek) suitable for cell free DNA isolation or any other manual or automated extraction method suitable for cell free DNA isolation.
[0248] Subsequent to isolation, the cell free DNA of the sample is end-prepped, followed by ligation of the DNA to the adapter. Subsequently, methylated and hydroxymethylated cytosines are firstly chemically modified to prevent their conversion into uracil. Successively, non-methylated cytosines are converted into uracil by a conversion reagent allowing thereafter differentiating cytosines from 5-methylcytosines and 5-hydroxymethylcytosines. PCR amplification is then performed followed by eventual assessment of the size distribution and concentration of the converted libraries prior to hybrid-capture enrichment using TAC oligonucleotides.In one embodiment, the methylated cytosines can be the ones converted leading to direct detection of said episignature.
[0249] In one embodiment, one sequencing library is constructed. In other embodiments, two or more sequencing libraries are constructed in parallel and are combined prior to hybridization.
[0250] Enrichment by hybridization
[0251] The region(s) of interest on the sequence(s) of interest is enriched by hybridizing the pool of TAC oligonucleotides to the converted sequencing library, accompanied by isolation of those sequences that bind to the TAC oligonucleotides. To achieve isolation of the selected, enriched sequences, typically the TAC oligonucleotide sequences are modified in such a way that sequences that hybridize to the TAC oligonucleotides can be separated from sequences that do not hybridize to the TAC oligonucleotides. Typically, this is accomplished by fixing the TAC oligonucleotides to a solid support, thus allowing for the physical separation of those sequences that bind the TAC oligonucleotides from those sequences that do not bind the TAC oligonucleotides. For example, each sequence within the pool of TAC oligonucleotides can be labelled with biotin and the pool can then be bound to beads coated with a biotin-binding substance, such as streptavidin or avidin. In a preferred embodiment, the TAC oligonucleotides are labelled with biotin and bound to streptavidin-coated magnetic beads. Nonetheless, the ordinarily skilled artisan will appreciate that other affinity binding systems are known in the art and can be used instead of biotin-streptavidin / avidin. For instance, an antibody-based system can be used in which the TAC oligonucleotides are labeled with an antigen and then bound to antibody-coated beads. Moreover, the TAC oligonucleotides can integrate on one end a sequence tag and can be bound to a solid support via a complementary sequence on the solid support that hybridizes to the sequence tag. Moreover, in addition to magnetic beads other beads can be used, such as polymer beads and the like. In another embodiment, the biotin can be chemically linked to the primer used to generate the TAC oligonucleotide. In a second embodiment, the TAC oligonucleotide can be generated by biotinylating the pool of sequences that can hybridize the target region while in another embodiment, the biotin can be incorporated during synthesis of the TAC oligonucleotides. In yet another embodiment, the TAC oligonucleotides are found free in solution and are not bound to any solid support.
[0252] In certain embodiments, the members of the converted sequencing library that bind to the pool of TAC oligonucleotides are fully complementary to the TAC oligonucleotides. In other embodiments, the members of the converted sequencing library that bind to the pool of TAC oligonucleotides are partially complementary to the TAC oligonucleotides allowing efficient capture even in the presence of smallsequence or conversion variations. For example, in specific circumstances it may be preferable to employ and analyze data that are from DNA fragments that are products of the enrichment process but that do not essentially belong to the genomic regions of interest. Such DNA fragments could bind to the TAC oligonucleotides because of partial complementarity with the TAC oligonucleotides and when sequenced would produce coverage throughout the genome in non-TAC oligonucleotide coordinates.
[0253] Following enrichment of the sequence(s) of interest using the TAC oligonucleotides, thus forming an enriched library, the members of the enriched library are optionally eluted from the solid support and are amplified and sequenced utilizing standard methods known in the art. Next Generation Sequencing is typically used, even though other sequencing technologies can also be used, which offers very precise counting in addition to sequence information and methylation estimation.
[0254] Data analysis
[0255] Next generation sequencing (NGS) data were processed using state of the art methods known to those skilled in the art. The NGS sequencing reads were subjected to a demultiplexing procedure whereby sequenced reads were assigned to their samples, as identified via an index sequence, producing FASTQ files. Each sample's FASTQ files were processed using the Bismark software suite, as part of which the Bowtie2 aligner is run on converted and non-converted versions of the human reference genome built 19 (hgl9). In one embodiment, said processing comprises, but is not limited to, removal of NGS sequencing reads with at least two consecutive non-converted cytosines (not in a CpG context). In one embodiment, NGS sequencing reads with at least one or two or three or four or five non-converted cytosines (not in a CpG context) are subsequently removed. The CpG context is subsequently extracted, i.e., the number of methylated and non-methylated cytosines for each targeted region.
[0256] In one embodiment, for each aligned sequencing read, the probability of the said sequencing read being derived from tissue T (e.g. hematopoietic cells or tumor tissue or liver tissue etc.) is estimated using a Poison distribution for the variable X denoting the number of times the sequencing read methylation pattern is observed, i.e X~Po(np), an approximation to a Binomial distribution when n is large and p is small, where n is the total number of molecules spanning the region and p is a parameter representing the expected proportion of molecules harboring the said pattern, estimated from a set of reference samples from tissue T. In other embodiments, the probability density function for X is estimated using a Negative Binomial distribution or a non-parametric kernel density estimator. In various embodiments, deep learning methods on large data sets can be applied to perform regression or risk classification of samples.Cell-free DNA fragmentation is a non-random process. Cell-free DNA in plasma consists of a mixture of fragmented DNA molecules released from various tissues within the body. Each cell-free DNA fragment bears molecular signatures of its cell of origin. The current invention could be used to deconvolve the signal from these unique signatures using deep learning. These features comprise DNA methylation motifs, nucleosome footprints, size of molecules, fragment-end sequence motifs and mutational signatures. In one embodiment of the invention, the origins of cfDNA molecules can be traced back to their corresponding tissues using a multiomics ensemble method since both genomic and epigenomic information are retained. Said information comprises methylation and / or fragmentation and / or topological and / or copy number patterns detected in cfDNA samples subjected to the methodology of the current invention. In some embodiments of the method, multiple deep learning and / or principal component analysis methods are integrated within a multi-dimensional feature space. Said analysis methods can be nested or ensemble or stacked regression or classification or deconvolution models, examples of which are logistic or multinomial regression, random decision forest, support vector machine, naive Bayes or artificial neural networks. In other embodiments of the method, an in-silico enrichment for fragments bearing unique signatures of methylation motifs and / or fragmentation patterns prior to model training was performed.
[0257] In one embodiment, the CpG methylation status pattern on each sequenced fragment from a plurality of sequenced fragments is computed in order to determine the tissue of origin. Said sequenced fragments are defined as pairs of sequenced reads, with said CpG methylation pattern being defined by the binary status (0: unmethylated; 1: methylated) of each CpG on said fragment, i.e. a fragment spanning four methylated CpGs in a row and two unmethylated CpGs in a row would have a CpG methylation pattern defined by 111100.
[0258] Examples of statistics that can be calculated using the current invention and can, subsequently, be used for biomarker discovery and / or within a classification model are listed below:
[0259] Statistic SI: M J / NJ, where MJ denotes the number of fragments having CpG j at a methylated state and NJ denotes the total number of fragments spanning CpG j.
[0260] Statistic S2: LFM_r / (LFM_r+LFU_r), where LFM_r denotes the number of fully methylated fragments and LFU_r denotes the number of fully unmethylated fragments in region r. Said fully methylated fragments are defined as those fragments having all CpGs on them at a methylated state. Said fully unmethylated fragments are defined as those fragments having none of the CpGs on them at a methylated state. In other embodiments, a fully methylated fragment is defined as the fragment with at least 70% or 75% or 80% or 85% or 90% or 95% of the CpGs at a methylated state. In otherembodiments, a fully unmethylated fragment is defined as the fragment with at most 5% or 10% or 15% or 20% of the CpGs at a non-methylated state.
[0261] Statistic S3: NCSJ / NJ, where NCS_i denotes the number of fragment end points at base i and NJ denotes the total number of fragments spanning base i.
[0262] Statistic S4: SF_r / LF_r, where SF_r denotes the number of fragments with size less than or equal to 148bp at region r (minimum size of region is lOObp) and LF_r denotes the number of fragments with size greater than 150bp and less than 200bp at region r.
[0263] Statistic S5: D_r / S_r, where D_r represents the diversity of methylation motifs on fragments aligned within region r. Said diversity is defined as the number of times a specific methylation motif (e.g., 1000001, where 1 denotes a methylated CpG and 0 denotes a non-methylated CpG for a fragment spanning 7 CpGs with only the first and last one being methylated). S_r represents the total number of fragments aligned in region r.
[0264] Statistic S6: -(l / C_r)*sum(D_r / S_r,*log(D_r / S_r„2)), where C_r is the total number of CpGs lying within region r.
[0265] Example 1
[0266] In one embodiment, control regions (having a known methylation status such as for housekeeping genes constantly transcribed) were tested on three plasma samples from healthy individuals not previously diagnosed with cancer. Figure 7A illustrates an example of a known hypomethylated region comprising 23 CpGs (expected methylation levels less than 1%) on x-axis. The methylation level (y-axis) for each CpG is represented by a dot. All CpG methylation levels of each sample are connected with dotted lines. Figure 7B illustrates an example of a known hypermethylated region, specifically of methylation level greater than 90% at 16 consecutive CpGs for 3 normal samples.
[0267] Example 2
[0268] The current invention can serve as a powerful tool for biomarker discovery. A total of 17 normal samples and 7 abnormal samples were compared to identify differentially methylated regions. Saidnormal samples comprise of 13 plasma and 4 buccal swab samples taken from healthy donors not previously diagnosed with cancer. Said abnormal samples include 7 tissue samples taken from patients diagnosed with non-small cell lung cancer. Figure 8 shows two regions identified to have a statistically different methylation level between normal and abnormal samples (p-value<0.0004). Differentially methylated regions are identified using the Hotelling's T-squared distribution. In other embodiments of the method, other statistics (apart from the methylation level per CpG) were used for biomarker discovery. Said markers can be discovered using any combination or function of statistics S1-S6.
[0269] Example 3
[0270] In another embodiment of the method, a stacked machine learning classifier was used (base models: random decision forest, KNN and Naive Bayes classifiers; meta-model: logistic regression) to predict the presence of ctDNA in plasma samples. A total of 2784 biomarkers were used. The covariates used in each of the base classification models are based on S1-S6. In other embodiments of the method, other covariates were used including, but not limited to, fragment size distribution differences and / or preferred fragment end sites distribution, and frequency of specific proportions of methylated CpGs on each fragment, said proportions being equal to 0, 0.4, 0.5, 0.7, 0.9 and 1. A cross-validation (leave one out method) was applied to assess the performance of the machine learning classifier using normal and abnormal plasma samples. Said normal samples comprise 18 plasma samples taken from healthy donors not previously diagnosed with cancer. Said abnormal samples comprise 11 plasma samples taken from patients diagnosed with colorectal cancer. A promising performance is achieved using the methodology taught in the current invention, said performance being estimated at 82% (95% Cl: 48-98%) sensitivity and 95% (95% Cl: 75-99.9%) specificity with an AUC being equal to 0.90325 (Figure 9).
[0271] Kits of the invention
[0272] In another embodiment, the disclosure provides kits for carrying out the methods of the invention. In one embodiment, the kit comprises a container consisting of the pool of TAC oligonucleotides and instructions for performing the method. In one embodiment, the TAC oligonucleotides are provided in a form that allows them to be bound to a solid support, such as biotinylated TAC oligonucleotides. In another embodiment, the TAC oligonucleotides are provided together with a solid support, such as biotinylated TAC oligonucleotides supplied together with streptavidin-coated magnetic beads. In various other embodiments, the kit can include additional components for carrying out other aspects of the method. For example, in addition to the pool of TAC oligonucleotides, the kit can consist of oneor more of the following: (i) one or more components for isolating DNA from a sample; (ii) one or more of components for preparing the converted sequencing library (e.g., primers, adapters, buffers, linkers, ligation reagents, polymerase reagents, DNA conversion reagents, DNA glycosylation reagents, DNA oxidation reagents and enhancers, DNA restriction reagents and so forth); (iii) one or more components for enriching the converted sequencing library (e.g., probes, hybridization reagents, washing buffers and so forth), and / or (iv) software and instructions for performing the statistical analysis for the determination of the methylation pattern of a specific region or a plurality of genomic regions of interest in a sample and assessment of clinical disease in the said sample. In yet another embodiment, the kit comprises a container comprising the pool of TAC oligonucleotides, library reagents and instructions for performing the method wherein (i) each TAC oligonucleotide is between 150-350 nucleotides in length, each member sequence having a 5' end and a 3' end; (ii) each TAC oligonucleotide binds to the sequence of interest at least 50 base pairs away, on both the 5' end and the 3' end, from regions harboring segmental duplications or repetitive DNA elements; (iii) the GC content of the TAC oligonucleotides is between 25% and 85%; (iv) the pool of TAC oligonucleotides comprises of 30 or more distinct sequences; and (v) the pool of TAC oligonucleotides comprises of converted TAC oligonucleotides alone or in a mixture including partially converted and / or nonconverted TAC oligonucleotides. In one embodiment, the TAC oligonucleotides are between 150-350 nucleotides in length. In another embodiment, the TAC oligonucleotides are between 150-250 nucleotides, 250-350 nucleotides or 350-450 nucleotides in length.
Claims
1. CLAIMS1. Method based on methylation assessment for cancer and disease risk prediction, cancer and disease diagnosis, cancer and disease prognosis, accurate cancer and disease subclassification, MRD detection, disease recurrence, drug response prediction, treatment planning, therapy selection and / or treatment efficacy monitoring, patient stratification for clinical trial enrollment by determining the methylation status of a single or a plurality of cytosine bases in genomic DNA molecules of a vertebrate sample comprising the steps of:a. obtaining a sample's DNA, preferably a blood sample from a human subject, more preferably a plasma sample;b. converting the tested sample's DNA molecules;c. amplifying the tested sample's converted DNA to create a sequencing library;d. hybridizing one or more Target Capture (TAC) oligonucleotides to the target DNA of the sequencing library, wherein the target DNA comprises both the plus and the minus strands, wherein the TAC oligonucleotides have been modified to also bind the target DNA's amplified converted strands or their complementary strands thereto;e. enriching the sequencing library by isolating the nucleic acids of the library that have bound the TAC oligonucleotides;f. sequencing the enriched library;g. performing statistical analysis on the enriched library sequences comprising the steps of:i) aligning all sequenced fragments to the vertebrate genome and to the converted vertebrate genome;ii) assigning the methylation status in the cytosines of the aligned fragment by detecting the induced conversions from step (b), where an induced conversion is exhibited as a genetic transition when compared to the vertebrate genome; andiii) determining the methylation pattern of a specific region or a plurality of genomic regions of interest; andh. determining cancer and disease risk prediction, cancer and disease diagnosis, cancer and disease prognosis, cancer and disease sub-classification, drug response prediction, MRD status, risk of disease recurrence, treatment planning, therapy selection and / or treatment efficacy and patient stratification for clinical trial enrollment of the tested sample by utilizing the information from step (g).
2. Method according to claim 1, wherein a TAC oligonucleotide that binds the target's plus or minus strand can be modified by converting all its cytosines into uracils or by converting only a portion of its cytosines into uracils.
3. Method according to claims 1 or 2, wherein the modification of the TAC oligonucleotides involves converting one or more of the unmethylated cytosines.
4. Method according to claims 1 to 3, whereina. each TAC oligonucleotide is between 150-350 nucleotides in length;b. each TAC oligonucleotide has a 5' end and a 3' end;c. each TAC oligonucleotide has a GC content of between 25% and 85%;d. each TAC oligonucleotide binds at least 50 base pairs away from regions harbouring segmental duplications or repetitive DNA elements; and / ore. there are 30 or more different TAC oligonucleotides used in the method.
5. Method according to claims 1 to 4, wherein the conversion of cytosines of the genomic DNA and / or the TAC oligonucleotides is performed enzymatically and / or chemically.
6. Method according to claims 1 to 5, wherein one or more of the TAC oligonucleotides has been modified only partially and not all cytosines are converted.
7. Method according to claim 6, wherein 100%, 80%, 70%, 60% 50 %, or less than 50% of the cytosines are converted.
8. Method according to claim 6, wherein different proportions of converted to unconverted and / or partially converted TAC oligonucleotides are used to enrich the plurality of target sequences including but not limited to 1:1:0, 1:0:1, 2:2:1, 0:1:1, 3:2:1, 2:1:2, 1:3:2.
9. Method according to claims 1 to 8 wherein, the fragments which have been enriched are amplified before they are sequenced.
10. Method according to claims 1 to 9, wherein the amplification of converted TAC oligonucleotides generates additional TAC oligonucleotide sequences leading to more optimal hybrid capture efficiency besides improving the efficacy of TAC oligonucleotide production.
11. Method according to claims 1 to 10, wherein the DNA sample is a vertebrate sample, preferably a human sample, preferably a blood sample, and more preferably a plasma sample.
12. Method according to claims 1 to 11, wherein the DNA sample is from a group comprising a peripheral blood sample, a plasma sample, a urine sample, a stool sample, a sputum sample, a cerebrospinal fluid sample, an ascites sample, a tear sample, a sweat sample, a saliva sample, a pleural effusion sample, a bronchoalveolar lavage sample, a buffy coat and / or a solid tissue sample.
13. Method according to claims 1 to 12, wherein the sample comprises cell-free DNA (cfDNA) or cell-free RNA (cfRNA) and the cfDNA or cfRNA is the nucleic acid preferentially analysed.
14. Method according to claims 1 to 13, wherein the DNA sample is combined with spike-in control DNAs comprising fully methylated or unmethylated cytosines, preferably from non-vertebrate genomes, wherein the TAC pool is also enriched by TAC oligonucleotides targeting the spike-in control DNAs.
15. Method according to claims 1 to 14, wherein adapters are added to the genomic DNA sample prior to a conversion of cytosines to uracil, wherein the cytosines of the adapter are protected from conversion or the conversion from cytosine to uracil is taken into account.
16. Method according to claims 1 to 15, wherein the TAC oligonucleotides contain a biotin modification wherein the biotin is incorporated during the generation of sequences that can hybridize the target regions or added through ligation of a DNA adapter containing a biotin modification.
17. Method according to claims 1 to 16, wherein determining the methylation pattern of the tested sample comprises steps selected from the group of:a. a demultiplexing procedure whereby sequenced reads are assigned to their samples, as identified via an index sequence, producing FASTQ. files;b. running an aligner that maps the sequenced reads on the genomic coordinates that have the highest mapping likelihood, said likelihood is calculated from mapping in both the converted and non-converted versions of the human genome;c. processing the aligned sequences using a software suite that determines cytosine methylation status from converted sequence reads;d. extracting the number of methylated and non-methylated cytosines for each target region;e. applying an exact binomial test to calculate the confidence interval for the methylation level of each CpG site;f. computing reference CpG methylation levels for various vertebrate tissues;g. quantifying the methylation level strand bias and / or the proportion of methylated cytosines (not CpG sites) and / or the coefficient of variation of each CpG;h. computing the depth of coverage (sequencing read depth) for each targeted CpG;i. computing the CpG methylation status pattern on each sequenced fragment from a plurality of sequenced fragments, said sequenced fragments being defined as pairs of sequenced reads; with said CpG methylation pattern being defined by the binary status (0: unmethylated; 1: methylated) of each CpG on said fragment; andj. determining cancer and disease risk prediction, cancer and disease prognosis, cancer and disease sub-classification, the tissue of origin for each sequenced fragment, drug response prediction, MRD status, treatment planning, therapy selection and / or treatment efficacy of the tested sample by utilizing the information from steps g) to i).
18. Method according to claims 1 to 17, wherein the method can also be applied for the determination of the combined methylation, hydroxymethylation or other episignatures' status of a single or a plurality of cytosine bases in DNA molecules of a vertebrate sample.
19. A kit for performing the method of claim 1, wherein the kit comprises a container comprising:a) the pool of TAC oligonucleotides wherein:i) each TAC oligonucleotide is between 150-350 nucleotides in length, each member sequence having a 5 end and a 3' end;ii) each TAC oligonucleotide binds to the sequence of interest at least 50 base pairs away, on both the 5' end and the 3' end, from regions harboring segmental duplications or repetitive DNA elements;iii) the GC content of the TAC oligonucleotides is between 25% and 85%;iv) the pool of TAC oligonucleotides comprises 30 or more distinct sequences;v) the pool of TAC oligonucleotides comprises converted TAC oligonucleotides alone or in a mixture including partially converted and / or non-converted TAC oligonucleotides;b) one or more components for isolating DNA from a sample;c) one or more components for preparing the converted sequencing library;d) one or more components for enriching the converted sequencing library;e) software and instructions for performing the statistical method for the determination of the methylation pattern of a specific region ora plurality of genomic regions of interest in a sample and clinical assessment of disease in the sample.