Methylation markers for coronary heart disease
Novel methylation markers at specific genomic loci enhance CHD risk prediction in T2D patients by accounting for genetic and environmental interactions, offering improved diagnostic accuracy and personalized treatment options.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE CHINESE UNIVERSITY OF HONG KONG
- Filing Date
- 2025-10-23
- Publication Date
- 2026-04-30
AI Technical Summary
Current methods for predicting the risk of coronary heart disease (CHD) in individuals with type 2 diabetes (T2D) are inadequate, as existing biomarkers lack replication in large-scale studies and fail to account for the interaction between genetic inheritability and environmental factors.
Identification of novel methylation markers at specific genomic loci, such as 13: 86824977-86825479, 3: 145957813-145958340, and others, to assess CHD risk by determining the methylation status of genomic DNA using reagents that differentially interact with methylated and non-methylated DNA, followed by sequencing or amplification.
The methylation markers provide a robust prediction of CHD risk, outperforming clinical risk factors with an improved AUC, enabling early medical intervention and personalized treatment strategies.
Smart Images

Figure CN2025129485_30042026_PF_FP_ABST
Abstract
Description
METHYLATION MARKERS FOR CORONARY HEART DISEASEBACKGROUND OF THE INVENTION
[0001] As living standards continue to improve globally, the number of individuals who are overweight or even obese is also rapidly increasing, which in turn has led to a notably higher incidence of many diseases including diabetes (especially type 2 diabetes, or T2D) , heart disease, hypertension, stroke, and the like. In developed nations, coronary heart disease (CHD) and cardiovascular disease (CVD) together account for the most frequent cause of mortality and morbidity, especially among patients who have already developed T2D. In clinical practice, however, it is difficult to identify which patients are more at risk of developing CHD so as to make early medical intervention available to them, especially among T2D patients. Several risk engines based on epidemiological studies are identified clinical risk factors exist, though the prediction performance are in general modest, with AUC around 0.75-0.8. Several novel biomarkers of cardiovascular risk have also been proposed, though these often lack replication in large-scale studies. Polygenic risk scores (PRS) , constructed from large-scale genetic studies, have also become increasingly popular, though they have limitations in reflecting only inherited risk and are fixed once an individual is bom. Methylation markers capture the interaction of the genetic inheritability with environmental factors and other modifiers including hyperglycemia, diet, and medication, and hence may be more suitable biomarkers to predict risk of adverse clinical outcomes and complications. This study identifies novel blood methylation markers for assessing risk of CHD, and these markers, with or without additional information on clinical risk factors, provide robust prediction of risk for coronary heart disease among individuals with type 2 diabetes. The performance of the methylation model is significantly better than that achieved using clinical risk factors, indicating its clinical utility as prognostic biomarkers for precision medicine in diabetes.
[0002] Because of the prevalence of T2D and CHD as well as their enormous social and economic impact globally, there exists a desire for new and more effective methods to diagnose, monitor, and treat CHD, especially among T2D patients who have a known propensity of developing cardiovascular diseases.BRIEF SUMMARY OF THE INVENTION
[0003] The present inventors have identified several genomic loci as novel biomarker for the risk assessing, diagnosis, and / or prognosis of coronary heart disease (CHD) , especially among patients who have already been diagnosed as suffering from type 2 diabetes (T2D) . More specifically, the inventors show that, compared with individuals without CHD, CpG islands of these genomic regions are less methylated in the genomic DNA obtained from blood cells taken from patients already suffering from CHD or known to have an increased risk of developing CHD at a later time.
[0004] Thus, in the first aspect, the present invention provides a method for assessing the methylation status of genomic DNA sequence comprising these steps: (a) contacting genomic DNA from blood cells taken from the subject with a reagent that differentially interacts with methylated and non-methylated DNA; (b) analyzing at least a portion of a genomic DNA sequence at a genomic locus selected from: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, that harbors a plurality of CpGs; and (c) determining nnmber of methylated CpGs among the plurality of CpGs. In an example embodiment, step (b) comprises analyzing at least a portion of a genomic DNA sequence at a genomic locus selected from Table 11.
[0005] In some embodiments, the portion of the genomic sequence being sequenced is at a genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663. In some embodiments, the method further comprises, prior to step (a) , isolating blood cells from a blood sample taken from the subject and then isolating genomic DNA from the blood cells. In some embodiments, the subject has been diagnosed with type 2 diabetes (T2D) or has known risk for T2D, e.g., due to a family history of diabetes or elevated blood sugar level (albeit not reaching a diagnostic threshold for diabetes) . In some embodiments, the analyzing in step (b) comprises (1) sequencing of the at least a portion of the genomic DNA or (2) amplification of the methylated or unmethylated version of the at least a portion of the genomic DNA. In some embodiments, wherein sequencing is encompassed in step (b) , this step may further comprise fragmentation of the genomic DNA prior to the sequencing. For example, the fragmentation of the genomic DNA may be achieved by sonication of the genomic DNA. In some embodiments, the fragmentation step generates genomic DNA fragments from about 100 to about 300 nucleotides in length, for example, about 200 or 250 nucleotides. In some embodiments, step (b) of the method comprises amplification of the methylated genomic DNA after enrichment and prior to the sequencing. For example, the amplification may be achieved by a polymerase chain reaction (PCR) . In some embodiments, the reagent that differentially interacts with methylated and non-methylated DNA comprises a bisulfite or a protein or protein conjugate comprising a protein or protein fragment that includes a methyl-CpG-binding domain (MBD) . In some embodiments, the method further comprises, after step (c) , comparing the number of methylated CpGs from step (c) with a standard control and identifying the subject as having coronary heart disease (CHD) or having an increased risk for CHD, upon determining the number ofmethylated CpGs from step (c) as less than the standard control, which reflects the number of methylated CpGs in the same portion of the genomic sequence at the same genomic locus found in the blood cells obtained by the same procedure from a healthy control subject, i.e., an average healthy individual who has no CHD nor any known elevated risk for CHD. In some embodiments, the method further comprises repeating steps (a) - (c) using a second preparation of genomic DNA from blood cells taken from the subject obtained by the same procedure at a later time, wherein an increase in the nnmber ofmethylated CpGs found in the genomic sequence at the later time as compared to the number ofmethylated CpGs from the original step (a) indicates a lessened risk for CHD for the subject over the intervening time, whereas a decrease indicates a heightened risk for CHD. In some embodiments, when the subject is identified as having CHD, having an increased risk for CHD, or having a worsening risk for CHD over an intervening time, further includes a step of administering to the subject an antiplatelet drug, a cholesterol-lowering drug, a blood pressure-lowering drug, nitroglycerin, a calcium channel blocker, or a β-blocker for the purpose of reducing CHD risk or treating CHD as deemed appropriate, e.g., by an attending physician.
[0006] In the second aspect, the present invention provides a method for assessing the relative risk for CHD among individuals who have not yet been diagnosed with CHD or for assessing CHtD progression among individuals who have already been diagnosed with CHD. The method includes these steps: (a) contacting genomic DNA from blood cells, taken from two subjects whose CHD risk is being assessed, with a reagent that differentially interacts with methylated and non-methylated DNA; (b) for the genomic DNA from both subjects, analyzing at least a portion of a genomic DNA sequence at a genomic locus selected from: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, this portion comprising a plurality of CpGs; (c) for both subjects, determining the number of methylated CpGs among the plurality of CpGs; (d) comparing the number of methylated CpGs from step (c) between the two subjects; and (e) determining the first subject who has more methylated CpGs in the portion of the genomic sequence as having a lower risk for CHD or as having CHD in a less progressed state when compared to the second subject. In an example embodiment, step (b) comprises analyzing at least a portion ofa genomic DNA sequence at a genomic locus selected from Table 11.
[0007] In some embodiments, the portion of the genomic sequence being sequenced in step (b) is at a genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663. In some embodiments, the method further comprises, prior to step (a) , isolating blood cells from a blood sample taken from the two subjects and then isolating genomic DNA from the blood cells. In some embodiments, the two subjects have been diagnosed with type 2 diabetes (T2D) . In some embodiments, the analyzing in step (b) comprises (1) sequencing of the at least a portion of the genomic DNA or (2) amplification of the methylated or unmethylated version of the at least a portion of the genomic DNA. In some embodiments, wherein sequencing is encompassed in step (b) , this step may further comprise fragmentation of the genomic DNA prior to the sequencing. In some embodiments, the fragmentation of the genomic DNA comprises sonication of the genomic DNA, for example, generating genomic DNA fragments in the length range of about 100 to about 300 nucleotides, for example, about 200 or about 250 nucleotides. In some embodiments, the method comprises amplification of the methylated genomic DNA after the enrichment step and prior to the sequencing step. The amplification may be achieved by a polymerase chain reaction (PCR) , for example. In some embodiments, the reagent that differentially interacts with methylated and non-methylated DNA comprises a bisulfite or a protein or protein conjugate comprising a protein or protein fragment that includes a methyl-CpG-binding domain (MBD) .
[0008] In the third aspect, a kit for detecting CHD or assessing CHD risk or monitoring CHD progression in a subject. The kit includes (1) a standard control that provides the number of methylated CpGs within a portion ora genomic DNA at a genomic locus selected from: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, that comprises a plurality of CpGs, the number reflecting the number of methylated CpGs found in blood cells taken from an average healthy individual who does not have CHD nor any known risk for CHD; (2) a reagent that differentially interacts with methylated and non-methylated DNA; and (3) an agent useful for analyzing the portion of the genomic DNA. In some embodiments, the reagent that differentially interacts with methylated and non-methylated DNA may comprise a bisulfite or a protein or protein conjugate comprising a protein or protein fragment that includes a methyl-CpG-binding domain (MBD) . In some embodiments, the agent in (3) is useful for sequencing the portion of the genomic DNA or for amplifying the methylated or unmethylated version of the portion of the genomic DNA. In some embodiments, where the agent in (3) is an agent for sequencing the portion of the genomic DNA, the kit may further comprise an agent useful for a DNA amplification reaction, for example, a polymerase chain reaction (PCR) . In some embodiments, the standard control provides the number ofmethylated CpGs found in an average healthy control subject's genomic sequence at the genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663. In an example, the genomic locus is selected from Table 11. Optionally, the kit in some cases may further include an instruction manual to provide instructions for the users.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1. HKDR methylation profiling design and analysis workfiow. Fig. 1 (a) Discovery of Differentially Methylated Regions (DMRs) as biomarkers for assessing risk of Coronary Heart Disease (CHD) in individuals with Type 2 Diabetes (T2D) . This study leveraged a clinical workflow for establishing a prospective cohort since 1995 as a quality-improvement program with detailed documentation of risk factors and clinical outcome. Cases and controls were selected to investigate the differential methylation patterns associated with incident and prevalent CHD in individuals with T2D from Hong Kong Diabetes register (HKDR) . Genome-wide DNA methylation analysis was conducted to identify Differentially Methylated Regions (DMRs) in blood samples using methyl-capture coupled with deep sequencing. The DMRs were identified using edgeR for various case-control groups including Non-diabetes control (NDC) vs Diabetes groups. The latter included patients with T2D without CHD after at least 10 years of disease duration (DC) and those with incident CHD (DI; individuals who developed CHD during study follow-up) and prevalent CHD (DP; individuals with CHD at recruitment to HKDR) . The inventors employed Receiver Operating Characteristic (ROC) analysis and Uniform Manifold Approximation and Projection (UMAP) on the DMRs to characterise DNA methylation-based biomarkers for classifying the risk of CHD, for patient stratification and early prevention for onset or recurrent CHD. Fig. 1 (b) Diagram for methylation biomarkers detected in various contrast analysis. Case-control groups included (i) DI (incident CHD in T2D) vs DC (T2D without CHD) ; (ii) DP vs DC (prevalent CHD in T2) vs DC) , (iii) DI+DP vs DC, and (iv) T2D cases vs NDC (non-diabetes control) . The Veun diagram illustrates the intersection of DMRs for each contrast. The highlighted overlapping area represents methylation changes exclusively detected in individuals with CHD, indicating their relevance to classifying incident and prevalent CHD (DI vs DC, DP vs DC, and DI+DP vs DC comparisons) . Overall, 20 methylation biomarkers were identified for downstream assessments including overlap with TFBS. Fig. 1 (e) Summary ofmethyl-TFBS associated with gain or loss ofmethylation at candidate biomarkers.
[0010] Figure 2. Accuracy of methylation biomarkers for incident and prevalent CHD in T2D. Fig. 1 (a) Combinatorial ROC analysis by generating a list of combinations comprising ofmethylation biomarkers (5mC) and clinical risk factors (CRf). The combined area under the curve (AUC) for these combinations were computed to measure the performance of DNA methylation in estimating risk and progression of CHD (DI and / or DP) compared with the DC group. In the depicted ROC plot, the green line represents the ROC curve (AUC) and accuracy (ACC) for the best combination of DMRs annotated to CHD2, DACH1, TRAM1L1, AKR1B1, and Intergenic region (DMR2; 13: 104486307-104486663) . The orange line represents the AUC / ACC of CRf including LDL-C, total cholesterol and eGFR at baseline for diagnosing CHD risk and progression. The red line represents the model combining the methylation biomarkers and CRf. P-values were computed using bootstrap in R. Fig. 2 (b-d) Clinical utility of indices for classifying CHD risk using clinical risk factors (CRf) , methylation biomarkers (5mC) and their combination (5mC + CRf) , respectively. The plot illustrates the prediction probabilities of False negative (FN) , False positive (FP) , True negative (TN) , and True positive (TP) classifications for those with and without CHD (DI, DP, and DC) using the optimal cutoff obtained from the corresponding ROC curve analysis. The optimal cutoff was calculated using the ROCR package prediction function. The optimal cut-off for the CRf is shown as a red line was 0.60 (CRf) , 0.64 (5mC) 0.64, and 0.52 (5mC + CRt) .
[0011] Figure 3. Systematic Workflow for Automated Methyl-CpG Enrichment and Sequencing. Step-by-step workflow of methyl-CpG enrichment including automated processes on the Tecan Freedom EVO 200 Liquid Handling platform. Sample gDNA was fragmented by sonication and subjected to methyl-CpG enrichment using components from the MethylMinerTM Kit (ME10025, ThermoFisher) in a 96-well format. The main reason for automating “Methylminer” was to efficiently manage a large number of clinical samples while minimising any technical variability associated with manual handling. This setup allows for the simultaneous processing of up to 48 samples within a two-day timeframe, including quality control assessments, prior to downstream library preparation and high-throughput sequencing.
[0012] Figure 4. Tecan Freedom Evo 200 configuration for methyl-seq automatation. The figure illustrates the layout of the Tecan Freedom Evo 200 worktable showing various components and accessories integral to automated methyl-seq. Tecan components utilised include, Robot arms: Integrated 8-channel liquid handling arm (LiHa) with additional 8-channel large volume dispense technology (Te-Fill technology) ; Robotic manipulator arm, standard z-rail (RoMa) ; Labware carriers: Low-profile 96-well microplate carriers (x 3) ; Device carriers: Te-ShakeTM Microplate Shaker; Magnetic plate: Agencourt SPRIPlate 96R Super Magnet Plate (Beckman-Coulter) .
[0013] Figure 5. DNA methylation clustering before and after GLM for clinical covariates. Panel 1: Uniform Manifold Approximation and Projection (UMAP) of unadjusted CHD methylation sequencing data obtained from blood DNA. Each symbol represents DNA methylation profile of each sample categorized by Non diabetes control (NDC, n=51) , Diabetes without CHD control (DC, n=55) , Diabetes with CHD incidence (DI, n=58) and Diabetes with CHD prevalence (DP, n=55) . Panel 2: UMAP plot of adjusted CHD methylation sequencing data clearly shows the importance of using generalised linear modelling (GLM) analyses to adjust for age variability and gender, disease duration, experimental covariates, and technical variability such as cell type and CpG enrichment ratios. The influence of clinical eovariates was assessed through regression modelling between the top principal components (UMAP 1 and 2) and the imputed covariates, revealing their impact on genome-wide DNA methylation.
[0014] Figure 6. HKDR genomic methylation indices. Fig. 6 (a) Graphical representation of genomic features and CpG islands, shelves and shores. Fig. 6 (b) Distribution of methylated peaks overlap with genomic features, including gene promoter regions, exons, coding exons, introns, and intergenic regions. Majority of DNA methylation differences are concentrated within gene intronic regions and intergenic regions between genes. Fig. 6 (c) Distribution of differential methylation (%) , indicating gain and loss of methylation at DMRs identified in various case-control groups using edgeR with a significance threshold of P <0.001 in T2D individuals with or without CHD at recruitment. The numbers above the bars indicate the total DMRs detected, indicating gain and loss of methylation. These methylation data were analysed in 219 individuals, including 51 non-diabetes controls (NDC) , 55 T2D individuals without CHD (DC) , 58 T2D individuals with incident CHD (Di) , and 55 T2D individuals with prevalent CHD (DP) . The number of DMRs across the DI and DP groups, either compared to NDC or DC groups are shown. Loss of methylation was more evident in the DI and DP groups compared to the DC group.
[0015] Figure 7. Discriminative performance of Methylation Biomarkers for diagnosing CHD. Fig. 7 (a) DMRs were identified using edgeR, with a significance threshold ofP < 0.001, in individuals with or without CHD and T2D from the HKDR. Clinical data were collected from 219 individuals, including 51 non-diabetes controls (NDC) , 55 T2D individuals without CHD (DC) , 58 T2D individuals with incident CHD after recruitment (DI) , and 55 T2D individuals with prevalent CHD at recruitment (DP) . Baseline clinical risk factors for CHD were determined using the Modem Applied Statistics with S (MASS) package in R. The initial classification model for DI and DP group included age, sex, duration of diabetes, smoking (yes / no) , diastolic / systolic blood pressure, BMI, HbA1c, lipids, eGFR and UACR. The stepAIC() function was employed for selection of risk factors which included total cholesterol, LDL-C and eGFR. The performance of individual DMRs and clinical risk factors in classifying CHD events in DI and DP individuals was evaluated using the Area Under the ROC curve (AUC) . Potential methylation biomarkers were selected based on specific criteria: AUC greater than 0.65 per contrast analysis, DMRs with more than 4 CG sites, and exclusivity for CHD (excluding DMRs detected in all T2D vs NDC) . Next, the inventors performed combiuatorial ROC analysis to generate a list of combinations for methylation biomarkers and clinical risk factors. The combined AUC ranges for different combinations were computed, and the best combination was identified. Subsequently, the clinical utility of the best combination was assessed for CHD risk prediction. Fig. 7 (b) Enrichment analysis of methylation biomarkers interaction with Transcription Factor binding Sites (TFBS) . Plot shows differential enrichment of direct interactions between candidate methylation biomarkers and TFBS. Intersects associated with a gain of methylation are represented in blue, while regions with a loss of methylation are shown in light blue. The enrichment analysis was conducted using R package LOLA, and differential enrichment scores were computed based on odd ratios using a list of background regions compared to methylation biomarkers. Fig. 7 (c) and (d) UMAP clustering of methylation biomarkers in T2D groups with or without CHD. The UMAP plots show the clustering of samples based on the methylation signals of the 20-candidate biomarkers, distinguishing between different diabetes groups.
[0016] Figure 8. DNA Methylation signal of candidate biomarkers. Fig. 8A. Methyl-biomarkers 1-10. Fig 8B. Methyl-biomarkers 11-20. Boxplots depicting the DNA methylation signal in individuals with T2D, both with and without CHD. The methylation signal is normalized to the input DNA and represented by the vertical axis. Each boxplot corresponds to a specific methylation biomarker. The presence of black points indicates the outliers within each group. All the methyl-biomarkers have a P <0.001 comparing DI and DP versus DC group. The comparison between the groups provides insights into the differential methylation patterns associated with CHD in individuals with diabetes.
[0017] Figure 9. The combinatorial ROC analysis for CHD integrates AUC scores for the 20 methylation biomarkers (5mC) and clinical risk factors (CRf) to assess diagnostic performance and accuracy. C (n, r) is the number of combinations, where (n) is the total number of set elements and (r) is the number of elements in each set with factorial notation (! ) . AUC scores were ranked for all combinations to distinguish biomarker accuracy and clinical utility.DEFINITIONS
[0018] As used herein, the term "coronary heart disease (CHD) " or "coronary artery disease (CAD) " refers to any heart disease involving restricted / reduced blood flow to the cardiac muscle due to build-up of atherosclerotic plaque in the heart arteries. Also known as ischemic heart disease or myocardial ischemia, CHD is the most common type among cardiovascular diseases.
[0019] The term “nucleic acid” or “polynucleotide” refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single-or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) , alleles, orthologs, single nucleotide polymorphisms (SNPs) , and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19: 5081 (1991) ; Ohtsuka et al., J. Biol. Chem. 260: 2605-2608 (1985) ; and Rossolini et al., Mol. Cell. Probes 8: 91-98 (1994) ) . The term nucleic acid is used interchangeably with gene, eDNA, and mRNA encoded by a gene.
[0020] The term “gene” means the segment of DNA involved in producing a polypeptide chain; it includes regions preceding and following the coding region (leader and trailer) involved in the transeription / translation of the gene product and the regulation of transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons) .
[0021] In this application, the terms “polypeptide, ” “peptide, ” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which at least one (potentially more) amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. As used herein, the terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens) , wherein the amino acid residues are linked by covalent peptide bonds.
[0022] In this disclosure the term "biological sample" or “sample” includes sections of tissues such as biopsy and autopsy samples, and frozen sections taken for histologic purposes, or processed forms of any of such samples. Biological samples include blood and blood fractions or products (e.g., acelhilar fractions such serum or plasma, cellular fractions such as all blood cells, certain specific blood cell types such as red blood cells, white blood cells, platelets, and the like) , sputum or saliva, lymph and tongue tissue, cultured cells, e.g., primary cultures, explants, and transformed cells, stool, urine, esophagus biopsy tissue etc. A biological sample is typically obtained from a eukaryotic organism, which may be a mammal, may be a primate, and may be a human subject.
[0023] In this disclosure the term "isolated" nucleic acid molecule means a nucleic acid molecule that is separated from other nucleic acid molecules that are usually associated with the isolated nucleic acid molecule. Thus, an "isolated" nucleic acid molecule includes, without limitation, a nucleic acid molecule that is free of nucleotide sequences that naturally flank one or both ends of the nucleic acid in the genome of the organism from which the isolated nucleic acid is derived (e.g., a cDNA or genomic DNA fragment produced by PCR or restriction endonuclease digestion) . Such an isolated nucleic acid molecule is generally introduced into a vector (e.g., a cloning vector or an expression vector) for convenience of manipulation or to generate a fusion nucleic acid molecule. In addition, an isolated nucleic acid molecule can include an engineered nucleic acid molecule such as a recombinant or a synthetic nucleic acid molecule. A nucleic acid molecule existing among hundreds to millions of other nucleic acid molecules within, for example, a nucleic acid library (e.g., a cDNA or genomic library) or a gel (e.g., agarose, or polyacrylamine) containing restriction-digested genomic DNA, is not an "isolated" nucleic acid.
[0024] The term "bisulfite" as used herein encompasses all types ofbisulfites, such as sodium bisulfite, that are capable of chemically converting a cytosine (C) to a uracil (U) without chemically modifying a methylated cytosine and therefore can be used to differentially modify a DNA sequence based on the methylation status of the DNA.
[0025] As used herein, a reagent that "differentially interacts with" methylated and non-methylated DNA encompasses any reagent that reacts or modifies differentially with methylated and unmethylated DNA in a process through which distinguishable products or quantitatively distinguishable results (e.g., degree of binding or precipitation; relative concentration) are generated from methylated and non-methylated DNA, thereby allowing the identification of the DNA methylation status. Such processes may include, but are not limited to, chemical reactions (such as an unmethylated C → U conversion by bisulfite) , enzymatic treatment (such as cleavage by a methylation-dependent endonuclease) , binding, precipitation, and selective enrichment of methylated or unmethylated version of DNA. Thus, an enzyme that preferentially cleaves methylated DNA is one capable of cleaving a DNA molecule at a much higher efficiency when the DNA is methylated, whereas an enzyme that preferentially cleaves unmethylated DNA exhibits a significantly higher efficiency when the DNA is not methylated. In the context of the present invention, a reagent that "differentially interacts with" methylated and unmethylated DNA also refers to any reagent that exhibits differential ability in its binding to DNA sequences or precipitation of DNA sequences depending on their methylation status. One class of such reagents consists of full-length, conjugates, or functional fragments of methylated DNA binding proteins, which may have been derived from different species, including mammalian species such as human.
[0026] A "CpG-containing genomic sequence" or "genomic sequence comprising a plurality of CpGs" as used herein refers to a segment of genomic DNA sequence at a defined location in the genome of an individual within which multiple CpG di-nucleotide pairs are present. Typically, a "CpG-containing genomic sequence" is at least 15 contiguous nucleotides in length and contains at least two CpG pairs and likely more CpG pairs. In some cases, it can be at 1east 18, 20, 25, 30, 50, 80, 100, 150, 200, 250, 300, 500, 800, 1000, 2000, 3000, 5000, 8000, 10,000, 20,000, 30,000, 50,000, 100,000 or more contiguous nucleotides in length and contains at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 50, 80, 100, 200, 500, or more CpG palrs. For any one "CpG-containing genomic sequence" at a given location, e.g., within a genomic locus named in Table 11, especially at the genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663, nucleotide sequence variations may exist from individual to individual and from allele to allele even for the same individual. Furthermore, a "CpG-containing genomic sequence" may encompass a nucleotide sequence transcribed or not transcribed for protein production, and the nucleotide sequence can be a protein-coding sequence, a non protein-coding sequence (such as a transcription promoter) , or a combination thereof.
[0027] As used in this application, an "increase" or a "decrease" refers to a detectable positive or negative change in quantity from a comparison control, e.g., an established standard control (such as the average number ofmethylated CpGs found in a genomic sequence at a pre-selected genomic locus as determined in the blood cells of a healthy control subject who does not have CHD nor any known risk for CHD) . An increase is a positive change that is typically at least 10%, or at least 20%, or 50%, or 90%, or 100%, and can be as high as at least 2-fold or at least 5-fold or even 10-fold of the standard control value. Similarly, a decrease is a negative change that is typically at least 10%, or at least 20%, 30%, or 50%, or even as high as at least 80%or 90%of the control value. Other terms indicating quantitative changes or differences from a comparative basis, such as "more, " "less, " "higher, " and "lower, " are used in this application in the same fashion as described above. In contrast, the term "substantially the same" or "substantially lack of change" indicates little to no change in quantity from the standard control value, typically within ± 10%of the standard control, or within ± 5%, 2%, or even less variation from the standard control value.
[0028] A "polynucleotide hybridization method" as used herein refers to a method for detecting the presence and / or quantity of a pre-determined polynucleotide sequence based on its ability to form Watson-Crick base-pairing, under appropriate hybridization conditions, with a polynuclentide probe of a known sequence. Examples of such hybridization methods include Southern blot, Northern blot, and in situ hybridization.
[0029] "Primers" as used herein refer to oligonucleotides that can be used in an amplification method, such as a polymerase chain reaction (PCR) , to amplify a nucleotide sequence based on the polynucleotide sequence corresponding to a gene of interest, e.g., the cDNA or genomic sequence at a pre-determined genomic locus or a portion thereof. Typically, at least one of the PCR primers for amplification of a polynucleotide sequence is sequence-specific for that polynucleotide sequence. The exact length of the primer will depend upon many factors, including temperature, source of the primer, and the method used. For example, for diagnostic and prognostic applications, depending on the complexity of the target sequence, the oligonucleotide primer typically contains at least 10, or 15, or 20, or 25 or more nucleotides, although it may contain fewer nucleotides or more nucleotides. The factors involved in determining the appropriate length of primer are readily known to one of ordinary skill in the art. In this disclosure the term "primer pair" means a pair of primers that hybridize to opposite strands a target DNA molecule or to regions of the target DNA which flank a nucleotide sequence to be amplified. In this disclosure the term "primer site" means the region of the target DNA or other nucleic acid to which a primer hybridizes.
[0030] A "label, " "detectable label, " or "detectable moiety" is a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include32p, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in an ELISA) , biotin, digoxigenin, or haptens and proteins that can be made detectable, e.g., by incorporating a radioactive component into the peptide or used to detect antibodies specifically reactive with the peptide. Typically, a detectable label is attached to a probe or a molecule with defined binding characteristics (e.g., a polypeptide with a known binding specificity or a polynucleotide) , so as to allow the presence of the probe (and therefore its binding target) to be readily detectable.
[0031] "Standard control" as used herein refers to the number of methylated CpGs or level or degree of methylation within a segment of a genomic DNA sequence at a pre-determined genomic locus, e.g., any one named in Table 11, especially at the genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663, that is present in the genomic DNA sequence from an established normal, CHD-free tissue sample, e.g., a normal all blood cell sample taken from an average, normal, healthy person who does not have CHD or any elevated risk for CHD. The standard control value is suitable for the use of a method of the present invention, to serve as a basis for comparing the number of methylated CpGs or the level of DNA methylation that is present in the same type of sample obtained from a test subject. An established sample serving as a standard control provides an average number of methylated CpGs, i.e., an average level of DNA methylation, that is typical for a particular genomic sequence at a particular genomic locus found in a particular sample type (e.g., all blood cells) obtained from an average, healthy human without any heart disease especially CHD as conventionally defined. A standard control value may vary depending on the nature of the sample (e.g., sample type and how sample has been processed) as well as other factors such as the gender, age, ethnicity, other medical condition (e.g., presence or absence of T2D) of the subjects based on whom such a control value is established.
[0032] The term "average, " as used in the context of describing a human who is healthy, free of any heart disease (especially CHD) as conventionally defined, refers to certain characteristics, especially the number of methylated CpGs, or the level of methylatiun, found within a segment of the genomic sequence at a specified genomic locus in the person′sbiological sample of a particular type, e.g., an all-blood cell sample, that are representative of a randomly selected group of healthy humans who are free of any heart diseases (especially CHD) . This selected group should comprise a sufficient number of human individuals such that the average number of methylated CpGs or level of DNA methylation in the specified genomic sequence reflects, with reasonable accuracy, the corresponding number ofmethylated CpGs or methylation level of that genomic sequence in the general population of healthy human subjects without CHD or with no known elevated risk for CHD. In addition, the selected group of humans generally have a similar age to that of a subject whose biological sample is tested for indication or risk of CHD. Moreover, other factors such as gender, ethnicity, medical history (e.g., whether or not a diagnosis of T2D has been received) are also considered and preferably closely matching between the profiles of the test subject and the selected group of individuals establishing the "average" value.
[0033] The term "amount" as used in this application refers to the quantity of a molecule of interest (for example, a polynucleotide of interest or a polypeptide of interest, e.g., methylated DNA sequence at a pre-determined locus) present in a sample. Such quantity may be expressed in the absolute terms, i.e., the total quantity of the molecule in the sample, or in the relative terms, i.e., the concentration of the molecule in the sample.
[0034] The term "treat" or "treating, " as used in this application, describes to an act that leads to the elimination, reduction, alleviation, reversal, or prevention or delay of onset or recurrence of any symptom of a relevant condition. In other words, "treating" a condition encompasses both therapeutic and prophylactic intervention against the condition.
[0035] The term "effective amount" as used herein refers to an amount of a given substance that is sufficient in quantity to produce a desired effect. For example, an effective amount of a blood pressure-lowing drag is the amount of said drug to achieve a desirable level of decrease in a recipient's blood pressure, such that the symptoms of an abnormal and unhealthy high blood pressure level are reduced, reversed, eliminated, prevented, or delayed of the onset in a patient who has been given the drug for therapeutic purposes. An amount adequate to accomplish this is defined as the "therapeutically effective dose. " The dosing range varies with the nature of the therapeutic agent being administered and other factors such as the route of administration, the patient's physical and medical condition, and the severity of the condition being treated.
[0036] The term "subject" or "subject in need of treatment, " as used herein, includes individuals who seek medical attention due to risk of, or actual suffering from, or progression of CHD. Subjects also include individuals currently undergoing therapy that seek manipulation of the therapeutic regimen. Subjects or individuals in need of treatment include those that demonstrate symptoms of CHD or are at risk of later developing CHD or its symptoms. For example, a subject in need of treatment includes individuals with a genetic predisposition or family history for CHD, those that have suffered relevant symptoms in the past, as well as those suffering from chronic or acute symptoms of the condition. A “subject in need of treatment” may be at any age of life and of any gender. In the context of diagnosis or treatment methods of this invention, when a subject is described as of a certain descent (e.g., Asian or Chinese) , the person is genetically identifiable due to the presence of genetic markers characteristics of the population native to that specific geographic location and / or whose ancestors have been residing in that geographic area for at least the last 8 generations or at least 200 years.
[0037] As used herein, the term “about” denotes a range of+ / -10%of a given value. For instance, “about 10” denotes a range of 10+ / -1, i.e., 9-11.DETAILED DESCRIPTION OF THE INVENTIONI. Introduction
[0038] Despite the rapid advancement in medical sciences and steady improvement in the treatment and maintenance of patients suffering from heart diseases, cardiovascular disease, especially coronary heart disease (CHD) , remains a significant health concern with grave implications in both developed countries as well as in developing countries. The insidious aspect of CHD is that, for many patients, the first definitive indication of their CHD is a potentially fatal heart attack. The benefits of receiving adequate effective and early medical intervention are therefore often diminished. As such, early detection of the presence of CHD or the risk of CHD is critical for improving patient prognosis and survival from this potentially deadly disease.
[0039] The present inventors set out to study changes in the methylation status of genomic DNA sequence among patients who have already been identified as suffering from type 2 diabetes (T2D) . This disclosure reports the results of such a discriminative biomarker study of circulating cells for the purpose of identifying methylation differences of coronary heart disease (CHD) in type 2 diabetes (T2D) . The present inventors developed an automated methylation enrichment and deep sequencing pipeline for CHD in Chinese patients with T2D (n=219) consisting of non-diabetic controls (NDC, n=51) , T2D subjects without CHD (DC, n=55) , T2D subjects with CHD before entering the Hong Kong Diabetes Register (HKDR) study (prevalent or DP, n=55) and T2D subjects that develop CHD during the HKDR study (incident or DI, n=58) . Liquid handling robotics was used to develop differentially methylated regions (DMRs) associated with CHD risk and progression. Diagnostic value and performance of methylation biomarkers were investigated with phenotypes and its interactions with total cholesterol, LDL and baseline eGFR in CHD.
[0040] The inventors identified five unique methylation biomarkers that annotate to CHD2 (15: 93021943-93022463) , DACH1 (13: 71548599-71548873) , TRAM1L1 (4: 116801311-116801910) , AKR1B1 (7: 134400095-134400659) , and DMR2, differentially methylated region 2 (13: 104486307-104486663) that classify CHD risk and progression. The AUC score for clinical risk factors alone (CRt) was 0.69 (P=0.002) for CHD in T2D. Methylation biomarkers improved the AUC score to 0.91 (P=3.3e-12) while the combined AUC score was 0.92 and accuracy score of 0.90 (P=1.4e-12) .
[0041] This study reveals novel genomic methylation indices that show diagnostic value on CHD risk in Chinese patients with T2D. These observations strengthen the role of DNA methylation derived from the HKDR study as a useful biomarker resource to the pathogenesis of CHD. Thus, this invention provides a method to detect the presence or risk of CHD based on changes in CpG methylation level within genomic sequence at certain specific genomic loci, especially among patients already diagnosed with T2D. The invention also provides a kit and device useful for practicing such a CHD detection method.II. General Methodology
[0042] Practicing this invention utilizes routine techniques in the field of molecular biology. Basic texts disclosing the general methods of use in this invention include Sambrook and Russell, Molecular Cloning, A Laboratory Manual (3rd ed. 2001) ; Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990) ; and Current Protocols in Molecular Biology (Ausubel et al., eds., 1994) ) .
[0043] For nucleic acids, sizes are given in either kilobases (kb) or base pairs (bp) . These are estimates derived from agarose or acrylamide gel electrophoresis, from sequenced nucleic acids, or from published DNA sequences. For proteins, sizes are given in kilodaltons (kDa) or amino acid residue numbers. Protein sizes are estimated from gel electrophoresis, from sequenced proteins, from derived amino acid sequences, or from published protein sequences.
[0044] Oligonucleotides that are not commercially available can be chemically synthesized, e.g., according to the solid phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Lett. 22: 1859-1862 (1981) , using an automated synthesizer, as described in Van Devanter et. al., Nucleic Acids Res. 12: 6159-6168 (1984) . Purification of oligonucleotides is performed using any art-recognized strategy, e.g., native acrylamide gel electrophoresis or anion-exchange high performance liquid chromatography (HPLC) as described in Pearson and Reanier, J. Chrom. 255: 137-149 (1983) .
[0045] The sequence of interest used in this invention, e.g., the polynucleotide sequence of a gene of interest, and synthetic oligonucleotides (e.g., primers) can be verified using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981) . III. Acquisition of Tissue Samples and Analysis of Genomic DNA
[0046] The present invention relates to analyzing the methylation pattern of genomic DNA found in a person's biological sample, especially blood cell sample, as a means to detect the presence of, to assess the risk of developing, and / or to monitor the progression or treatment efficacy of coronary heart disease (CHD) . Thus, the first steps of practicing this invention are to obtain a suitable tissue sample from a test subject and extract genomie DNA from the sample.A. Acquisition and Preparation of Tissue Samples
[0047] A biological sample is obtained from a person to be tested or monitored for CHD using a method of the present invention. Collection of a suitable tissue sample from an individual is performed in accordance with the standard protocol hospitals or clinics generally follow, such as during a blood draw or a biopsy. An appropriate amount of blood or other tissues is collected and may be stored according to standard procedures prior to further preparation.
[0048] The analysis of genomic DNA found in a patient′s tissue (e.g., blood cell) sample according to the present invention may be performed using, e.g., whole blood or cellular fraction thereof. The methods for preparing tissue samples for nucleic acid extraction are well known among those of skill in the art. For example, a subject′s blood cell sample should be first treated to disrupt cellular membrane so as to release nucleic acids contained within the cells.B. Detection of Methylation in Genomic Sequence
[0049] Methylation status of a segment of genomic sequence containing one or more CpG (cytosine-guanine dinucleotide) pairs is investigated to provide indication as to whether a test subject is suffering from heart disease (especially CHD) , whether the subject is at risk of developing heart disease (especially CHD) , or whether the subject's heart disease (especially CHD) is worsening or improving.
[0050] Typically, a segment of the genomic sequence that harbors one or more CpG dinucleotide pairs is analyzed for methylation pattern, for example, a portion of a genomic sequence at a genomic locus 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143 (named in Table 11) , e.g., at the genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663, used to determine how many of the CpG pairs within the portion of sequence are methylated and how many are not methylated. The sequence being analyzed should be long enough to contain at least 1 CpG dinucleotide pair, typically 2 or more CpG pairs, and detection of changes in the methylation status at the CpG site (s) is typically adequate indication of the presence of CHD or an elevated risk for CHID. The length of the sequence being analyzed is usually at least 100 or 200 contiguous nucleotides, and may be longer with at least 250, 300, 500, 1000, 2000, 3000, 4000, 5000, 10,000, 20,000, 50,000, 100,000 or more contiguous nucleotides. At least one, typically 2 or more, often 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 500, 1000 or more, CpG nucleotide pairs are present within the sequence. A decrease in the number ofmethylated CpG or decreased methylation level of a genomic DNA sequence as specified above indicates either the presence or an elevated risk for heart disease, especially CHD, in a test subject, when such a decrease is found when a comparison is made to a standard control value. For the purpose of determining the methylation level of a particular genomic sequence, a reagent that differentially interacts with a DNA sequence depending on the methylation status of the CpG within the sequence is often used. For example, bisulfite treatment followed by DNA sequencing is particularly useful, since bisulfite converts an unmethylated cytosine (C) to a uracil (U) while leaving methylated cytosines unchanged, allowing immediate identification through a DNA sequencing process. Optionally, an amplification process such as PCR is included after a step of bisulfite conversion or enrichment of methylated DNA by way of using a protein containing a methyl-CpG-binding domain and before the DNA sequencing.1. DNA Extraction and Treatment
[0051] Methods for extracting DNA from a biological sample are well known and routinely practiced in the art of molecular biology, see, e.g., Sambrook and Russell, supra. RNA contamination should be eliminated to avoid interference with DNA analysis. The DNA is then treated with a reagent capable of modifying or interacting with DNA in a methylation differential manner, i.e., different and distinguishable chemical structures will result from a methylated cytosine (C) residue and an unmethylated C residue following the treatment. For example, such a reagent reacts with the unmethylated C residue (s) in a DNA molecule and converts each unmethylated C residue to a uracil (U) residue, whereas the methylated C residues remain unchanged. This unmethylated C → U conversion allows detection and comparison of methylation status based on changes in the primary sequence of the nucleic acid. An exemplary reagent suitable for this purpose is bisulfite, such as sodium bisulfite. Methods for using bisuifite for chemical modification of DNA are well known in the art (see, e.g., Herman et al., Proc. Natl. Acad. Sci. USA 93: 9821-9826, 1996) .
[0052] As a skilled artisan will recognize, any other reagents that are unnamed here but have the same property of chemically (or through any other mechanism) modifying methylated and unmethylated DNA differentially can be used for practicing the present invention. For instance, methylation-specific modification of DNA may also be accomplished by methylation-sensitive restriction enzymes, some of which typically cleave an unmethylated DNA fragment but not a methylated DNA fragment, while others (e.g., methylation-dependent endonuclease McrBC) cleave DNA containing methylated cytosines but not unmethylated DNA. In addition, a combination of chemical modification and restriction enzyme treatment, e.g., combined bisulfite restriction analysis (COBRA) (Xiong et al. 1997 Nucleic Acids Res. 25 (12) : 2532-2534) , is useful for practicing the present invention. Other available methods for detecting DNA methylation include, for example, methylation-sensitive restriction endonucleases (MSREs) assay by either Southern blot or PCR analysis, methylation specific or methylation sensitive-PCR (MS-PCR) , methylation-sensitive single nucleotide primer extension (Ms-SnuPE) , high resolution melting (HRM) analysis, bisuifite sequencing, pyrosequencing, methylation-specific single-strand conformation analysis (MS-SSCA) , methylation-specific denaturing gradient gel electrophoresis (MS-DGGE) , methylation-specific melting curve analysis (MS-MCA) , methylation-specific denaturing high-performance liquid chromatography (MS-DHPLC) , methylation-specific microarray (MSO) . These assays can be either PCR analysis, quantitative analysis with fluorescence labelling, Southern blot analysis, or mass spectrometry analysis. Commercially available products include EpiTYPER (Agena Bioscience) , GoldenGate Assay (Illumina) , and Infinium MethylationEPIC Kit (Ilhimina) . Exemplary methylation sensitive DNA cleaving reagent such as restriction enzymes include AatII, AciI, AclI, AgeI, AscI, Asp718, AvaI, BhrP1, BceAI, BmgBI, BsaAI, BsaHI, BsiEI, BsiWI, BsmBI, BspDI, BsrFI, BssHII, BstBI, BstUI, ClaI, EagI, EagI-HFTM, FauI, FseI, FspI, HaeII, HgaI, HhaI, HinP1I, HpaII, Hpy99I, HpyCH4IV, KasI, MluI, NarI, NgoMIV, NotI, NotI-HFTM, NruI, Nt. B smAI, PaeR7I, PspXI, PvuI, RsrII, SacII, SalI, SalI-HFTM, SfoI, SgrAI, Sinai, SnaBI or TspMI.
[0053] Moreover, a class of proteins that preferentially binds methylated DNA via methyl-CpG sites has been identified in various species, including mammalian species such as human. These proteins, each containing a methyl-CpG-binding domain (MBD) , are particularly useful for selectively binding and therefore specifically enriching methylated genomic DNA fragments in connection with the subsequent use of next-generation sequencing technology-based analysis in practicing the method of this invention. See, e.g., Trimarchi et al., BMC Genomics. 2012; 13 (Suppl 8) : S6.2. Optional Amplification and Sequence Analysis
[0054] Following the modification or enrichment of DNA genomic sequence in a methylation-differential manner, the treated DNA is then subjected to sequence-based analysis, such that the methylation status of a selected genomic sequence may be determined. An amplification reaction is optional prior to the sequence analysis after methylation specific modification. A variety of polynucleotide amplification methods are well established and frequently used in research. For instance, the general methods of polymerase chain reaction (PCR) for polynucleotide sequence amplification are well known in the art and are thus not described in detail herein. For a review of PCR methods, protocols, and principles in designing primers, see, e.g., Innis, et al., PCR Protocols: A Guide to Methods and Applications, Academic Press, Inc. N. Y., 1990. PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems.
[0055] Although PCR amplification is typically used in practicing the present invention, one of skill in the art will recognize that amplification of the relevant genomic sequence may be accomplished by any known method, such as the ligase chain reaction (LCR) , transcription-mediated amplification, and self-sustained sequence replication or nucleic acid sequence-based amplification (NASBA) , each of which provides sufficient amplification.
[0056] Techniques for polynucleotide sequence determination are also well established and widely practiced in the relevant research field. For instance, the basic principles and general techniques for polynucleotide sequencing are described in various research reports and treatises on molecular biology and recombinant genetics. DNA sequencing methods routinely practiced in research laboratories, either manual or automated, can be used for practicing the present invention. Additional means suitable for detecting changes (e.g., C → U) in a polynucleotide sequence for practicing the methods of the present invention include but are not limited to mass spectrometry, primer extension, polynucleotide hybridization, real-time PCR, melting curve analysis, high resolution melting analysis, heteroduplex analysis, pyrosequencing, and electrophoresis. The more recently developed next-generation sequencing technology is a particularly powerful tool for analysis such as in the context of the method of this invention.IV. Establishing a Standard Control
[0057] In order to establish a standard control for practicing the method of this invention, a group of healthy persons free of any heart disease (especially any form of coronary heart disease) as conventionally defined is first selected. These individuals are within the appropriate parameters, if applicable, for the purpose of detecting, assessing risk for, and / or monitoring heart disease (especially CHD) using the methods of the present invention. Optionally, the individuals are of same gender, similar age, or similar ethnic background. In some cases, the individuals also share the common medical history, e.g., with or without being diagnosed with type 2 diabetes (T2D) .
[0058] The healthy status and medical conditions of the selected individuals can be confirmed by well established, routinely employed methods including but not limited to general physical examination of the individuals and general review of their medical history.
[0059] Furthermore, the selected group of healthy individuals must be of a reasonable size, such that the average number of methylated CpG sites or methylation level within a specific genomic sequence such as: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143 (named in Table 11) , especially at the genomic locus 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663, in the biological samples obtained from the group can be reasonably regarded as representative of the normal or average number / level found among the general population of healthy people. Preferably, the selected group comprises at least 10 human subjects, more typically at least 20 or more subjects.
[0060] Once an average value for the methylated CpG number within a particular genomic sequence is established based on the individual values found in each subject of the selected healthy control group, this average or median or representative value is considered a standard control. A standard deviation is also determined during the same process. In some cases, separate standard controls may be established for separately defined groups having distinct characteristics such as age, gender, ethnic background, or even a pertinent medical condition, for example, whether or not a T2D diagnosis has been given to the subjects.V. Diagnosis and Treatment of Coronary Heart Disease
[0061] By illustrating the correlation of a reduced level of methylation within certain genomic DNA sequence (such as any one located at a genomic locus 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, named in Table 11, especially at the genomic locus 15: 93021943-93022463, 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663) and the presence or risk of heart disease such as CHD, the present invention provides a new and effective diagnostic method, which may achieve improved efficacy when used in connection with conventional diagnostic techniques for CHD, including assessing risk factors such as high blood pressure, smoking, diabetes, lack of exercise, obesity, high blood cholesterol, poor diet, depression, and excessive alcohol consumption for individuals without any heart disease symptoms, or performing blood tests (e.g., for cardiac troponins) , cardiac stress tests, echocardiography, or electrocardiogram (ECG / EKG) , computed tomography angiography (CTA) , coronary angiogram, positron emission tomography (PET) , single-photon emission computed tomography (SPECT) , intravascular ultrasound, or magnetic resonance imaging (MRI) for those who exhibit pertinent symptoms such as chest pain, shortness of breath, sweating, nausea or vomiting, lightheadedness, and the like, in order to confirm the disease diagnosis.
[0062] The present invention further provides a means for treating patients suffering from the condition or at a heightened risk of developing the condition at a later time. As used herein, treatment of CHD encompasses reducing, reversing, lessening, or eliminating one or more of the symptoms of CHD, as well as preventing or delaying the onset of one or more of the relevant symptoms. Additionally, since certain risk factors for CHD (smoking, high blood pressure, high cholesterol, type 2 diabetes, obesity due to inadequate physical exercise and a poor diet, etc. ) are well known, preventive measures can be prescribed to patients at risk of developing CHD such as cessation of smoking, increasing physical exercise, and adopting a healthy diet, in addition to proper medications for reducing blood pressure and cholesterol level. For individuals who have been deemed to have an increased risk of developing CHD by the method of this invention and who are then diagnosed as actually having already developed CHD (e.g., by conventional diagnostic methods named above and well known in the medical field) , various treatment strategies are available for treating and managing CHD in these patients, including but not limited to, lifestyle changes (relating to diet and exercise) , medications (cholesterol lowering drugs, hypertension drags, beta-blockers, nitroglycerin, calcium channel blockers, etc. ) , coronary interventions (angioplasty and coronary stent) , coronary artery bypass grafting (CABG) , or any combination thereof, as deemed appropriate by the attending physician.VI. Kits and Devices
[0063] The invention provides compositions and kits for practicing the methods described herein to assess the level of DNA methylation within a specified genomic sequence in a subject, which can be used for various purposes such as detecting or diagnosing the presence of heart disease (especially CHD) , determining the risk of developing heart disease (especially CHD) , and monitoring the progression of heart disease (especially CHD) over time in a patient, such as one who has already received a T2D diagnosis.
[0064] Kits for carrying out assays for determining DNA methylation level typically include at least one reagent that is capable of differentially interacting with methylated and unmethylated DNA, including any one that preferentially chemically reacts with, including cleaves, or specifically binds one version (e.g., methylated version) of CpG pairs over the other version (e.g., unmethylated version) of CpG pairs. Some examples include bisulfites, methylation or non-methylation specific endonucleases, and methyl-CpG-binding proteins.
[0065] Kits for carrying out assays for determining DNA methylation level typically also include at least one agent for DNA analysis (which may include sequencing of the pertinent genomic DNA or specific amplification of the methylated or unmethylated version of the pertinent gnnomic DNA) . Where a kit contains an agent of DNA sequencing, it optionally contains a further agent useful for DNA sequence amplification (e.g., PCR) . While DNA sequencing techniques have been well known and in routine practice for decades, the more recently developed next-generation sequencing methods are particularly useful for the method of this invention due to their exceptional capacity of processing a large number of samples and generating a large quantity of information.
[0066] Typically, the kits also include an appropriate standard control. The standard controls indicate the average number of methylated CpG pairs (thus reflecting the average methylation level) in a specific genomic DNA sequence: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143 (any one named in Table 11, especially at the genomic locus 15: 93021943-93022463, 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663) found in healthy subjects not suffering from heart disease (especially CHD) , including those who have already been diagnosed with T2D. In some cases, such standard control may be provided in the form of a set value. In addition, the kits of this invention may provide instruction manuals to guide users in analyzing test samples and assessing the presence, risk, or state of heart disease (especially CHD) in a test subject.
[0067] In a further aspect, the present invention can also be embodied in a device or a system comprising one or more such devices, which is capable of carrying out all or some of the method steps described herein. For instance, in some cases, the device or system performs the following steps upon receiving a biological sample, e.g., a sample of all blood cells isolated from a whole blood sample taken from a subject being tested for detecting heart disease (especially CHD) , assessing the risk of developing heart disease (especially CHD) , or monitored for progression of the condition: (1) determining in sample the number of methylated CloGs or level ofmethylation in a specified genomic DNA sequence; (2) comparing the number or level from step (1) with a standard control value; and (3) providing an output indicating whether heart disease (especially CHD) is present in the subject or whether the subject is at risk of developing heart disease (especially CHD) , or whether there is a change, i.e., worsening or improvement, in the subject′sheart disease (especially CHD) risk or severity. In other cases, the device or system of the invention performs the task of steps (2) and (3) , after step (1) has been performed and the number of methylated CpGs or DNA methylation level from (1) has been entered into the device. Preferably, the device or system is partially or fully automated.EXAMPLES
[0068] The following examples are provided by way of illustration only and not by way of limitation. Those of skill in the art will readily recognize a variety of non-critical parameters that could be changed or modified to yield essentially the same or similar results.INTRODUCTION
[0069] Coronary heart disease (CHD) in people with type 2 diabetes (T2D) is significantly higher when compared to individuals without diabetes. CHD remains the leading cause of morbidity and mortality in people with T2D. The mechanism by which T2D accelerates CHD remains poorly understood. Indeed, the relationship between tight glycaemic control and the risk of CHD in T2D is complex and multifactorial. While achieving and maintaining tight glycaemic control is important for overall diabetes management and reducing the risk of microvascular complications (such as nephropathy, retinopathy, and neuropathy) , its impact on reducing the risk of CHD has been less clear-cut. For instance, several large clinical trials, such as the ACCORD1, ADVANCE2, and the VADT3, have investigated the influence of intensive glycaemic control on cardiovascular outcomes in people with T2D. These trials have shown mixed results, with some indicating a potential benefit of tight glycaemic control on reducing cardiovascular events, while others have not shown significant reductions in cardiovascular outcomes.
[0070] Factors such as the duration of diabetes, baseline cardiovascular risk, individual patient characteristics, and the presence of comorbidities may influence the relationship between glycaemic control and cardiovascular risk. Additionally, interventions targeting multiple cardiovascular risk factors, including blood pressure control, lipid management, smoking cessation, and lifestyle modifications, may have a more significant impact on reducing the risk of CHD in patients with T2D. Overall, while achieving and maintaining tight glycaemic control is an essential component of diabetes management, its direct impact on reducing the risk of CHD in patients with T2D may be modest thus emphasising the benefit of early detection to improve the understanding of risk prediction.
[0071] Despite the technological advances in genetic mapping and genome wide association studies (GWAS) by international consortia, including the recent Hong Kong Diabetes Register (HKDR) and Hong Kong Diabetes Biobank (HKDB)4, only a few genes have been identified and account for disease susceptibility. The UK Biobank5, EPIC-Interact Consortium6 and the CARDIoGRAM plusC4D comprising ENGAGE and SUMMIT consortia of European and South Asian descent (n=66, 643) found no genetic association of genome-wide significance of CHD in those with T2D when compared to people without T2D7. It is increasingly appreciated from these studies that genetic factors cannot fully explain susceptibility to diabetic kidney disease and a more integrative approach that addresses multiple cardiovascular risk factors is recommended for optimizing cardiovascular outcomes in individuals with T2D8.
[0072] The covalent modification of DNA provides a direct mechanism to control gene expression with previous studies showing a direct correlation for DNA methylation indices with diabetic complications9, 10. The effectiveness of the popular methylation BeadChip arrays relies on their capacity to characterise sites of DNA methylation has been widely embraced by the research community11. Of the approaches for detecting DNA methylation the BeadChip arrays dominate population studies involving people living with diabetes12-15. However, reduced genome coverage and methylation resolution are compromised because of confined probe annotation the current estimates of DNA methylation changes by arrays are likely to be of small effect size16 with limited trait-associated genetic variants identified from much larger genome-wide association studies17. Indeed, recent studies have shown the technical limitations ascribed with BeadChlp arrays18, 19. To date no published CHD risk studies in T2D have used sequencing to define gene methylation for biomarker discovery20. This is particularly important because contiguous genomic regions outside promoters such as the gene body, introns and enhancer elements are important in gene regulation and may represent biomarkers for CHD risk in diabetes21.
[0073] Because CHD is accelerated in people with T2D disease the inventors used sequencing to characterize methylation-based biomarkers from peripheral blood. With the application of sequencing technologies, the inventors show for the frrst time that it is possible to determine the CG methylation landscape showing transcription factor binding sites that uncover network relationships in CHD. The inventors developed an automated methylation enrichment and sequencing pipeline using a standard operating procedure involving conventional liquid handling platform that was configured specifically for large population studies. To this end, participants from the Hong Kong Diabetes Register (HKDR) study were selected based on the presence or absence of T2D and CHD4. Because it is unknown whether a multifactorial network for CHD risk can be resolved by DNA methylation sequencing, the inventors constructed differentially methylated regions (DMRs) that were associated with T2D. The ineventors defined novel methylation-based biomarkers for incident and prevalent CHD groups that are associated with increased risk in T2D.RESEARCH DESIGN AND METHODSStudy design and participants
[0074] The Hong Kong Diabetes Register (HKDR) Study, cnndueted from 1995 to 2014, included >10,000 individuals with diabetes4, established as a quality improvemeut program and outcomes study for Chinese patients with diabetes at the Prince of Wales Hospital in Hong Kong. Participants were referred from various healthcare settings, including hospital-based specialty clinics, community clinics, and general practitioners. Type 2 Diabetes (T2D) diagnosis followed the World Health Organization (WHO) criteria. Exclusions comprised individuals with type 1 diabetes, non-Chinese or unknown nationality, missing diabetes type data, or continuous insulin requirement within one year of diaguosis. In addition to detailed clinical information and a comprehensive assessment of diabetes complications at baseline following the EURODIAB protocol22, participants underwent regular repeat assessments for diabetes complications. Additional data on hospitalisations, prescriptions, health outcomes and biochemical investigations were collected. Follow-up time was calculated as the period from recruitment to the first occurrence of the endpoint, the date of death, or 31 December 2015 -whichever came first23. Non-diabetes subjects were recruited from a community-based health assessment programme24. All participants provided written informed consent upon recruitment, and ethical approval was obtained from the Clinical Research Ethics Committee of the Chinese University of Hong Kong. For the purposes of this study, the presence of coronary heart disease (CHD) and related death was defined as a history of myocardial infarction, unstable angina requiring hospitalization and a history of coronary revascularization procedures25.
[0075] The inventors established a nested case-control longitudinal cohort from a subgroup of participants within the HKDR study who developed incident CHD but were free of overt CHD at the time of reeruitment (DI group, n=58) , along with age-matched individuals with a minimum of 10 years of diabetes duration but no clinical manifestation of CHD (DC group, n=55) . Additionally, the inventors included, for comparison, individuals with previous CHD diagnosis (DP group, n=55) and control subjects without diabetes and CHD (NDC group, n=51) . In order to achieve better matching of the cases and controls, each subject with the outcome was matched for age and sex with a subject within each of the other 3 groups. To minimize the confounding effect of renal dysfunction or proteinuria, individuals with chronic kidney disease, defined as eGFR <60 ml / min / 1.73m2 at baseline, or during follow-up, were excluded from the matching in this study. As far as possible in the matching, the inventors also focused on including subjects with type 2 diabetes who did not have microalbuminuria or macroalbuminutria at baseline. The clinical characteristics of participants in this study are summarised in Table 1.DNA isolation and fragmentation
[0076] Genomic DNA (gDNA) was extracted from venous blood, specifically isolating the white blood cell layer using a phenol-chloroform based method. The purified gDNA (1 μg in 50μl of TE buffer) was fragmented into a median length of 250bp using the Qsonica sonicator (Q800R2) . The sonication was performed in pulse mode, with intervals of 30 seconds on and 30 seconds off, over a duration of 10 minutes at 75%amplitude and the water temperature was maintained at 4 ℃. Samples were quality checked for fragment size and fragment uniformity by capillary electrophoresis on the MultiNA (DNA-500 kit, Shimadzu) .Automated methylation enrichment
[0077] Following sonication, 500ng of fragmented genomic DNA was used for methyl-CpG eurichment18 through an automated workflow described in Figures 3-4. Briefly, the inventors programmed a Tecan Freedom EVO 200 liquid handling robot to conduct automated MBD-based methylation enrichment using components from the MethylMinerTM Kit (ME10025, ThermoFisher) in a 96-well format (the program can be made available upon request) . The main reason for automating Methylminer was to efficiently manage a large number of clinical samples while minimising any technical variability associated with manual handling26. This setup allows for the simultaneous processing of up to 48 samples within a two-day timeframe, including quality control assessments, prior to downstream library preparation and high-throughput sequencing.Methylation sequencing (methyl-seq)
[0078] Methyl-binding domain enrichment sequencing (Methyl-seq) was used to investigate DNA methylation in patients with T2D from the HKDR study categorised into DC, DI and DP groups along with controls without diabetes from the community-based cohort. Following methyl-CpG enrichment, eluted methylated-DNA was quantified, and 10ng of this methylated DNA and matching input DNA was used to generate Illumina sequencing libraries. The inventors used the UltraTM DNA Library Prep Kit and indexes from the Index Primer Set 1 (all New England Biolabs) following the manufacturer's protocol with amplification at the last sample preparation stage performed in ten PCR cycles. Following library preparation and 4-plex pooling, high-throughput sequencing was carried out using the Illumina HiSeq 2500 platform (paired-end 150bp) at Novogene Co (China) .Methyl-seq data mapping and peak calling
[0079] Raw Input-DNA and methyl-seq reads per sample were examined for quality assurance. Briefly, Fastx (version 0.0.13) quality trimmer was used to remove low quality bases from the 3' end of the sequence read at a base quality threshold of 20. After data quality control, the sequenced reads were aligned to the human reference genome (hg38) using BWA-MEM27 with default alignment parameters. Subsequently, duplicate reads were removed, and the sequence alignments for each contig were quantified, resulting in a median of 83 million reads per sample.
[0080] After alignment, peak calling was performed on the BAlM files generated from the methylated DNA reads, which involved comparing profiles between all possible pairs of samples using the MACS peak calling sofiware28 (version 2.1.1) . Only peaks called with a fold-enrichment greater than 4 and P-value < 0.01 were retained from each sample, the selected peaks were merged into a consensus peak set using Bedtools muhiinter tool29. The inventors further filtered consensus peaks to avoid likely false positives by only including those peaks overlapping more than 2 samples and excluded peaks overlapping ENCODE blacklisted reginus30. For the genomic annotation of peaks, the inventors used a custom python script to compute CpG counts within the regions. Next, ChlPseeker31 was utilised to assign genomic information such as region length, identifying the nearest genes and overlap with functional elements from genome assembly: GRCh38.101. Finally, the annotated peaks were quantified per sample (in both methylated DNA and input BAM files) using Bedtools multicov. The resulting count matrix was then filtered for peaks with a mean read coverage of greater than 10 to account for potential artifact regions.Differential methylation analysis
[0081] The count matrix was used as the input file for the R package edgeR (REF) , which employed a Generalized Linear Model (GLM) fitted for biological and technical variation to identify Differentially Methylated Regions (DMRs) . To address inter-sample variations, methylation data was corrected for biological variability caused by differences in age, gender, disease duration and experimental classifiers including comorbidities (hypertension) , smoking, as well as medications, which were also integrated into the GLM. Any technical variances were adjusted by incorporating the enrichment ratios (calculated for each sample) of the “spike” in control-methylated-DNA as part of the modelling. Additionally, to account for differences in cell heterogeneity, the inventors employed a custom reference-based cell-type deconvolution algorithm (implemented in R) originally developed by Houseman et al. in 201232. This method leverages cell-type specific DMRs to estimate cell type proportions. The inventors applied this analysis to their data and corrected the methylation data for variance associated with cell-type markers for three major types of white blood cells (B-cell; CD19, T-cell; CD3D, Monocyte; CD14) . Differentially methylated regions were determined by conducting comparisons between controls and cases groups (edgeR) . The following contrasts were included; DC vs NDC, DI vs NDC, DP vs NDC, DI vs DC, DP vs DC, DI + DP vs DC and DP vs DI, with a significance threshold of P-value < 0.001.
[0082] Next, DMRs identified by edgeR were evaluated using Receiver Operating Characteristics (ROC) analysis, conducted with the ROCR package in R33. This analysis employed a binary classification model for individual DMRs to classify CHD risk in DI and DP individuals. The Area Under the Curve (AUC) was computed to measure the performance of DNA methylation in estimating participants with CHD (DI and / or DP) in comparison to the DC group.
[0083] A higher AUC indicates a better model for identifying individuals from the DI and DP groups. The performance of clinical risk factor modelling for ROC analysis was assessed using the Modem Applied Statistics with S (MASS) package in R. The initial model for DI group classification included age, sex, duration of T2D, smoking (yes / no) , and baseline clinical measurements such as HbA1c, eGFR, UACR, diastolic / systolic blood pressure, BMI, and lipid profiles. The stepAIC() function was employed for covariate selection, which resulted in total cholesterol, LDL-C and baseline eGFR remaining in the model. Consequently, these clinical indices were evaluated for their capability in classifying participants with CHD by ROC analysis.Functional analysis
[0084] The functional analysis of the DMRs following ROC analysis, involved multiple steps. Initially, the inventors extracted the positions of the DMRs in BED format. These positions were then inputed into Unibind34, a web-based tool specializing in mapping direct interactions between transcription factors (TFs) and identified DMRs. This enabled the exploration of potential regulatory connections between the methylated regions and specific TFs, providing insights into the functional implications of the observed DNA methylation changes. In the final step, the inventors conducted pathway analysis using annotated Eusembl gene symbols. This analysis was carried out using the Enrichr platform35, which is a widely recognized tool for functional enrichment analysis. To ensure robustness, the inventors applied a significance threshold of P-value < 0.05 when selecting gene sets derived from the Gene Ontology (GO) biological process (BP) database -2023.RESULTSDNA methylation profiling in T2D groups with or without CtlD
[0085] This study was performed using CHD ease-control analyses of participants with T2D recruited for the HKDR study (n=219) . The first group consisted of people living without diabetes and no CHD (NDC group, n=51) . The second group, age-matched diabetes controls (DC) , comprised individuals with a minimum duration of 10 years of T2D without CHD (n=55) . The case group consisted of T2D participants who were free of overt CHD at recruitment but developed CHD during the study and they are referred as the DI group (incident for CHD, n=58) . This group was compared with T2D participants with pre-existing CHD at study recruitment and referred to as the DP group (prevalent for CHD, n=55) . The workflow for DNA methylation profiling in the HKDR study is summarised in Figure la 100. Participants in the T2D groups, ranged from 33 to 73 years of age with an average of 57 ± 7.8 years. The mean duration of diabetes was 5 ± 5.4 years, and the fasting plasma glucose level at the time of sample collection averaged 8.3 ±2.8 mmol / l and HbA1c (%) of 7.4 ± 1.6, consistent with moderate glycaemic control. The clinical characteristics of the HKDR are summarised in Table 1.
[0086] Leukocyte DNA was isolated from case and control groups and characterised for DNA methylation using MBD-capture sequencing (Methyl-seq)18, generating on average 83 million (M) mapped reads per sample and totalling 20 billion reads (Figure 1a 100) . Peak calling identified 842, 478 methylated regions, with a mean length of 835bp. These regions mapped to 17M from a total of 28.3M CG sites in the human genome36, providing a CG coverage of 60%. Furthermore, CG site coverage was comparable between the four groups. The methylation data was adjusted for age, gender, duration of T2D including other comorbidities such as smoking, medications, cell type heterogeneity and DNA methylation enrichment efficiency. UMAP analyses show clustering of control and case HKDR groups following GLM modelling (Figure 5 500) . Genomic mapping show methylation changes were predominantly intronic and in the genebody (Figure 6a 600) . DNA methylation localisation was also assessed at other genomic features including gene promoters, exons, 3' UTR, 5' UTR, CpG islands (CGIs) , and CpG shores (± lkb from CGIs) . CpG islands and shores accounted for < 4%of differential methylation identified in the HKDR study (Figure 6b 610) .
[0087] Next, the inventors investigated DNA methylation in the development and progression of CHD in T2D. Pairwise comparisons of case and control groups to identify differentially methylated regions (DMRs) was performed using edgeR with a significance threshold of P < 0.001. The inventors' analyses focused on samples from the DI group (incident for CHD: developed CHD during study follow-up) and those with prevalent CHD, i.e., a history of CHD before study recruitment (DP group) . DMRs were defined by comparing T2D groups living with or without CHD at recruitment, as well as with controls without diabetes (Figure 6c 620) . Sequencing analysis identified 4, 734 DMRs in the DIvs DC contrast and a 77%reduction in methylation in the DI group (P < 0.001) . Similarly, the DP vs DC contrast identified 4,031 DMRs, with a 70%reduction in methylation observed in the DP group (P < 0.001) .Identification of DMRs for CHD incident and prevalent groups in T2D
[0088] To assess the role of DNA methylation in identifying incident and prevalent CHD, the inventors compared the DI and DC groups to determine the difference associated with CHD risk, and then examined the DP and DC groups for changes related to disease progression (described in Figure 7a 700) . ROC curve analysis (n=219) of DMRs with participant clinical features were used to classify CHD cases with T2D. Among the clinical risk factors assessed, LDL-C presented the highest AUC score of 0.652 for CHD classification in T2D (Table 3) . The inventors used a threshold of AUC >0.65, including the CG overlap criteria of>4 CG sites per region. Based on these parameters, the inventors identified 597 DMRs that define CHD risk in T2D (DI group) . This contrast comprised 288 regions associated with a gain in methylation and 309 regions with a loss in methylation (Table 4) . Similarly, DMRs that define prevalent CHD (DP vs DC) , indicating advanced disease, in T2D. The inventors mapped 505 DMRs in this contrast, with 278 genomic sequences associated with a gain in methylation and 227 showing a loss of methylation. Furthermore, DMRs from the combined DI and DP vs DC comparison were examined to identify changes associated with any CHD case with T2D, revealing 152 DMRs associated with gain in methylation and 117 with loss of methylation. The results of these analyses show dynamic changes in DNA methylation associated with the risk and advancement of CHD in individuals with T2D.Differential methylation defines CHD networks associated with T2D
[0089] Gene set enrichment analysis (GSEA) was used to understand whether DNA methylation influence on CHD pathways that are associated with T2D (Table 4) . The inventors show highly connected pathways with reduced methylation in the incident (DI) and prevalent (DP) CHD groups involved in cardiac muscle function, ion transport, angiogenesis, cellular signaling, and regulation of neural control of the heart. This information is summarised in Tables 5-10. The inventors show that the HKDR DMRs delineate clinically significant CHD networks, with high-confidence mapping underscoring the polygenic nature of the T2D phenotype37, 38.Distinguishing CHD methylation biomarkers in T2D
[0090] To improve the discriminative performance of DMRs for CHD risk and progression in T2D, the inventors focussed on the methylation indices intersecting from DI, DP, and the combined DI+DP in comparison to the DC group, as shown in Figure lb 110. Additionally, a complementary analysis was conducted to exclude methylation changes associated with T2D development (i.e., all T2D cases vs NDC, Figure 1c 120) . This integrative approach identified 20 methylation biomarkers in the incident (DI) and the prevalent (DP) CHD groups (Table 2) . Moreover, 9 of the 20 methylation biomarkers were located in the gene body (comprising 3 within intronic regions, 3 at promoters, and 3 in other genomic features, Table 11) . The remaining 11 methylation biomarkers overlapped intergenic regions, with 8 regions located within 30kb of a known gene. DNA methylation regulates gene expression by direct modification of CG sites that are positioned at transcription factor binding sites39 (TFBS) . To test this hypothesis, the inventors performed integrative analyses to map the methylation biomarkers with known TFBS (Figure 7b 710) .
[0091] The inventors identified binding sites that are subject to gain in DNA methylation for FOXA1 and CEBPB, as well as binding sites for MAX and MAFK that corresponded with the loss of methylation (Figure 7b 710) . These analyses indicate changes in DNA methylation converge on binding sites of genes that are implicated in the development and progression of CHD in T2D. Next, the inventors used UMAP analysis to evaluate the inter-sample profile of the methylation biomarkers across the DI, DP, and DC groups. The inventors observed distinct clustering of the 20 methylation biomarkers in T2D cases with CHD (DI and DP) compared to those without CHD (DC) (Figure 7c 720) . Interestingly, these biomarkers showed improved separation of samples associated with CHD risk (DI) from those without CHD in individuals with T2D (DC) (Figure 7d 730) . Taken together, these findings indicate that DNA methylation may influence CHD gene networks in T2D. These changes in candidate methylation biomarkers involve biological processes relevant to CHD (Table 2) . Furthermore, distinct methylation patterns associated with CHD risk (DI) and progression (DP) were observed compared to diabetes controls without CHD (DC) , highlighting 14 methylation biomarkers with gain in methylation and 6 with loss in methylation (Figure 8 800) .Diagnostic value of methylation biomarkers for CHD in T2D
[0092] A significant challenge is accurately defining biomarkers that interact with T2D to modify the risk and progression of CHD. To address this, the inventors performed combinatorial ROC analysis to prioritise methylation biomarkers for risk (DI) and progression (DP) of CHD in T2D. This involved integrating the 20 methylation biomarkers (5mC) with clinical risk factors (Cf) which generated 33,649 interactive combinations to rank AUC ranges for biomarker accuracy and clinical utility (Figure 9 900) .
[0093] Figure 2a 200 shows the diagnostic performance of the methylation biomarkers when compared to Cf (LDL-C, total cholesterol, and baseline eGFR) . The AUC score for clinical risk factors (Cf) was 0.69 (P=0.002) for CHD risk and progression in T2D. The top five CHD methylation biomarkers (5mC: CHD2, DACH1, TRAM1L1, AKR1B1 and DMR2) improved the AUC score to 0.91 (P=3.3e-12) . The collective AUC score for 5mC+Cf was 0.92 with an improved accuracy score of 0.90 (P=1.4e-12) . Next, the inventors assessed the accuracy of clinical risk factors obtained from the HKDR along with the five CHD methylation biomarkers using prediction probability analyses to distinguish true negatives (TN) from false positives (FP) and true positives (TP) from false negatives (FN) for risk (DI) and progression (DP) of CHD in T2D. Clinical risk factor accuracy was estimated at 70%for identifying true positives CHD cases (DI + DP) in T2D (Figure 2b 210) . The inclusion ofmethylation biomarkers (5mC) improved the accuracy to 85%of for identifying true positives for CHD in T2D (Figure 2c 220) , while the combined 5mC+Cf accuracy was 93%(Figure 2d 230) . These results indicate that methylation scores enhance the performance and accuracy of detecting CHD risk and progression in T2D, beyond what clinical factors alone can achieve.DISCUSSIONS
[0094] The HKDR study examined 842, 478 methylated regions spanning 17 million CG sites in 219 individuals with T2D, including those with and without CHD. The inventors compared methylation difference among individuals with incident (DI) and prevalent CHD (DP) to those without CHD (DC) , revealing unique methylation signatures associated with CHD risk and progression in T2D. After adjusting for age, gender, T2D duration, smoking and cell type heterogeneity, the inventors observed that a predominant loss ofmethylation is associated with CHD in T2D. The inventors'study builds on identifying numerous DMRs, genes, and pathways previously unexplored in incident CHD. In contrast to previous epigenome-wide studies of CHD, which were conducted in non-diabetic populations, the inventors show distinct methylation changes that are specific to CHD in the context of T2D. This study prioritises incident CHD with a longitudinal clinical follow-up period of 10 years, contrasting with predominantly cross-sectional studies40-43. The inventors also present the largest methylation dataset of its type, representing a powerful method for unbiased DNA methylation profiling which enhanced sensitivity and coverage when compared to array-based technologies44, 45.
[0095] The inventors'objective was to profile biomarkers during the development and progression of CHD, enabling risk prediction and understanding of disease advancement. Methylation sequencing identified 20 methylation biomarkers common to incident and prevalent CHD in T2D. The DNA Methylation signals associated with these biomarkers map to genes involved in cardiac function, angiogenesis, and cellular signalling pathways. These associations are not merely coincidental; as the inventors'biomarker analysis identifies methylation changes at genes implicated in the pathophysiology of CHD in T2D46, underscoring the significance of their findings.
[0096] Furthermore, the inventors show five high-confidence methylation biomarkers, four of these annotate to genes CHD2, DACH1, TRAM1L1, AKR1B1 and one located in an intergenic region (DMR2) . These methylation biomarkers, derived from the incident (DI) and prevalent (DP) groups, respectively, are the most reliable diagnostic indicators of CHD risk and progression in T2D. The combined AUC score and accuracy of these biomarkers was 25%higher for identifying CHD risk and progression in T2D when compared to using clinical risk factors alone. Additionally, integrating methylation biomarkers with clinical risk factors further enhanced the accuracy of CHD risk prediction in individuals with T2D. Previous studies have examined the function of these targets in the context of CHD. For example, a genetic variant in DACH1 (Dachshund Family Transcription Factor 1) was previously linked with familial young-onset diabetes, pre-diabetes, and cardiovascular disease47. Genetic deletion of the DACH1 gene (at exon 2) has been shown to impair arterial development48 and pancreas islet cell development49. In diabetic kidney disease, reduced expression of DACH1 is also associated with poor clinical outcomes50, 51. Furthermore, elevated DACH1 expression was observed at branch points of human coronary arteries, indicative of angiogenesis48, whereas DACH1 overexpression in mice is linked to cardiac remodelling and influencing recovery from myocardial infarction52. In this study, the inventors show a gain of DACH1 methylation was associated with CHD development and progression, thereby representing a novel methylation biomarker for assessing CHD risk in T2D.
[0097] This study also identified reduced gene methylation of AKR1B1 (encoding aldose reductase) , which has been implicated in the pathogenesis of diabetic complications53. AKR1B1 gene polymorphisms have been associated with diabetic nephropathy and retinopathy54. The activity of the AKR1B1 gene activity is linked to cardio-renal complications55, 56, facilitated by the activation of the polyol pathway and the production of advanced glycation end products57. While these studies have established a genetic link between AKR1B1 and diabetic vascular complications, they do not directly address the role of DNA methylation. The inventors report reduced AKR1B1 methylation was associated with CHD risk and progression in T2D, indicating a potential mechanism involving oxidative stress and inflammation pathways. Notably, in the most recent meta-analysis of DNA methylation comprising more than 11K non-diabetic CHD participants, the association of AKR1B1 methylation was not detected, indicating its importance specifically in hyperglycemic conditions associated with T2D58.
[0098] This study contributes to the inventors' understanding of DNA methylation associated with CHD in T2D, identifying potential biomarkers and regulatory mechanisms implicated in CHD development and progression. These findings provide insights that could enhance CHD risk assessment and personalized management strategies for individuals with cardiovascular complications in diabetes.
[0099] All patents, patent applications, and other publications, including GenBank Accession Numbers, cited in this application are incorporated by reference in the entirety for all purposes.References1. Action to Control Cardiovascular Risk in Diabetes Study G, Gerstein HC, Miller ME, Byington RP, Goff DC, Jr., Bigger JT, Buse JB, Cushman WC, Genuth S, Ismail-Beigi F, et al. Effects of intensive glucose lowering in type 2 diabetes. N Engl JMed. 2008; 358: 2545-2559. doi: 10.1056 / NEJMoa08027432. Group AC, Patel A, MacMahon S, Chalmers J, Neal B, Billot L, Woodward M, Matte M, Cooper M, Glasziou P, et al. Intensive blood glucose control and vascular outcomes in patients with type 2 diabetes. NEnglJMed. 2008; 358: 2560-2572. doi: 10.1056 / NEJMoa0802987 3. Duckworth W, Abraira C, Moritz T, Reda D, Emanuele N, Reaven PD, Zieve FJ, Marks J, Davis SN, Hayward R, et al. Glucose control and vascular complications in veterans with type 2 diabetes. N Engl J Med. 2009; 360: 129-139. doi: 10.1056 / NEJMoa08084314. Tam CHT, Lim CKP, Lnk AOY, Shi M, Man Cheung H, Ng ACW, Lee HM, Lau ESH, Fan B, Jiang G, et al. Identification ora Common Variant for Coronary Heart Disease at PDE1A Contributes to Individualized Treatment Goals and Risk Stratification of Cardiovascular Complications in Chinese Patients With Type 2 Diabetes. Diabetes Care. 2023; 46: 1271-1281. doi: 10.2337 / dc22-23315. Fall T, Gustafsson S, Orho-Melander M, Ingelsson E. Genome-wide association study of coronary artery disease among individuals with diabetes: the UK Biobank. Diabetologia. 2018; 61: 2174-2179. doi: 10.1007 / s00125-018-4686-z6. Zhao W, Rasheed A, Tikkanen E, Lee JJ, Butterworth AS, Howson JMM, Assimes TL, Chowdhury R, Orho-Melander M, Damrauer S, et al. Identification of new susceptibility loci for type 2 diabetes and shared etiological pathways with coronary heart disease. Nat Genet. 2017; 49: 1450-1457. doi: 10.1038 / ng. 39437. van Zuydam NR, Ladenvall C, Voight BF, Strawbridge RJ, Fernandez-Tajes J, Rayner NW, Robertson NR, Mahajan A, Vlachopoulou E, Goel A, et al. Genetic Predisposition to Coronary Artery Disease in Type 2 Diabetes Mellitus. Circ Genom Precis Med. 2020; 13: e002769. doi: 10.1161 / CIRCGEN. 119.0027698. American Diabetes Association Professional Practice C. 2. Classification and Diagnosis of Diabetes: Standards of Medical Care in Diabetes-2022. Diabetes Care. 2022; 45: S17-S38. doi: 10.2337 / dc22-S0029. Park J, Guan Y, Sheng X, Gluck C, Seasock MJ, Hakimi AA, Qiu C, Pullman J, Verma A, Li H, et al. Functional methylome analysis of human diabetic kidney disease. JCI Insight. 2019; 4. doi: 10.1172 / jci. insight. 12888610. Bansal A, Balasubramanian S, Dhawan S, Leung A, Chen Z, Natarajan R. Integrative Omics Analyses Reveal Epigenetic Memory in Diabetic Renal Cells Regulating Genes Associated With Kidney Dysfunction. Diabetes. 2020; 69: 2490-2502. doi: 10.2337 / db20-038211. Laird PW. Principles and challenges of genomewide DNA methylation analysis. Nat Rev Genet. 2010; 11: 191-203. doi: 10.1038 / urg273212. Sapienza C, Lee J, Powell J, Erinle O, Yafai F, Reichert J, Siraj ES, Madaio M. DNA methylation profiling identifies epigenetic differences between diabetes patients with ESRD and diabetes patients without nephropathy. Epigenetics. 2011 ; 6: 20-28.13. Bell CG, Teschendorff AE, Rakyan VK, Maxwell AP, Beck S, Savage DA. Genome-wide DNA methylation analysis for diabetic nephropathy in type 1 diabetes mellitus. BMC Med Genomics. 2010; 3: 33. doi: 10.1186 / 1755-8794-3-3314. Stefan M, Zhang W, Concepcion E, Yi Z, Tomer Y. DNA methylation profiles in type 1 diabetes twins point to strong epigenetic effects on etiology. J Autoimmun. 2014; 50: 33-37. doi: 10.1016 / j. jaut. 2013.10.00115. Swan EJ, Maxwell AP, McKnight AJ. Distinct methylation patterns in genes that affect mitochondrial function are associated with kidney disease in blood-derived DNA from individuals with Type 1 diabetes. Diabet Med. 2015; 32: 1110-1115. doi: 10.1111 / dme. 1277516. Paul DS, Teschendorff AE, Dang MA, Lowe R, Hawa MI, Ecker S, Beyan H, Cunningham S, Fouts AR, Ramelius A, et al. Increased DNA methylation variability in type 1 diabetes across three immune effector cell types. Nat Commun. 2016; 7: 13555. doi: 10.1038 / ncomms1355517. Ahlqvist E, van Zuydam NR, Groop LC, McCarthy MI. The genetics of diabetic complications. Nat Rev Nephrol. 2015; 11: 277-287. doi: 10.1038 / nmeph. 2015.3718. Khurana I, Kaipananickal H, Maxwell S, Birkelund S, Syreeni A, Forsblom C, Okabe J, Ziemann M, Kaspi A, Rafehi H, et al. Reduced methylation correlates with diabetic nephropathy risk in type 1 diabetes. J Clin Invest. 2023; 133. doi: 10.1172 / JCI16095919. Khurana I, Howard NJ, Maxwell S, Du Preez A, Kaipananickal H, Breen J, Buckberry S, Okabe J, A1-Hasani K, Nakasatien S, et al. Circulating epigenomic biomarkers correspond with kidney disease susceptibility in high-risk populations with type 2 diabetes mellitus. Diabetes Res Clin Pract. 2023 ; 204: 110918. doi: 10.1016 / j. diabres. 2023.11091820. Pirola L, Balcerczyk A, Tothill RW, Haviv I, Kaspi A, Lunke S, Ziemann M, Karagiannis T, Tonna S, Kowalczyk A, et al. Genome-wide analysis distinguishes hyperglycemia regulated epigenetic signatures of primary vascular cells. Genome Res. 2011; 21: 1601-1615. doi: 10.1101 / gr. 116095.11021. Keating ST, Plutzky J, El-Osta A. Epigenetic Changes in Diabetes and Cardiovascular Risk. Circ Res. 2016; 118: 1706-1722. doi: 10.1161 / CIRCRESAHA. 116.30681922. Porta M, Sjoelie AK, Chaturvedi N, Stevens L, Rottiers R, Veglio M, Fuller JH, Group EPCS. Risk factors for progression to proliferative diabetic retinopathy in the EURODIAB Prospective Complications Study. Diabetologia. 2001; 44: 2203-2209. doi: 10.1007 / s00125010003023. Chan JC, So W, Ma RC, Tong PC, Wong R, Yang X. The Complexity of Vascular and Non-Vascular Complications of Diabetes: The Hong Kong Diabetes Registry. Curr Cardiovasc Risk Rep. 2011; 5: 230-239. doi: 10.1007 / s12170-011-0172-624. Zhang Y, Luk AOY, Chow E, Ko GTC, Chan MHM, Ng M, Kong APS, Ma RCW, Ozaki R, So WY, et al. High risk of conversion to diabetes in first-degree relatives of individuals with young-onset type 2 diabetes: a 12-year follow-up analysis. DiabetMed. 2017; 34: 1701-1709. doi: 10.1111 / dme. 1351625. Yang X, So WY, Kong AP, Ma RC, Ko GT, Ho CS, Lam CW, Cockram CS, Chan JC, Tong PC. Development and validation of a total coronary heart disease risk score in type 2 diabetes mellitus. Am J Cardiol. 2008; 101: 596-601. doi: 10.1016 / j. amjcard. 2007.10.01926. Aberg KA, Xie L, Chan RF, Zhao M, Pandey AK, Kumar G, Clark SL, van den Oord EJ. Evaluation of Methyl-Binding Domain Based Enrichment Approaches Revisited. PLoS One. 2015; 10: e0132205. doi: 10.1371 / journal. pone. 013220527. Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009; 25: 1754-1760. doi: 10.1093 / bioinformatics / btp32428. Zhang Y, Liu T, Meyer CA, Eeckhoute J, Johnson DS, Bernstein BE, Nusbaum C, Myers RM, Brown M, Li W, et al. Model-based analysis of ChIP-Seq (MACS) . Genome Biol. 2008; 9: R137. doi: 10.1186 / gb-2008-9-9-r13729. Quinlan AR. BEDTools: The Swiss-Army Tool for Genome Feature Analysis. Curr Protoc Bioinformatics. 2014; 47: 11 12 11-34. doi: 10.1002 / 0471250953. bi1112s4730. Amemiya HM, Kundaje A, Boyle AP. The ENCODE Blacklist: Identification of Problematic Regions of the Genome. Sci Rep. 2019; 9: 9354. doi: 10.1038 / s41598-019-45839-z31. Yu G, Wang LG, He QY. ChiPseeker: an R / Bioconductor package for ChIP peak annotation, comparison and visualization. Bioinformatics. 2015; 31: 2382-2383. doi: 10.1093 / bioinformatics / btv14532. Houseman EA, Accomando WP, Koestler DC, Christensen BC, Marsit CJ, Nelson HH, Wiencke JK, Kelsey KT. DNA methylation arrays as surrogate measures of cell mixture distribution. BMC Bioinformatics. 2012; 13: 86. doi: 10.1186 / 1471-2105-13-8633. Sing T, Sander O, Beerenwinkel N, Lengauer T. ROCR: visualizing classifier performance in R. Bioinformatics. 2005; 21: 3940-3941. doi: 10.1093 / bioinformatics / bti62334. Gheorghe M, Sandve GK, Khan A, Cheneby J, Ballester B, Mathelier A. A map of direct TF-DNA interactions in the hman genome. Nucleic Acids Res. 2019; 47: e21. doi: 10.1093 / nar / gky121035. Chen EY, Tan CM, Kou Y, Duan Q, Wang z, Meirelles GV, Clark NR, Ma′ayan A. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinformatics. 2013; 14: 128. doi: 10.1186 / 1471-2105-14-12836. Babenko VN, Chadaeva IV, Orlov YL. Genomic landscape of CpG rich elements in hman. BMC Evol Biol. 2017; 17: 19. doi: 10.1186 / sl 2862-016-0864-037. Goodarzi MO, Rotter JI. Genetics Insights in the Relationship Between Type 2 Diabetes and Coronary Heart Disease. Circ Res. 2020; 126: 1526-1548. doi: 10.1161 / CIRCRESAHA. 119.31606538. Zarkasi KA, Abdul Murad NA, Ahmad N, Jamal R, Abdullah N. Coronary Heart Disease in Type 2 Diabetes Mellitus: Genetic Factors and Their Mechanisms, Gene-Gene, and Gene-Environment Interactions in the Asian Populations. Int J Environ Res Public Health. 2022; 19. doi: 10.3390 / ijerph1902064739. Ahmed SAH, Ansari SA, Mensah-Brown EPK, Emerald BS. The role of DNA methylation in the pathogenesis of type 2 diabetes mellitus. Clin Epigenetics. 2020; 12: 104. doi: 10.1186 / s13148-020-00896-440. Sharma P, Garg G, Kumar A, Mohammad F, Kumar SR, Tanwar VS, Sati S, Sharma A, Karthikeyan G, Brahrnachari V, et al. Genome wide DNA methylation profiling for epigenetic alteration in coronary artery disease patients. Gene. 2014; 541: 31-40. doi: 10.1016 / j. gene. 2014.02.03441. Westerman K, Sebastiani P, Jacques P, Liu S, DeMeo D, Ordovas JM. DNA methylation modules associate with incident cardiovascular disease and cumulative risk factor exposure. Clin Epigenetics. 2019; 11: 142. doi: 10.1186 / s13148-019-0705-242. Si J, Yang S, Sun D, Yu C, Guo Y, Lin Y, Millwood IY, Walters RG, Yang L, Chen Y, et al. Epigenome-wide analysis of DNA methylation and coronary heart disease: a nested case-control study. Elife. 2021; 10. doi: 10.7554 / eLife. 6867143. Navas-Acien A, Domingo-Relloso A, Subedi P, Riffo-Campos AL, Xia R, Gomez L, Haack K, Goldsmith J, Howard BV, Best LG, et al. Blood DNA Methylation and Incident Coronary Heart Disease: Evidence From the Strong Heart Study. JAMA Cardiol. 2021 ; 6: 1237-1246. doi: 10.1001 / jamacardio. 2021.270444. De Meyer T, Bady P, Trooskens G, Kurscheid S, Bloch J, Kros JM, Halnfellner JA, Stupp R, Delorenzi M, Hegi ME, et al. Genome-wide DNA methylation detection by MethylCap-seq and Infinium HumanMethylation450 BeadChips: an independent large-scale comparison. Sci Rep. 2015; 5: 15375. doi: 10.1038 / srep1537545. Shu C, Zhang X, Aouizerat BE, Xu K. Comparison of methylation capture sequencing and Infinium MethylationEPIC array in peripheral blood mononuclear cells. Epigenetics Chromatin. 2020; 13: 51. doi: 10.1186 / sl 3072-020-00372-646. De Rosa S, Arcidiacono B, Chiefari E, Brunetti A, Indolfi C, Foti DP. Type 2 Diabetes Mellitus and Cardiovascular Disease: Genetic and Epigenetic Links. Front Endocrinol (Lausanne) . 2018; 9: 2. doi: 10.3389 / fendo. 2018.0000247. Ma RC, Lee HM, Lam VK, Tam CH, Ho JS, Zhao HL, Guan J, Kong AP, Lau E, Zhang G, et al. Familial young-onset diabetes, pre-diabetes and cardiovascular disease are associated with genetic variants of DACH1 in Chinese. PLoS One. 2014; 9: e84770. doi: 10.1371 / journal. pone. 008477048. Chang AH, Raftrey BC, D′Amato G, Surya VN, Poduri A, Chen HI, Goldstone AB, Woo J, Fuller GG, Dunn AR, et al. DACH1 stimulates shear stress-guided endothelial cell migration and coronary artery growth through the CXCL12-CXCR4 signaling axis. Genes Dev. 2017; 31: 1308-1324. doi: 10.1101 / gad. 301549.11749. Yang L, Webb SE, Jin N, Lee HM, Chan TF, Xu G, Chan JC, Miller AL, Ma RC. Investigating the role of dachshund b in the development of the pancreatic islet in zebrafish. J Diabetes Investig. 2021; 12: 710-727. doi: 10.1111 / jdi. 1350350. Cao A, Li J, Asadi M, Basgen JM, Zhu B, Yi Z, Jiang S, Doke T, El Shamy O, Patel N, et al. DACH 1 protects podocytes from experimental diabetic injury and modulates PTIP-H3K4Me3 activity. J Clin Invest. 2021; 131. doi: 10.1172 / JCI14127951. Doke T, Huang S, Qin C, Lin H, Guan Y, Hu H, Ma Z, Wu J, Miao Z, Sheng X, et al. Transcriptome-wide association analysis identifies DACH1 as a kidney disease risk gene that contributes to fibrosis. J Clin Invest. 2021; 131. doi: 10.1172 / JCI14180152. Raftrey B, Williams M, Rios Coronado PE, Fan X, Chang AH, Zhao M, Roth R, Trimm E, Racelis R, D′Amato G, et al. Dach1 Extends Artery Networks and Protects Against Cardiac Injury. Circ Res. 2021; 129: 702-716. doi: 10.1161 / CIRCRESAHA. 120.31827153. Tang WH, Martin KA, Hwa J. Aldose reductase, oxidative stress, and diabetic mellitus. Front Pharmacol. 2012; 3: 87. doi: 10.3389 / fphar. 2012.0008754. Wang Y, Ng MC, Lee SC, So WY, Tong PC, Cockram CS, Critchley JA, Chan JC. Phenotypic heterogeneity and associations of two aldose reductase gene polymorphisms with nephropathy and retinopathy in type 2 diabetes. Diabetes Care. 2003; 26: 2410-2415. doi: 10.2337 / diacare. 26.8.241055. So WY, Wang Y, Ng MC, Yang X, Ma RC, Lam V, Kong AP, Tong PC, Chan JC. Aldose reductase genotypes and cardiorenal complications: an 8-year prospective analysis of 1, 074 type 2 diabetic patients. Diabetes Care. 2008; 31: 2148-2153. doi: 10.2337 / dc08-071256. Kaneko M, Bucciarelli L, Hwang YC, Lee L, Yan SF, Schmidt AM, Ramasamy R. Aldose reductase and AGE-RAGE pathways: key players in myocardial ischemic injury. Ann NY4cadSci. 2005; 1043: 702-709. doi: 10.1196 / annals. 1333.08157. Vedantham S, Ananthakrishnan R, Schrnidt AM, Ramasamy R. Aldose reductase, oxidative stress and diabetic cardiovascular complications. Cardiovasc Hematol Agents Med Chem. 2012; 10: 234-240. doi: 10.2174 / 18715251280265109758. Agha G, Mendelson MM, Ward-Caviness CK, Joehanes R, Huan T, Gondalia R, Salfati E, Brody JA, Fiorito G, Bressler J, et al. Blood Leukocyte DNA Methylation Predicts Risk of Future Myocardial Infarction and Coronary Heart Disease. Circulation. 2019; 140: 645-657.doi: 10.1161 / CIRCULATIONAHA. 118.03935759. Qi L, Teschendorff AE. Cell-type heterogeneity: Why we should adjust for it in epigenome and biomarker studies. Clin Epigenetics. 2022; 14: 31. doi: 10.1186 / s13148-022-01253-360. Qu Y, Luo J. Estimation of group means when adjusting for covariates in generalized linear models. Pharm Stat. 2015; 14: 56-62. doi: 10.1002 / pst. 165861. Fernandez-Sanles A, Sayols-Baixeras S, Subirana I, Senti M, Perez-Fernandez S, de Castro Moura M, Esteller M, Marrugat J, Elosua R. DNA methylation biomarkers of myocardial infarction and cardiovascular disease. Clin Epigenetics. 2021; 13: 86. doi: 10.1186 / s13148-021-01078-662. Ho FK, Gray SR, Welsh P, Gill JMR, Sattar N, Pell JP, Celis-Morales C. Ethnic differences in cardiovascular risk: examining differential exposure and susceptibility to risk factors. BMCMed. 2022; 20: 149. doi: 10.1186 / s12916-022-02337-w63. Ke C, Stukel TA, Thiruchelvam D, Shah BR. Ethnic differences in the association between age at diagnosis of diabetes and the risk of cardiovascular complications: a population-based cohort study. Cardiovasc Diabetol. 2023 ; 22: 241. doi: 10.1186 / s12933-023 -01951-zTable 1Clinical characteristics of participants recruited for DNA methylation profiling Data are reported as mean ±standard deviation.*HbA1c (%) was not used as a diagnostic criterion for diabetes, and therefore, it was not measured in individuals without diabetes.Smoking has been recorded as a categorical variable, Current, Former, and Non-smoker.eGFR; estimated glomerular filtration rate calculated with the CKD-EPI formulaACR; Albumin / Creatinine ratioACEi; Angiotensin-converting enzyme inhibitorsARB; Angiotensin receptor blockersTable 2Candidate diagnostic methylation biomarkers for CHD risk The table presents the 20-candidate methylation biomarkers associated with incident and prevalent Coronary Heart Disease (CHD) .For each biomarker, the Area Under the Curve (AUC) per contrast and functional annotation is provided, based on the closest gene annotated to the DMRs.Table is ranked based on highest to lowest AUC in contrasts: DI vs DC, DP vs DC, and DI+DP vs DC.Table 3Performance of clinical risk factors classifying CHD The area under the curve (AUCs) and accuracy (ACC) metrics for clinical risk factors computed to measure the performance of estimating participants with CHD (DI and / or DP) in comparison to the DC group.This analysis was conducted with the ROCR package in R (PMID: 16096348) for binary classification of DI and DP (Yes) versus DC groups (No) . The highest AUC score of 0.652 was observed for LDL, and based on this, a threshold of >0.65 was set to distinguish methylation biomarkers.Table 4Summary of DMRs distinguishing diabetes and CHD risk CHD; Coronary heart diseaseDI = incident CHD in T2DDP = prevalent (history) of CHD in T2DDC =T2D without CHD after 10 years of disease duration;NDC=control subjects without CHD and diabetes.Number of individual DMRs with an AUC > 0.65 and CpG count per region > 4, filtered after ROC analysis with binary classification of DI and DP versus DC group.Table 5Reactome Pathways associated with Gain-DNAm in DI vs DC (P-value <0.05) . Table 6Reactome Pathways associated with Loss-DNAm in DI vs DC (P-value <0.05) . Table 7Reactome Pathways associated with Gain-DNAm in DP vs DC (P-value <0.05) . Table 8Reactome Pathways associated with Loss-DNAm in DP vs DC (P-value <0.05) . Table 9Reactome Pathways associated with Gain-DNAm in DP and DI vs DC (P-value <0.05) . Table 10Reactome Pathways associated with Loss-DNAm in DP and DI vs DC (P-value <0.05) . Table 11Candidate Methylation Biomarkers for CHD risk and progression in T2D The table presents the candidate Methyl-biomarkers (atotal of 20) associated with the risk and progression of Coronary Heart Disease (CHD) . For each Methylation biomarker, the Area Under the Curve (AUC) per contrast and functional annotation is provided, based on the closest gene annotated to the DMRs.Table is ranked based on highest to lowest AUC in DI and DP vs DC contrast.Furthermore, the table reports any overlaps between the detected DMRs and the EPIC Beadchip array, with three DMRs showing intersections with the EPIC array. The table also includes a list of overlapping CG loci and their corresponding names.The identification of these methylation biomarkers and their functional annotations contributes to our understanding of the molecular mechanisms underlying CHD and provides potential targets for further research and clinical applications.TABLE 12DNA SEQUENCE OF THE 20 CANDIDATE METHYLATION BIOMARKERS. DNA SEQUENCES WERE OBTAINED FROM THE HUMAN REFERENCE GENOME (GRCH38 / HG38) BUILD. COORDINATES CORRESPOND TO GENOMIC LOCI BASED ON THE UCSC GENOME BROWSER ASSEMBLY.
Claims
1.A method for assessing genomic DNA methylation status in a subject, comprising the steps of:(a) contacting genomic DNA from blood cells taken from the subject with a reagent that differentially interacts with methylated and non-methylated DNA;(b) analyzing at least a portion of a genomic DNA sequence at a genomic locus selected from: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, that harbors a plurality of CpGs; and(c) determining number of methylated CpGs among the plurality of CpGs.2.The method of claim 1, wherein the genomic locus is 15: 93021943-93022463, 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663.3.The method of claim 1, further comprising, prior to step (a) , isolating blood cells from a blood sample taken from the subject and then isolating genomic DNA from the blood cells.4.The method of claim 1, wherein the subject has been diagnosed with type 2 diabetes (T2D) .5.The method of claim 1, wherein step (b) comprises sequencing of the at least a portion of the genomic DNA.6.The method of claim 5, wherein step (b) comprises fragmentation of the genomic DNA prior to the sequencing.7.The method of claim 6, wherein the fragmentation of the genomic DNA comprises sonication of the genomic DNA.8.The method of claim 6, wherein the genomic DNA fragments are about 100 to about 300 nucleotides in length.9.The method of claim 6, wherein step (b) comprises amplification of the methylated genomic DNA prior to the sequencing.10.The method of claim 1, wherein the reagent that differentially interacts with methylated and non-methylated DNA comprises a bisulfite or a protein comprising a methyl-CpG-binding domain (MBD) .11.The method of claim 1, further comprising, after step (c) , comparing the number of methylated CpGs from step (c) with a standard control and identifying the subject as having an increased risk for coronary heart disease (CHD) upon determining the number ofmethylated CpGs from step (c) as less than the standard control.12.The method of claim 11, further comprising repeating steps (a) - (c) using genomic DNA from blood cells taken from the subject at a later time, wherein an increase in the number of methylated CpGs at the later time as compared to the number of methylated CpGs from the original step (a) indicates a lessened risk for CHD, and wherein a decrease indicates a heightened risk for CHD.13.The method of claim 1, when the subject is identified as having an increased risk for CHD, further comprising administering to the subject an antiplatelet drug, a cholesterol-lowering drug, a blood pressure-lowering drug, nitroglycerin, a calcium channel blocker, or a β-blocker.14.The method of claim 1, wherein step (b) comprises analyzing at least a portion of a genomic DNA sequence at a genomic locus selected from Table 11.15.A method for assessing risk for CHD, comprising the steps of:(a) contacting genomic DNA from blood cells taken from two subjects with a reagent that differentially interacts with methylated and non-methylated DNA;(b) analyzing at least a portion of a genomic DNA sequence at a gnnomic locus selected from: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, that comprises a plurality of CpGs;(c) determining number of methylated CpGs among the plurality of CpGs;(d) comparing the number of methylated CpGs from step (c) between the two subjects; and(e) determining the subject who has more methylated CpGs in the portion of the genomic sequence as having a lower risk for CHD compared with the other subject.16.The method of claim 15, wherein the genomic locus is 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663.17.The method of claim 15, further comprising, prior to step (a) , isolating blood cells from a blood sample taken from the subjects and then isolating genomic DNA from the blood cells.18.The method of claim 15, wherein the subjects have been diagnosed with type 2 diabetes (T2D) .19.The method of claim 15, wherein step (b) comprises sequencing of the at least a portion of the genomic DNA.20.The method of claim 19, wherein step (b) comprises fragmentation of the genomic DNA prior to the sequencing.21.The method of claim 20, wherein the fragmentation of the genomic DNA comprises sonication of the genomic DNA.22.The method of claim 20, wherein the genomic DNA fragments are about 100 to about 300 nucleotides in length.23.The method of claim 20, wherein step (b) comprises amplification of the methylated genomic DNA prior to the sequencing.24.The method of claim 15, wherein the reagent that differentially interacts with methylated and non-methylated DNA comprises a bisulfite or a protein comprising a methyl-CpG-binding domain (MBD) .25.The method of claim 15, wherein step (b) comprises analyzing at least a portion of a genomic DNA sequence at a genomic locus selected from Table 11.26.A kit for detecting CHD or assessing CHD risk in a subject, comprising (1) a standard control that provides, in blood cells taken from an average healthy individual without CHD, number of methylated CpGs within a portion of a genomic DNA at a genomic locus selected from: 13: 86824977-86825479, 3: 145957813-145958340, 15: 93021943-93022463, 16: 51276840-51277172, 13: 104486307-104486663, 13: 71548599-71548873, 4: 130161611-130162143, 11: 82724287-82724832, 3: 169596013-169596618, 4: 90260660-90261167, 8: 123042830-123044737, 1: 159952921-159953349, 11: 11963071-11963607, 12: 67295079-67295806, 3: 81460371-81460911, 4: 116801311-116801910, 6: 71169743-71170162, 7: 134400095-134400659, 8: 144994587-144995131, or 1: 230576774-230577143, that comprises a plurality of CpGs; (2) a reagent that differentially interacts with methylated and non-methylated DNA; and (3) an agent useful for analyzing the portion of the genomic DNA.27.The kit of claim 26, wherein the agent in (3) is useful for sequencing the portion of the genomic DNA.28.The kit of claim 26, wherein the reagent that differentially interacts with methylated and non-methylated DNA comprises a bisulfite or a protein comprising a methyl-CpG-binding domain (MBD) .29.The kit of claim 26, further comprising an agent useful for a DNA amplification reaction.30.The kit of claim 27, wherein the DNA amplification reaction is a polymerase chain reaction (PCR) .31.The kit of claim 26, wherein the genomic locus is 15: 93021943-93022463) , 13: 71548599-71548873, 4: 116801311-116801910, 7: 134400095-134400659, or 13: 104486307-104486663.32.The kit of claim 26, wherein the genomic locus is selected from Table 11.
Citation Information
Patent Citations
DNA methylation composition related to death risk of coronary heart disease patients and screening method and application thereof
CN111850108A
Diagnosing, prognosing, and early detection of cancers by DNA methylation profiling
US20110028333A1
Diagnostic markers
US20130084287A1
Variability single nucleotide polymorphisms linking stochastic epigenetic variation and common disease
US20130296182A1
Methods for determining the likelihood of a malignant disease responding to treatment with a pharmaceutical inhibitor
WO2023062115A1